Senior Computational Scientist, Quantum-Si
I design algorithms — statistical, Bayesian, and machine learning — that turn raw sequencing data into biological insight, from research prototypes to production pipelines processing 500GB+ datasets.
About
I'm a computational scientist on the Data Science & Algorithms team at Quantum-Si, where I design the statistical, Bayesian, and machine learning algorithms that power our proteomics sequencing platform.
Before Quantum-Si, I completed my PhD in Electrical and Computer Engineering at UT Austin under Prof. Haris Vikalo, where my research focused on applying machine learning — including deep learning and, more recently, genomic foundation models — to genomic and oncoviral sequencing data. I'm also keenly interested in ethical machine learning, particularly privacy.
Outside of work, I'm usually reading a good fantasy or mystery novel, playing tennis or ultimate frisbee, or nerding out over trivia and board games.
Experience
Selected highlights — full role history and metrics in my resume.
I've led the redesign of Quantum-Si's protein-inference pipeline — taking it from a slower, less accurate Bayesian model to a system that's 5× faster and meaningfully more accurate, then pushing it further with a Rust rewrite that now handles 500GB+ runs. I've also built an interactive motif dashboard adopted directly by R&D scientists and customers, and regularly brief C-suite leadership on algorithm strategy.
Designed the model-audit framework within Cortex Certifai — scoring black-box models on robustness, fairness, explainability, and performance — and won first place in a company-wide Shark Tank for responsible-AI product pitches.
Research & Publications
Click a title to expand its abstract.
A genomic foundation model framework — fine-tuning DNABERT-S, Nucleotide Transformer, and HyenaDNA via LoRA — and the first method to jointly identify and classify oncoviral sequencing reads within a single model. Achieves approximately 95% accuracy across seven oncoviral families and human background, outperforming three state-of-the-art baselines on realistic, non-uniform-coverage sequencing data.
Many cancers can be linked to viral infections. With the advent of modern sequencing technologies, we now have access to massive amounts of tumor data, which in turn have allowed studies of the associations between viruses and cancers. However, the high diversity of oncoviral families makes detecting viral DNA challenging, thereby complicating the study of these associations. We propose XVir, a transformer-based deep learning architecture to reliably identify viral DNA present in human tumors, achieving high detection accuracy while being more compact and less computationally demanding than competing methods.
Reconstructing haplotypes of an organism from a set of sequencing reads is a computationally challenging (NP-hard) problem. At the core of haplotype assembly is the task of clustering reads according to their origin — grouping together reads that sample the same haplotype. Read length limitations and sequencing errors render this problem difficult even for diploids; complexity grows with ploidy. XHap learns correlations between pairs of sequencing reads, including those separated by large genomic distances, by leveraging transformers, and uses these learned correlations to assemble haplotypes. Experiments on semi-experimental and real data show XHap significantly outperforms state-of-the-art techniques in diploid and polyploid haplotype assembly on both short and long sequencing reads.
We consider RF-based network inference based on channel usage. The proposed approaches rely on distributed spectrum sensing and are agnostic to content and communication protocols. Inference based solely on observing nodes' channel usage is shown to be equivalent to a Boolean matrix decomposition problem, which in general does not have a unique solution and is NP-hard. We provide necessary and sufficient conditions for the decomposition to have a unique solution — i.e., for the network to be recoverable — along with a low-complexity recovery algorithm and an analysis of the observation time required.
Decision forests are popular for both regression and classification but require a large number of queries for training, making differential privacy especially challenging to attain. We propose DiPriMe forests, a novel scheme that ensures differential privacy while maintaining high utility.
Cortex Certifai audits black-box classification and regression models on tabular data across four dimensions — robustness, fairness, explainability, and performance — without requiring access to a model's internals. I designed the regression-model audit framework within Certifai as part of CognitiveScale's Machine Learning team.
Reconstructing tumor populations from heterogeneous samples using high-throughput sequencing data is a highly valuable area of study due to its potential to inform targeted studies and treatments. This is challenging due to complex mutations and read lengths too short to span regions exhibiting structural variations. We present AMTHet, a novel algorithmic framework to infer tumor clonal populations and their frequencies from a heterogeneous sample based on copy number variations.
We propose MAP-AOLS, an algorithm that leverages statistical information about the sensing matrix and signal to greedily reconstruct sparse binary signals from compressed measurements — in contrast to conventional greedy algorithms, such as OLS and OMP, that reconstruct without using any knowledge of the underlying statistical distributions.
Skills
Education