Shorya Consul

Senior Computational Scientist, Quantum-Si

I design algorithms — statistical, Bayesian, and machine learning — that turn raw sequencing data into biological insight, from research prototypes to production pipelines processing 500GB+ datasets.

About

I'm a computational scientist on the Data Science & Algorithms team at Quantum-Si, where I design the statistical, Bayesian, and machine learning algorithms that power our proteomics sequencing platform.

Before Quantum-Si, I completed my PhD in Electrical and Computer Engineering at UT Austin under Prof. Haris Vikalo, where my research focused on applying machine learning — including deep learning and, more recently, genomic foundation models — to genomic and oncoviral sequencing data. I'm also keenly interested in ethical machine learning, particularly privacy.

Outside of work, I'm usually reading a good fantasy or mystery novel, playing tennis or ultimate frisbee, or nerding out over trivia and board games.

Experience

Selected highlights — full role history and metrics in my resume.

Quantum-Si Dec 2024 — Present
Data Science & Algorithms Team · San Diego, CA

I've led the redesign of Quantum-Si's protein-inference pipeline — taking it from a slower, less accurate Bayesian model to a system that's 5× faster and meaningfully more accurate, then pushing it further with a Rust rewrite that now handles 500GB+ runs. I've also built an interactive motif dashboard adopted directly by R&D scientists and customers, and regularly brief C-suite leadership on algorithm strategy.

CognitiveScale Jun — Aug 2019
Machine Learning Team · Austin, TX

Designed the model-audit framework within Cortex Certifai — scoring black-box models on robustness, fairness, explainability, and performance — and won first place in a company-wide Shark Tank for responsible-AI product pitches.

Research & Publications

Click a title to expand its abstract.

NextVir: Enabling classification of tumor-causing viruses with genomic foundation models2025
Robertson, J., Consul, S., Vikalo, H. — PLOS Computational Biology

A genomic foundation model framework — fine-tuning DNABERT-S, Nucleotide Transformer, and HyenaDNA via LoRA — and the first method to jointly identify and classify oncoviral sequencing reads within a single model. Achieves approximately 95% accuracy across seven oncoviral families and human background, outperforming three state-of-the-art baselines on realistic, non-uniform-coverage sequencing data.

XVir: A Transformer-Based Architecture for Identifying Viral Reads from Cancer Samples2025
Consul, S., Robertson, J., Vikalo, H. — Journal of Computational Biology

Many cancers can be linked to viral infections. With the advent of modern sequencing technologies, we now have access to massive amounts of tumor data, which in turn have allowed studies of the associations between viruses and cancers. However, the high diversity of oncoviral families makes detecting viral DNA challenging, thereby complicating the study of these associations. We propose XVir, a transformer-based deep learning architecture to reliably identify viral DNA present in human tumors, achieving high detection accuracy while being more compact and less computationally demanding than competing methods.

XHap: Haplotype Assembly using Long-distance Read Correlations learned by Transformers2023
Consul, S., Ke, Z., Vikalo, H. — Bioinformatics Advances

Reconstructing haplotypes of an organism from a set of sequencing reads is a computationally challenging (NP-hard) problem. At the core of haplotype assembly is the task of clustering reads according to their origin — grouping together reads that sample the same haplotype. Read length limitations and sequencing errors render this problem difficult even for diploids; complexity grows with ploidy. XHap learns correlations between pairs of sequencing reads, including those separated by large genomic distances, by leveraging transformers, and uses these learned correlations to assemble haplotypes. Experiments on semi-experimental and real data show XHap significantly outperforms state-of-the-art techniques in diploid and polyploid haplotype assembly on both short and long sequencing reads.

RF-based Network Inference: Theoretical Foundations2021
Beytur, H.B., Consul, S., de Veciana, G., Vikalo, H. — IEEE MILCOM

We consider RF-based network inference based on channel usage. The proposed approaches rely on distributed spectrum sensing and are agnostic to content and communication protocols. Inference based solely on observing nodes' channel usage is shown to be equivalent to a Boolean matrix decomposition problem, which in general does not have a unique solution and is NP-hard. We provide necessary and sufficient conditions for the decomposition to have a unique solution — i.e., for the network to be recoverable — along with a low-complexity recovery algorithm and an analysis of the observation time required.

Balance is Key: Private Median Splits Yield High-Utility Random Trees2020
Consul, S., Williamson, S. — arXiv preprint

Decision forests are popular for both regression and classification but require a large number of queries for training, making differential privacy especially challenging to attain. We propose DiPriMe forests, a novel scheme that ensures differential privacy while maintaining high utility.

Certifai: A Toolkit for Building Trust in AI Systems2020
Henderson, J., et al. (contributor) — IJCAI-20 Demonstrations Track

Cortex Certifai audits black-box classification and regression models on tabular data across four dimensions — robustness, fairness, explainability, and performance — without requiring access to a model's internals. I designed the regression-model audit framework within Certifai as part of CognitiveScale's Machine Learning team.

Reconstructing Intra-tumor Heterogeneity via Convex Optimization and Branch-and-Bound Search2019
Consul, S., Vikalo, H. — ACM-BCB

Reconstructing tumor populations from heterogeneous samples using high-throughput sequencing data is a highly valuable area of study due to its potential to inform targeted studies and treatments. This is challenging due to complex mutations and read lengths too short to span regions exhibiting structural variations. We present AMTHet, a novel algorithmic framework to infer tumor clonal populations and their frequencies from a heterogeneous sample based on copy number variations.

A MAP Framework for Support Recovery of Sparse Signals Using Orthogonal Least Squares2018
Consul, S., Hashemi, A., Vikalo, H. — ICASSP

We propose MAP-AOLS, an algorithm that leverages statistical information about the sensing matrix and signal to greedily reconstruct sparse binary signals from compressed measurements — in contrast to conventional greedy algorithms, such as OLS and OMP, that reconstruct without using any knowledge of the underlying statistical distributions.

Skills

Languages & Tools
Python Rust C++ MATLAB Bash PyTorch TensorFlow Transformers / LLMs Samtools Docker Cwltool
Domains
Biological Sequence Modeling Representation Learning Self-Supervised Learning Bayesian Inference Computational Biology Genomics

Education

The University of Texas at Austin
PhD, Electrical and Computer Engineering
The University of Texas at Austin
MS, Electrical and Computer Engineering
Indian Institute of Technology Bombay
B.Tech, Electrical Engineering (Minor in Computer Science)