Most people think that sequencing a genome tells us everything about an organism. What inspired your team to look beyond the genome itself, and why is genome annotation still an important scientific challenge?
Our motivation came from years of proteomic and proteogenomic work across parasites including Leishmania, Trypanosoma, Entamoeba and other pathogens at the Institute of Bioinformatics, Bangalore. We kept meeting the same problem: a reference genome never fully captures an organism’s true coding potential, however long it has been sequenced. Annotation relies heavily on computational prediction, so genes get missed, boundaries misassigned, and proteins stay “hypothetical” for lack of evidence, especially in parasites, whose genomes are often divergent or species-specific.
Proteomics adds an experimental layer. Identifying peptides by mass spectrometry and mapping them to the genome lets us validate predicted genes and uncover overlooked coding regions. Sequencing gives the scaffold; protein evidence shows how complete our reading of it is. Annotation is evolving, not a one-time exercise; as technologies advance, hidden layers of genome biology become accessible.
Your study uses a proteogenomic approach to identify previously hidden protein-coding genes. Could you explain what proteogenomics is and how combining proteins with genomic information provides a more complete picture than DNA alone?
Proteogenomics integrates genomics and proteomics to characterise a genome’s coding potential using computational prediction plus experimental evidence. Sequencing gives the blueprint, but its accuracy depends on whether predicted genes are actually expressed. Across several parasites, well-sequenced genomes still carry incomplete annotations, missed genes, wrong boundaries, and “hypothetical” proteins since prediction struggles with proteins that are short, divergent or absent from reference databases.
Proteomics offers an independent check: mass spectrometry identifies peptides in a sample, and mapping these to the genome reveals coding regions prediction missed or refines known boundaries. This has uncovered unannotated genes and corrected gene models, giving a stronger basis for studying protein roles in host-parasite interactions.
DNA shows what an organism can potentially encode; proteins show what is actually produced. Proteogenomics moves us from a predicted genome to an experimentally supported one, asking not “what genes are present?” but “what coding potential can we demonstrate at the protein level?”
Your work revealed new genes and corrected existing gene annotations in Entamoeba histolytica. Why are these seemingly small improvements so important for understanding the biology and disease-causing mechanisms of this parasite?
Annotation underlies nearly every downstream analysis, so small errors distort our understanding of pathways and host–pathogen interactions. We’ve seen this recur across parasites: initial annotations rarely give a complete picture. A missing gene becomes invisible downstream; an incorrect boundary produces a wrong protein sequence, skewing predicted function and localisation.
In our E. histolytica study, combining mass spectrometry with genomic data identified 41 previously unrecognised protein-coding genes and corrected 18 existing gene models, providing evidence that even a well-studied genome holds coding information revealed only by pairing prediction with experiment. Many carried conserved domains hinting at function, some with orthologs in related Entamoeba species.
A more accurate coding repertoire lets researchers examine pathways tied to survival and pathogenicity, and revisit datasets for previously missed signals. A new gene adds a missing puzzle piece; a corrected model ensures researchers study the right protein.
Many newly identified proteins were previously labelled as “hypothetical.” How can confirming that these proteins actually exist change the way scientists study parasite biology and identify potential therapeutic targets?
“Hypothetical protein” reflects a prediction lacking experimental support common across parasite genomes, and heavy in Trypanosoma and Leishmania. When mass spectrometry detects a matching peptide, that protein moves from computational guess to confirmed entity. This doesn’t reveal function immediately, but opens questions: where and when is it expressed, and what role might it play in metabolism or host interaction?
This matters especially for parasites, whose divergent proteins are easily missed by similarity-based annotation. In our E. histolytica study, several new proteins carried conserved domains offering functional clues, moving them toward candidates for functional research.
Confirmation also strengthens target prioritisation: a protein essential for survival or infection becomes a stronger drug-target candidate, though existence alone isn’t proof of suitability. Divergence from human counterparts can open opportunities for selective therapies. The greatest value may be expanding the search space: overlooked hypothetical proteins, once confirmed, become starting points for new biology and drug targets.
Beyond amoebiasis, how could proteogenomic approaches benefit research on other infectious organisms, agricultural pathogens, or even human diseases?
Proteogenomics addresses a universal gap between genomic prediction and experimentally observed proteins, making it valuable for any pathogen with divergent or poorly characterised proteins, including bacteria, fungi and viruses with incomplete annotation.
In agriculture, pathogens rely on protein networks for invasion, colonisation and immune evasion; proteogenomic can identify these and inform surveillance, diagnostics and control strategies.
In human disease, though the genome is well annotated, the proteome is still being understood; variants often reveal functional consequences only at the protein level. This matters in cancer, where alterations produce tumour-specific proteins; integrating genomic and proteomic data can identify these as biomarkers or targets for precision medicine.
What were the biggest technical and computational challenges in integrating large-scale proteomic and genomic datasets, and how is this field continuing to evolve with advances in sequencing and mass spectrometry?
The core challenge is connecting genomic predictions with mass-spectrometry peptide spectra accurately enough to know which signals genuinely indicate coding regions. Proteogenomics searches beyond annotated databases into regions not flagged as coding, raising discovery potential but also false-positive risk. Confidence requires a match to be statistically robust, biologically sensible and to support a new gene, and whether conserved domains support it.
Genome assembly quality is another constraint, as misassembled repetitive regions make even strong proteomic evidence hard to interpret. No single computational approach works universally, requiring careful integration of genomic, proteomic and comparative evidence.
The field is advancing quickly. Long-read sequencing now produces more complete genome assemblies; mass spectrometry sensitivity keeps improving. Integration with transcriptomics adds independent evidence, and when approaches converge, confidence rises.
Looking ahead, how do you envision proteogenomics transforming biological research over the next decade? What exciting discoveries do you think this approach will enable in genome science and precision medicine?
We expect proteogenomics to move from a specialised annotation tool to a core part of biological research. Our parasite work shows substantial biology still hidden between genomic prediction and protein-level evidence. As assemblies and mass spectrometry improve, we’ll probe previously inaccessible regions, revealing new genes and lineage-specific proteins.
Annotation should shift toward an evolving, evidence-based model that is important for infectious disease, where proteins involved in immune evasion or drug resistance are hard to find through comparison alone. Comparative proteogenomics across strains could reveal which proteins distinguish more pathogenic organisms.
For precision medicine, proteins usually translate genetic change into biological effect. Linking genotype to protein sequence and function could be transformative in cancer, offering new biomarkers and targets. Integration with spatial and single-cell biology will reveal not just which proteins exist, but where and when. The deeper shift is conceptual: from “what does the genome contain?” to “which regions are expressed and functional?” The most exciting discoveries ahead will likely come from these unexplored connections – moving genome science from reading genomes to understanding what they do.











