Finding the Genes We Missed: The Next Frontier in Genome Annotation

Published on
September 15, 2026

Institute of Bioinformatics, Bangalore, Karnataka, India

Areas of Expertise
Parasitology, Proteogenomic, Proteomics, Vector Biology

Our motivation came from years of proteomic and proteogenomic work across parasites including Leishmania, Trypanosoma, Entamoeba and other pathogens at the Institute of Bioinformatics, Bangalore. We kept meeting the same problem: a reference genome never fully captures an organism’s true coding potential, however long it has been sequenced. Annotation relies heavily on computational prediction, so genes get missed, boundaries misassigned, and proteins stay “hypothetical” for lack of evidence, especially in parasites, whose genomes are often divergent or species-specific.

Proteomics adds an experimental layer. Identifying peptides by mass spectrometry and mapping them to the genome lets us validate predicted genes and uncover overlooked coding regions. Sequencing gives the scaffold; protein evidence shows how complete our reading of it is. Annotation is evolving, not a one-time exercise; as technologies advance, hidden layers of genome biology become accessible.

Proteogenomics integrates genomics and proteomics to characterise a genome’s coding potential using computational prediction plus experimental evidence. Sequencing gives the blueprint, but its accuracy depends on whether predicted genes are actually expressed. Across several parasites, well-sequenced genomes still carry incomplete annotations, missed genes, wrong boundaries, and “hypothetical” proteins since prediction struggles with proteins that are short, divergent or absent from reference databases.

Proteomics offers an independent check: mass spectrometry identifies peptides in a sample, and mapping these to the genome reveals coding regions prediction missed or refines known boundaries. This has uncovered unannotated genes and corrected gene models, giving a stronger basis for studying protein roles in host-parasite interactions.

DNA shows what an organism can potentially encode; proteins show what is actually produced. Proteogenomics moves us from a predicted genome to an experimentally supported one, asking not “what genes are present?” but “what coding potential can we demonstrate at the protein level?”

Annotation underlies nearly every downstream analysis, so small errors distort our understanding of pathways and host–pathogen interactions. We’ve seen this recur across parasites: initial annotations rarely give a complete picture. A missing gene becomes invisible downstream; an incorrect boundary produces a wrong protein sequence, skewing predicted function and localisation.

In our E. histolytica study, combining mass spectrometry with genomic data identified 41 previously unrecognised protein-coding genes and corrected 18 existing gene models, providing evidence that even a well-studied genome holds coding information revealed only by pairing prediction with experiment. Many carried conserved domains hinting at function, some with orthologs in related Entamoeba species.

A more accurate coding repertoire lets researchers examine pathways tied to survival and pathogenicity, and revisit datasets for previously missed signals. A new gene adds a missing puzzle piece; a corrected model ensures researchers study the right protein.

“Hypothetical protein” reflects a prediction lacking experimental support common across parasite genomes, and heavy in Trypanosoma and Leishmania. When mass spectrometry detects a matching peptide, that protein moves from computational guess to confirmed entity. This doesn’t reveal function immediately, but opens questions: where and when is it expressed, and what role might it play in metabolism or host interaction?

This matters especially for parasites, whose divergent proteins are easily missed by similarity-based annotation. In our E. histolytica study, several new proteins carried conserved domains offering functional clues, moving them toward candidates for functional research.

Confirmation also strengthens target prioritisation: a protein essential for survival or infection becomes a stronger drug-target candidate, though existence alone isn’t proof of suitability. Divergence from human counterparts can open opportunities for selective therapies. The greatest value may be expanding the search space: overlooked hypothetical proteins, once confirmed, become starting points for new biology and drug targets.

Proteogenomics addresses a universal gap between genomic prediction and experimentally observed proteins, making it valuable for any pathogen with divergent or poorly characterised proteins, including bacteria, fungi and viruses with incomplete annotation.

In agriculture, pathogens rely on protein networks for invasion, colonisation and immune evasion; proteogenomic can identify these and inform surveillance, diagnostics and control strategies.

In human disease, though the genome is well annotated, the proteome is still being understood; variants often reveal functional consequences only at the protein level. This matters in cancer, where alterations produce tumour-specific proteins; integrating genomic and proteomic data can identify these as biomarkers or targets for precision medicine.

The core challenge is connecting genomic predictions with mass-spectrometry peptide spectra accurately enough to know which signals genuinely indicate coding regions. Proteogenomics searches beyond annotated databases into regions not flagged as coding, raising discovery potential but also false-positive risk. Confidence requires a match to be statistically robust, biologically sensible and to support a new gene, and whether conserved domains support it.

Genome assembly quality is another constraint, as misassembled repetitive regions make even strong proteomic evidence hard to interpret. No single computational approach works universally, requiring careful integration of genomic, proteomic and comparative evidence.

The field is advancing quickly. Long-read sequencing now produces more complete genome assemblies; mass spectrometry sensitivity keeps improving. Integration with transcriptomics adds independent evidence, and when approaches converge, confidence rises.

We expect proteogenomics to move from a specialised annotation tool to a core part of biological research. Our parasite work shows substantial biology still hidden between genomic prediction and protein-level evidence. As assemblies and mass spectrometry improve, we’ll probe previously inaccessible regions, revealing new genes and lineage-specific proteins.

Annotation should shift toward an evolving, evidence-based model that is important for infectious disease, where proteins involved in immune evasion or drug resistance are hard to find through comparison alone. Comparative proteogenomics across strains could reveal which proteins distinguish more pathogenic organisms.

For precision medicine, proteins usually translate genetic change into biological effect. Linking genotype to protein sequence and function could be transformative in cancer, offering new biomarkers and targets. Integration with spatial and single-cell biology will reveal not just which proteins exist, but where and when. The deeper shift is conceptual: from “what does the genome contain?” to “which regions are expressed and functional?” The most exciting discoveries ahead will likely come from these unexplored connections – moving genome science from reading genomes to understanding what they do.

References

Malshetty MB, Chowdhury S, Pawar S, Jaiwar M, Kumar P, Pawar H. Application of Proteogenomic Approaches for Refinement of the Entamoeba histolytica Reference Genome Using High Resolution Mass-Spectrometry Data. Acta Parasitologica. 2026 Aug;71(4):166.
Article DOI

Chowdhury S, Pawar S, Kumar P, Vasudevan K, Jamdhade M, Pawar H. Applying a proteogenomic approach for improving genome annotation in Leishmania panamensis using high-resolution mass spectrometry data. Acta Tropica. 2026 Aug 1:108267.
Article DOI

Chowdhury S, Vasudevan K, Pawar S, Jaikumar M, Mohapatra P, Tupperwar N, Pawar H. Mapping the hidden secretome in Leishmania parasites using a proteogenomics approach. Journal of Parasitic Diseases. 2026 Jun 23:1-6.
Article DOI

Chowdhury S, Pawar S, Vasudevan K, Mishra N, Jamdhade M, Pawar H. Identification of new protein-coding potential in Leishmania donovani using a proteogenomics approach. Current Microbiology. 2026 Oct;83(10):507.
Article DOI

Science Factors.

Beyond Pneumonia: Recognising the Hidden Danger of Leptospirosis

0
What is leptospirosis, how is it transmitted, and why is it an important public health concern? Leptospirosis is a bacterial infection caused by Leptospira species....

How Do Plants Survive Stress? The Science Behind Stronger Crops

0
Plants cannot move away from heat, drought or salty soil. How do they protect themselves when conditions become difficult? Plants are remarkably adaptable organisms. Unlike...

Tiny Plastics, Big Consequences: What Fish Reveal About Freshwater Pollution

0
Microplastics have become a growing environmental concern worldwide. What inspired your team to investigate their presence in the Golden Mahseer, and why is this...

Shining Light on Artificial Enzymes: How Supramolecular Science Makes Aqueous Organocatalysis Switchable

0
What inspired your team to develop a light-switchable artificial enzyme, and what scientific challenge were you aiming to address? Our inspiration came from nature’s remarkable...

GABARAPL2 and Alix mediate reciprocal regulation of autophagy and exosome pathways to facilitate cellular homeostasis

0
Cancer cells often survive treatments that would normally kill healthy cells. What makes cancer cells so resilient, and why is understanding their survival mechanisms...

Engineering Gold Nanoparticles for Smarter Blood Typing

0
Blood transfusions save millions of lives every year, yet ensuring the right blood match can still be challenging. What inspired your team to develop...

Engineering Peptide Nanofibrils to Outsmart Superbugs-Toward Targeted Antibacterial Strategies for Drug-Resistant Infections

0
What inspired your team to explore self-assembling peptide nanofibrils as a new strategy to combat drug-resistant bacteria? The rapid spread of antibiotic-resistant bacteria has made...

Can Genetic Testing Predict Who Will Respond to Leukemia Treatment? New Insights into Chronic Myeloid Leukemia

0
What inspired you to study why some patients with chronic myeloid leukemia respond well to treatment while others develop drug resistance? There are two unanswered...