BDH-Based Bidirectional Protein Model
Protein modelling is useful when it helps a team decide what to inspect next. Our BDH-based direction connects sequence and structure representations in both directions, while keeping conformational uncertainty and downstream validation visible.
Two directions, one review loop
Sequence information can constrain plausible structural states, while structural context can expose sequence changes worth evaluating. Treating the two representations as separate one-way tasks leaves useful agreement and disagreement on the table.
The physical link is indirect: a sequence defines a chain of residues, but the observed structure also depends on solvent, ionic conditions, temperature, ligands, oligomeric state, and conformational dynamics. A representation can narrow plausible states without uniquely determining every state present in solution.
A bidirectional model is intended to surface compatible states and task-specific scores. It is not a claim that a single prediction captures every dynamic state a protein can occupy.
The output is a hypothesis
The useful output is a ranked set of states, residues, or candidate changes with enough context for a scientist to review. Confidence, input provenance, and the limits of the representation should remain alongside the score.
Higher-cost structure refinement, molecular dynamics, quantum chemistry, or an experiment can then focus on a smaller and better-justified set of hypotheses.
For a mutation or design proposal, the review should still consider residue environment, hydrogen-bond networks, electrostatics, packing, solvent exposure, and any known ligand or partner interactions. A high model score alone is not evidence that a protein will fold, bind, or remain stable.
What we are evaluating
The development work is centred on reproducible inputs, clear task definitions, held-out evaluation, and failure analysis. Progress is measured by whether the model improves the next decision, not by a single headline metric.
Evaluation should separate memorisation from generalisation. Families or near-duplicate structures should not leak across train and test splits, and performance should be reported by task and protein class. A useful failure report explains where the representation loses information and which follow-up calculation or experiment is needed.