Generative models in structural biology are easy to make impressive and hard to make trustworthy. The gap between the two is method validation — the unglamorous work of proving a system can recognize what's already known before letting it speak about what isn't. This is how TriaX is held to account.
A result is only as trustworthy as the proof that the method could have found the opposite.
Six commitments govern every prediction we make — controls before claims, geometry before compute, and calibration before confidence.
We validate against experimentally-characterized proximity complexes before trusting any prediction. The logic is unforgiving in the right direction: a method that can't recognize a known interface can't be trusted to rule one out. Recovering an interface that crystallography or cryo-EM has already settled is not a demonstration of cleverness — it is the minimum entry fee for having an opinion about the unknown.
This makes falsifiability a design requirement rather than a virtue we claim after the fact. A pipeline tuned until it produces attractive answers has no way to distinguish a discovery from a preference. We would rather know that a method fails on a control than discover later that it never could have failed at all.
Before committing expensive compute, a fast geometric screen asks whether a proposed proximity is even physically possible — so we never mistake a geometric impossibility for a chemical one. Without that gate, a negative result is ambiguous in the worst way: the model may have rejected a molecule on chemical grounds, or the arrangement may simply never have been buildable in three dimensions. Those two failures demand completely different responses, and a pipeline that can't tell them apart will keep drawing the wrong lesson.
The screen encodes a substantive point about how proximity therapeutics work, not just a shortcut. A productive molecular glue must present an exposed, partner-bindable face; proximity buried in an enclosed pocket isn't productive proximity. A molecule can bind its target with excellent affinity and still be useless as a glue, because nothing is left facing outward for a second protein to engage. We screen for that exposure up front, where it costs almost nothing to ask.
Every score is calibrated against reference complexes and reported with the settings that produced it. A number without its calibration isn't evidence — it's a figure with a unit attached, and it cannot tell you whether the value in front of you is remarkable or unexceptional.
So scores travel with their provenance. Two predictions are comparable only when they were produced under identical conditions, and we treat any comparison across differing settings as invalid until re-run. The purpose is to make a score decision-grade: something a scientist can weigh against an experiment and act on, rather than a ranking that looks precise and means little.
This is the FLEXIBILITY axis of the model, stated as a methodological commitment. Interfaces that exist only in a minor conformational state are invisible to single-structure methods. A protein's deposited structure is one frame from a moving ensemble, and the pose that opens a druggable interface is often not that frame — it's a transient or lightly-populated state that never appears in the picture.
Which is why modeling the full ensemble isn't a refinement — it's the difference between finding an interface and never seeing it. Treating flexibility as a late-stage correction applied to a rigid answer inverts the problem: by then the search has already been run in a space that excluded the interface worth finding. A rigid-structure method that reports nothing has not shown that nothing is there.
We distinguish “not found under current assumptions” from “impossible,” and we name what a model hasn't yet tested. The two get conflated constantly, and the conflation is expensive — a programme abandoned because a search came up empty is a different decision from one abandoned because the chemistry rules it out.
Stating the limits of a method is what makes its positive results credible. A system described as universally capable gives you no way to calibrate the claims it makes, because nothing it reports could ever count against it. Marking the edges of what has been examined is not a hedge; it is what makes the interior worth believing.
A claim clears a fixed bar before we make it. The bar has three parts, and the standard is fixed in advance so that it constrains conclusions rather than adapting to them.
Exhaustive coverage — the search space was examined completely, not sampled until something promising appeared. Controls passing under identical settings — the positive controls held under the same configuration that produced the result, not a more permissive one. Explicit attribution — we can say what is driving the result, because a correct prediction obtained for the wrong reason will not survive contact with the next target.
A claim that clears all three is one we're prepared to defend. Anything that doesn't stays internal until it does.
Six commitments, applied to every prediction before it leaves the building.
Validate against known proximity complexes first. A method that can't recognize a known interface can't be trusted to rule one out.
Screen physical possibility up front, so a geometric impossibility is never mistaken for a chemical one.
Scores are calibrated against reference complexes and reported with the settings that produced them.
Minor conformational states are where many interfaces live. Single-structure methods cannot see them at all.
“Not found under current assumptions” is not “impossible.” We name what hasn't been tested.
Exhaustive coverage, controls passing under identical settings, and explicit attribution of what drives the result.
If you evaluate computational methods for a living, the questions you'd want answered are the ones we'd rather discuss early. Tell us the interaction you care about and the standard of evidence you hold methods to.
Talk to the team →