Research

Published,
with the work attached.

Every result Omic reports comes with its sources, its code and what it does not show. Benchmarks, technical reports and case studies, as they are released.

technical report

Source assertions grouping into propositions, with two opposed predicates over the same concept pair.

Attest: a claim-native data layer for shared knowledge

Organizations share documents rather than facts, so every reader re-derives what a corpus says. Attest stores one row per source assertion, which makes corroboration a group-by and disagreement a first-class record. Measured on 150 biomedical concept pairs and a 50-paper corpus, with the costs reported.

technical report

Detection status for twelve archaeal compound classes and the composition of 3,935 classified clusters.

A survey of 3,974 archaeal biosynthetic gene clusters, and a substrate predictor for the fraction tools can see

Archaea are treated as an underexplored source of natural products. We surveyed 1,681 archaeal genomes, catalogued 3,974 clusters, built a substrate predictor for them, and tested the two usual explanations for the detection gap against matched bacterial controls. Neither survived.

technical report

Ranked binder candidates against scrambled controls, and absolute affinity accuracy against PRODIGY.

PRESTO: one model for binding affinity, and where it stops working

Affinity prediction, de novo binder ranking and mutational effects are treated as one problem by practitioners and as three by the tools. PRESTO is evaluated on all three, roughly halves the error of the established contact-based predictor, and reports the two tasks where it fails as findings rather than omissions.

technical report

An annotation metric that counts vocabulary, and two tools whose agreement is indistinguishable from chance.

Seven silent failure modes in genome-mining pipelines

We re-derived every number in one of our own working pipelines from the files that produced it. Seven stages were wrong in ways that looked exactly like being right: five were invisible until recomputation, and one had survived an earlier audit of ours that read the same evidence and drew the opposite conclusion.

technical report

Change in AUC-ROC as the partner is rotated, scrambled and swapped, across frozen and LoRA models.

Near-cognate decoys hide partner identity in disordered-region benchmarks

Benchmarks for partner-conditioned prediction of disordered regions share a recipe, and a model is credited with using the partner when it beats a partner-free ablation. We built one to current best practice, 5,683 complexes, then audited every step of the recipe. Four of six steps are wrong in a direction that flatters the benchmark.

case study

Omic Discover mid-run: six databases searched, 391 results ranked to 25 papers, with the literature panel open.

Can you stop onion tears without touching the onion? A recorded Omic investigation

Ten stages, unattended: 391 search results ranked to 25 papers, two real crystal structures pulled from the Protein Data Bank, three food-safe candidate inhibitors, and a 3,896-word manuscript with thirteen references.

case study

Omic Discover mid-run: three mechanisms for the flavor question, scored 8.0, 7.3 and 7.0 in a hypothesis tournament.

Can you engineer a no-cry onion that tastes the same? An end-to-end Omic investigation

One question, start to finish: 842 candidate papers, 25 kept, three scored hypotheses, and an answer that says what is established and what is not.

preprint

Trial-selection funnel narrowing 428,377 records to 3,133 trials, and the nine-module model architecture.

MERIT: predicting trial outcomes from mechanism, not from compound memorization

Most trial-outcome models can recognize the drug in front of them. Score them only on drugs they have never seen and the reported numbers fall. MERIT predicts from biology instead, reaches an honest AUROC of 0.770 across 753 drugs and 3,133 trials, and has 55 locked Phase III predictions on the record.