How each capability is evaluated
Each repo is checked on two axes: whether the science holds and whether it runs unattended. Physical execution is measured with Rhodamine B, Qubit, and computer vision, never inferred from a software log.
The physical biology and chemistry executed correctly, and the bioinformatics recovers ground truth, against a metric fixed before the run.
Physical-AI QC in the loop, error handling that fails closed, and closed-loop feedback where the agent reads QC, corrects, and re-runs.
| Capability | Science axis | Autonomy axis | Graded by | Expert in loop |
|---|---|---|---|---|
| fullstack-omics FLASH-seq scRNA + scWGS |
UMI counts, low-input recovery, coverage; wet chemistry to run on deckBuilding |
End to end in the PyLabRobot simulator; scWGS 10/10 testsVerified |
scanpy, BJ-WGS concordance | operator at deck |
| plr-mcp Instruments as agent tools |
Actions execute on liquid handler, reader, cycler, shakerVerified |
MCP tool-call round-trip; CI green on 3.10-3.13Verified |
simulator + hardware handshake | agent proposes, operator confirms |
| omics-demos Nine scored demos |
Recovery vs. planted ground truth across nine assaysVerified |
Blind run, per-assay recovery scoreVerified |
recovery score, threshold preset | none, automated eval |
| plr-minimum-effective Minimal recipe search |
Yield held at reduced reagent; wet confirmation pendingBuilding |
Bayesian optimization, held-out recoveryVerified |
cross-validated recovery | expert sets acceptance floor |
| plr-clarity (sci)TIP-seq on Hamilton STAR |
Rhodamine-B validation ladder on a real deck; biovalidated tierBuilding |
Plan text compiled into a runnable methodVerified |
validation tier reached | expert signs each tier |
| lab-cv Spatial AI on bench video |
Wells filled vs. empty, pipetting correctVerified |
ROI motion QC + detection on protocol video, CPU-onlyVerified |
ground-truth-validated frames | reviewer adjudicates |