Models Can't Learn from What Science Doesn't Record
Execution data for model training, evaluation, and robot learning in science
Aug 25, 2026

Scientific AI has a data problem at the point where instructions become physical work. Publications are a static record of conclusions, protocols record intended steps, and laboratory systems record selected results. Papers rarely, if ever, preserve the atomic actions, object states, timing, deviations, judgments, and recoveries that produced those results.
Transfyr records that execution layer, working towards a lossless record of scientific work. Transfyr’s sensor suite for robotics/machine-readable data capture combines time-synchronized egocentric, overhead, and side view high-resolution video capture to allow for 3D reconstruction of the scene and detailed analysis of the scientific action space. When combined with experimental data and metadata, including the protocol, equipment, sample, environmental, and outcome context, this provides a rich training ground for models and robots to understand how specific physical actions translate to experimental results. Transfyr’s perception systems then provide detailed hierarchical labeling to allow different systems to interact with this data.

The bridge from the reality to simulation
AI and robotics need a way to interpret complex physical realities. Transfyr passively captures real scientific work at robotic quality without interrupting a scientist’s workflow. Synchronized views show the operator's perspective and the surrounding workspace. The perception layer identifies objects, contacts, motion, actions, and changes in laboratory state. Each item remains linked to the protocol and experimental context.

We collaborate to determine the appropriate vocabulary, annotation depth, time resolution, privacy controls, and quality thresholds for a specific task. A log for an automation engineer may only need task-level descriptions, whereas a robotic policy may need millisecond-level action boundaries and an evaluation set may need selected edge cases with evidence-backed scoring. Data can be adjudicated by multiple subject-matter experts to create a high fidelity key that can be used to label related content reliably and with confidence scores, at scale.
We leverage these datasets in our work with frontier AI labs, robotics teams, and lab automation engineers across:
- Robot learning from human demonstrations, including examples near the edges of successful execution
- Traditional lab automation deployment, giving automation engineers an atomic-action level description of protocol steps.
- Model evaluation for capability and safety risk: evaluation of object and action recognition, error detection, and human uplift
Automation / Robotics
There are two main approaches to scientific automation today: traditional workcell-based automation and generalized robotics (for an excellent primer, we love Abhi’s coverage on Owl Posting).
Traditional automation translates a scientific method into a detailed machine program, executed on a series of specialized machines, programmed to execute a desired task. An automation engineer (whether human or robotic), decomposes a method or protocol into actions such as moving to a coordinate, gripping an object, aspirating a volume, waiting for a set time, and verifying a state.
The bottleneck most automation engineers experience is that protocols are rarely specified sufficiently - “mix” can mean any number of different things and sometimes it matters which one was selected. This type of automation also tends to come with some amount of sample multiplexing, which can easily change the “interstitial timing” between steps for the Nth sample. While coding agents can reduce the programming burden, the debugging process between scientist and engineer or machine is still long and costly, often taking 6 months or longer.
Vision-language-action (VLA) models, imitation learning, and related methods offer a second route towards more generalizable (or even humanoid) robotics. These systems learn from human demonstrations and feedback rather than relying on a fully specified program for every action. Raw video does not contain the scientific interpretation on its own. Robot-learning data benefits from context - the relevant objects, protocol step, sample state, action boundary, outcome, error, and recovery. It also needs variation across operators, equipment, environments, and acceptable ways to complete the task.
Transfyr makes both paths more tractable. We capture what actually happens during scientific work: the operator’s actions, objects and samples, protocol step, timing, equipment state, deviations, recovery, and outcome context. We then convert that record into structured, time-synchronized execution data.

For traditional automation, this creates an evidence-backed process specification. Instead of translating an underspecified protocol from memory and observation alone, engineers can see the real workflow, identify consequential variation, and validate whether a workcell is reproducing the intended process. For general robotics, the same record becomes training and evaluation data: demonstrations linked to task intent, action boundaries, laboratory state, errors, and acceptable recovery.
Transfyr is not a workcell integrator or a robot manufacturer, but we help companies deploy automation by providing the ground-truth execution layer that helps automation teams specify, debug, train, and evaluate systems against the reality of scientific work.
Model training & evaluation
Science is not a multiple-choice question. Yet most biology benchmarks are built as if it were: a model chooses one answer from a short list, usually using text derived from published papers or protocols. MCQs are cheap, reproducible, and easy to score. But they also reveal the shape of the answer, constrain the search space, and let partial pattern matching look like scientific competence.
Real science is both open-ended and grounded in physical reality. The most interesting questions don’t have a known answer but many can be explored with an experiment. Thus the most relevant scientific benchmarks must allow a model to interrogate routes previously explored and test the validity of that route in the real world. Unfortunately, a “real” scientific benchmark would require vast investments in responsive physical exploration of open-ended hypotheses. We offer a useful bridge: real-world evaluations that test the model’s ability to troubleshoot science in the real-world, responding to the same multimodal signals that an expert scientist would leverage. We ask models to identify sources of process variance or error and establish causality when combined with scientific results, supporting separation of biological variance from procedural explainability.
Open-ended does not mean unscorable. Transfyr evaluates models at several levels - basic multimodal reasoning (whether the model can correctly identify objects and actions), scientific reasoning (whether models can identify errors and link them to experimental outcome deviation), and uplift. Individual steps can be scored for object, action, state, timing, and protocol adherence. A full episode can be scored against subject-matter-expert feedback, experimental validity, and outcome. Inputs can be masked, perturbed, or withheld to test which evidence the model uses. Uplift trials can measure whether AI assistance improves expert performance (accelerating scientific discovery) or helps a novice approach an expert baseline (a critical safety metric). In addition to our field-deployed work, we leverage our labs in Cambridge, MA to design and execute experiments, allowing us to tightly control experimental design and perturbations to maximize the variance observed by a model.
Biosecurity and safety evaluation for frontier models
As generative AI and physical automation converge in life sciences, safety and biosecurity evaluations are becoming critical. Assessing whether AI models lower the technical barrier to dual-use capabilities - or whether safety guardrails hold firm under real-world conditions - requires observing and measuring the real-world physical execution of experiments.
Transfyr’s lab in Cambridge, MA doubles as a physical evaluation environment where frontier AI models, operators, and bio-automation systems are benchmarked against high-consequence biological workflows:
- Multimodal Capability Uplift & Failure Analysis: We evaluate how AI assistance alters the performance of operators across the expertise spectrum—from novices to domain experts. By tracking physical actions, timing, and errors, we identify the exact friction points where non-experts fail, establishing an empirical baseline for the physical capability uplift provided by a model.
- Controlled Perturbations for AI Supervision: We deliberately execute controlled protocol perturbations and edge-case deviations. This data feeds directly into model training and evaluation, teaching models to identify sources of process variance, establish biological causality, and detect procedural anomalies.
- Real-world Multimodal Benchmarks: Replacing static text-based multiple-choice tests, Transfyr generates physical evaluation sets to score models on real-world troubleshooting, error identification, hazard recognition, and adherence to containment standards.
- Guardrail Testing Environment: We can deploy live interaction loops between AI models and human operators to stress-test safety controls and embedded technological guardrails for efficacy. Discovered vulnerabilities can be systematically tracked to inform future iterations of guardrail design.
By capturing the physical execution layer where biological risks actually materialize, while using biosafe surrogates and proxies, Transfyr provides frontier AI labs with the ground-truth observability required to measure physical capability uplift, map misuse friction points, and embed safety directly into next-generation physical AI and autonomous lab infrastructure.


