Every scientist has lived this. You run the same protocol twice with the same kit. Monday's data looks great and Friday's doesn't. You open your notebook and find a few lines you wrote three days later, from memory.
The difference is almost always in the execution. An incubation that ran long, a reagent that sat on the bench, a step you repeated. But you can't compare what you didn't record.
We document plans, not runs
Our publishing system is built around intent. A methods section describes the protocol as designed. A protocol repository stores the protocol as written. What happened during a particular run mostly lives in someone's head.
That is why, in 2026, we still email corresponding authors to ask what they actually did. And it's why so many of those emails go unanswered. The graduate student who ran the experiment may have left years ago.
Every other industry has a flight recorder
I keep coming back to how unusual life science is here. Airplanes have had flight data recorders for decades. Cars have telematics. Websites know exactly how people move through them. Television has audience measurement.
Research tools are one of the few products we send into the world and then learn almost nothing about how they're used. Most of the feedback that does come back arrives as a complaint, long after the run is over.
Why this is possible now
Asking scientists to write more documentation has never worked. People skim protocols and skip notes, and I don't blame them. They have gloves on and a timer running.
What has changed is that capturing a record no longer has to be a separate task. Voice interfaces and AI can turn the running conversation at the bench into structured data as the work happens. When capturing a record takes no extra effort, it can finally become routine instead of a burden.
What a useful record answers
A protocol execution record should answer a short list of questions:
Protocol execution record
- Which protocol and version was this run based on?
- Which reagents, lots, and instruments were used?
- In what order were the steps done, and how long did each take?
- Where did the run deviate from the protocol, and why?
It deliberately leaves out results. The record describes how the experiment was run, not what it found.
The payoff
Imagine submitting execution records alongside your sequencing data, the way we already submit structured metadata to public repositories. A reader could see exactly how your run compared with the reference protocol, or with their own. Your Friday run could be compared with your Monday run step by step. Two labs getting different results could see where their executions diverged.
That's the standard I want to help build. If you missed it, here's why it has to be open.