Before I drafted a single field of a protocol execution standard, I talked with people who run research identifier registries. Their advice was blunt: don't reinvent anything that already exists. Plug into it.

That turns out to be both good manners and good engineering. A lot of the hard work is already done.

RRIDs: knowing what was on the bench

Research Resource Identifiers (RRIDs) give reagents and tools a persistent ID that anyone can look up. They come from existing authoritative databases: the Antibody Registry for antibodies, Addgene for plasmids, and Cellosaurus for cell lines, with SciCrunch aggregating them (RRID citation guidelines). Publishers have built them into their workflows. Cell Press made its STAR Methods format mandatory, and it asks authors to list resources with identifiers, including RRIDs (overview of the RRID initiative, PMC).

This is a growing resource and we can all contribute. If you have a kit, molecule or other entity not currently covered by an RRID, you can submit the resource for consideration here.

PIDINST: knowing which instrument

A Research Data Alliance working group developed PIDINST, a metadata schema for persistently identifying instruments, and RDA endorsed it as a recommendation. It is built around individual instruments, not just instrument types (PIDINST working group).

That distinction matters. Knowing someone used a particular model of flow cytometer gets you to the right neighborhood. Knowing which specific machine they used gets you to the right house.

Five layers, one new piece

Here is how I think a protocol execution record should fit together:

Protocol execution record: five layers

  1. Reference protocol: the versioned protocol the run was based on, identified with a DOI
  2. Resources: reagents, kits, and lots, cited with RRIDs wherever they exist
  3. Instrument: the specific instrument, described with PIDINST
  4. Execution: step order, timing, pauses, repeats, and deviations. This is the only layer that is genuinely new
  5. Links: pointers to the resulting datasets and publications

My working proposal is that the standard defines layer 4 and nothing else. Everything else points to identifiers the community already maintains. That keeps the standard small, and it means a record can plug into tools people already use. The group will pressure-test this.

The gaps worth filling together

This approach also shows where the holes are. The RRID guidelines cover antibodies, cell lines, organisms, plasmids, and software tools. Everyday kit-based reagents, like nucleic acid extraction and PCR kits, aren't among the listed types. On the instrument side, I've heard in my conversations that capturing detailed settings and calibration history is still unfinished. A standard for execution records gives these efforts a concrete reason to connect.

Sources

  1. RRID citation guidelines
  2. Overview of the RRID initiative, PMC
  3. RDA recommendation: PIDINST metadata schema
  4. RDA PIDINST working group
  5. Submit a resource to RRID