ATTENTION: To use this site, it is necessary to enable JavaScript in your browser.
Here are the Instructions on how to enable JavaScript in your web browser.

Razor Evaluation Protocol — Empirical Testing Framework for AI Labs

Preregistered Empirical Evaluation Framework

Razor Evaluation Protocol

Empirical Testing Framework for AI Labs

A quality-gated, architecture-agnostic protocol for testing whether a specified implementation of Robbie’s Razor changes the performance, resource use, reusable state, or stability of an AI system under equivalent, preregistered conditions.

Protocol boundary: This framework does not assume that Razor-guided reasoning is more efficient. It defines how that proposition can be tested, what evidence must be collected, and what would count as failure or an inconclusive result.

Visual introduction to the Razor Evaluation Protocol for controlled AI-system testing.
Evaluation is performed on an explicitly registered system, workload, configuration, baseline, measurement period, and evidence boundary.
Document version: 2.0 Governing authority: GC-MRD-v2.0 Status: Public evaluation protocol Canonical law: Robbie’s Razor

Canonical and Evidentiary Notice

The Protocol Tests a Proposition; It Does Not Presume the Result

This protocol is governed by GC-MRD-v2.0 and should be read alongside the Grand Compression Canonical Claims Register. The governing framework defines the theory and its claim boundaries; this page defines a controlled process for evaluating specified implementations under observable conditions.

“When competing explanations exist, prefer the model that follows compression → expression → memory → recursion.”

Canonical formulation of Robbie’s Razor. This sentence should not be paraphrased when presented as the canonical law.

The law supplies the proposition to be examined. It does not, by itself, establish that a prompt, controller, memory system, registry, retrieval process, or other implementation will reduce cost, improve quality, preserve state, or remain stable in a particular model or workload.

Canon defines

GC-MRD-v2.0 defines the governing framework, terminology, scope, claim boundaries, and formal evidence states.

Protocol tests

This protocol preregisters the system, workload, baseline, quality threshold, metrics, failure conditions, and analysis plan.

Evidence records

Results receive scoped finding labels and may inform a separate formal evidence-state review. A documented observation is not automatically a Supported claim.

Implementation is not validation. Building a tool, publishing code, integrating the framework, licensing it, accepting payment, appearing in a repository, or being used in an AI interface can demonstrate implementation, publication, access, or adoption. None of those events independently demonstrates effectiveness.

Results apply only to the registered system, version, configuration, workload, constraints, measurement period, and evidence boundary. Transfer to another model, domain, scale, or environment requires a new test.

Relationship to the Evaluation System

The Razor Auditor identifies questions, gaps, and possible test targets. This protocol converts those questions into preregistered experiments. The Robbie’s Razor Benchmark Hub organizes public evaluation pathways, while the GitHub repository serves as a public reproducibility layer for applicable code, fixtures, schemas, and result assets.

The governing sequence is: canon defines → claims register tracks → Auditor diagnoses → protocol tests → GitHub preserves → replication checks → evidence states record.

Reasoning-privacy boundary: The protocol evaluates observable inputs, outputs, tool calls, stored state, retrieval behavior, execution results, telemetry, and accepted outcomes. It neither requires nor claims access to hidden chain-of-thought.

Protocol Foundation

1. Purpose and Scope

The Razor Evaluation Protocol provides AI laboratories, research groups, platform teams, and independent evaluators with a controlled method for testing a specified implementation of Robbie’s Razor. Its purpose is to determine whether an observed result survives equivalent conditions, a preregistered quality threshold, total-cost accounting, and repeatable measurement.

The protocol can be applied to prompt-layer experiments, controller logic, retrieval systems, agent workflows, memory architectures, registries, tool-using systems, or other observable implementations of the sequence compression → expression → memory → recursion. These four dimensions are referred to throughout this page as CEMR.

Required Unit of Evaluation

Every evaluation must identify the complete unit being tested:

System + Version + Configuration + Workload + Constraints + Measurement Period + Evidence Boundary

A result that omits any material part of this unit may be useful as an observation, but it is not sufficient for a formal evidence-state decision. Changes to model weights, model version, system prompts, decoding settings, retrieval configuration, tool access, hardware allocation, batching, workload composition, or acceptance criteria can create a new evaluation unit.

Architecture-Agnostic

The protocol may be adapted to different model architectures and system designs because it evaluates registered behavior, outcomes, costs, and state rather than requiring one proprietary implementation.

Workload-Specific

Each conclusion is limited to the registered workload and acceptance standard. Results from reasoning, retrieval, coding, planning, or agent tasks should not be treated as interchangeable.

Evidence-Bounded

Findings must remain inside the available evidence boundary. Unmeasured behavior, unavailable telemetry, hidden processes, and missing cost categories should be reported as Unknown.

Appropriate Evaluation Contexts

The protocol is most useful where a system produces observable differences that can be compared under controlled conditions, including:

  • long-context analysis and complex question answering;
  • retrieval-augmented generation and multi-stage evidence synthesis;
  • tool-using agents, planners, and coordinated workflows;
  • code generation, execution, verification, and correction;
  • memory, registry, and reusable-state experiments;
  • controller or prompt-layer CEMR implementations; and
  • bounded recursive tasks in which identity, relationships, provenance, constraints, or version state must persist.

Scope rule: Architecture-agnostic does not mean universally transferable. Every new model, configuration, workload, domain, scale, or operational environment requires its own applicability decision and, where claims are made, a new test.

Testable Questions

2. What the Protocol Tests

The protocol compares a registered baseline with a registered intervention under equivalent conditions. It asks whether the intervention produces a measurable difference—and whether that difference remains meaningful after quality, verification, correction, failed runs, and other material costs are included.

Each question selected for evaluation must be converted into a preregistered prediction with defined metrics, units, quality thresholds, success thresholds, failure conditions, and an analysis plan.

Quality Gate

Are the outcomes acceptable?

Does the intervention meet the same preregistered standards for correctness, completeness, safety, usefulness, constraint compliance, format compliance, verification burden, correction burden, and memory fidelity?

Compression

Is relevant structure represented more economically?

Does the intervention reduce unnecessary representation, retrieval, repetition, or processing while preserving the information required to produce an accepted outcome?

Expression

Is the compressed structure expressed effectively?

Does the system convert its registered inputs and reusable state into an output, action, tool sequence, or decision that satisfies the same acceptance criteria?

Memory

Is reusable state preserved faithfully?

Does stored or retrieved state preserve the relevant identity, relationships, provenance, constraints, version state, and retrieval pathways needed by later operations?

Recursion

Can the preserved state be reused without unacceptable drift?

Across repeated or nested operations, does the system preserve required facts and constraints, avoid material contradiction, and stop or escalate according to registered rules?

Total Cost

What is the cost per accepted outcome?

After including construction, model calls, storage, indexing, retrieval, tools, verification, human review, correction, failures, retries, coordination, infrastructure, and maintenance, does total cost change?

Example of a Properly Scoped Test Question

Under a named model version, fixed configuration, registered workload, equivalent tool access, and identical acceptance threshold, does a specified CEMR intervention change total measured cost per accepted outcome during the stated observation period?

Quality-first rule: A shorter, faster, or less expensive output is not an efficiency improvement if it fails the same acceptance standard or shifts material work into verification, correction, retries, human review, or another uncounted system component.

Non-Claim Boundary

3. What the Protocol Does Not Establish

Running this protocol does not guarantee a favorable result. It also does not convert an implementation, metric change, pilot, publication, or commercial relationship into proof that Robbie’s Razor is effective.

The protocol does not automatically establish Required boundary
Universal efficiency Any efficiency finding is limited to the registered system, configuration, workload, constraints, measurement period, and acceptance standard.
Automatic token, FLOP, latency, or KV-cache reduction These are separate measurable variables. A change in one does not prove a change in another, and each requires appropriate instrumentation or a clearly labeled proxy.
Hallucination mitigation Factual error, unsupported assertion, contradiction, or another failure type must be operationally defined, labeled, and measured before a scoped finding can be made.
Energy, cooling, water, carbon, or emissions savings Environmental conclusions require direct telemetry or a defensible accounting model that includes hardware, utilization, batching, geography, energy supply, cooling design, and measurement boundaries.
Transfer to other domains or systems Success in one model, workload, domain, scale, or environment cannot be generalized to another without a new applicability decision and test.
Certification or compliance A completed evaluation may produce evidence about defined requirements. It does not create certification unless a separate, documented certification process and authorized decision exist.
Independent validation Author-created tools, tests, examples, or results must remain distinguishable from independent evaluation and replication.
Access to hidden chain-of-thought Evaluation should rely on observable inputs, outputs, tool calls, stored state, retrieval behavior, execution results, telemetry, and accepted outcomes.
Effectiveness through licensing or adoption Licensing, integration, payment, publication, indexing, publicity, or use demonstrates access, activity, or adoption—not validated performance.

Missing Measurement

If the necessary telemetry, provenance, cost record, or comparison data are unavailable, the appropriate finding label is Unknown.

Failed Prediction

If the intervention misses its quality threshold, success threshold, or another preregistered condition, that outcome must be recorded rather than redefined after observation.

Mixed Result

A favorable change in one measure does not erase a regression elsewhere. Tradeoffs, uncertainty, quality failures, and shifted costs must remain visible.

Interpretation rule: The protocol is designed to permit favorable, unfavorable, mixed, inconclusive, and failed outcomes. A framework that cannot record failure cannot provide a credible empirical test.

Evaluation Registration

4. System-Boundary Registration

Before predictions, metrics, or trials are selected, the evaluator must define exactly what system is being tested. The system boundary determines which components, behaviors, resource costs, and failure modes belong inside the comparison.

A boundary that includes model execution but excludes retrieval, tools, verification, correction, storage, human review, or maintenance can create a misleading result. Any component necessary to produce an accepted outcome should be included or explicitly identified as excluded.

Register the Complete Evaluation Unit

System + Version + Configuration + Workload + Constraints + Measurement Period + Evidence Boundary

Registration field Required record Why it matters
System Model, application, agent, retrieval system, controller, tools, memory layer, and human operations included in the test. Prevents the comparison from counting benefits in one component while hiding costs shifted elsewhere.
Version Model identifier, weights or release, application build, prompt version, controller version, dataset version, and dependency versions. Makes the result reproducible and prevents silent version changes from entering the comparison.
Configuration System instructions, decoding settings, context limits, retrieval parameters, memory rules, tool permissions, retry policies, batching, routing, and hardware allocation. Separates the tested intervention from unrelated configuration effects.
Workload Named task set, dataset or prompt source, sampling procedure, task classes, difficulty strata, and number of evaluation units. Defines the population to which the result may apply.
Constraints Time, context, compute, storage, tool, privacy, safety, format, cost, and stopping constraints applied to each arm. Ensures the comparison occurs under equivalent operating requirements.
Measurement period Start and end times, trial order, warm-up treatment, repeated-run schedule, and any known infrastructure conditions. Makes time-dependent variation and operational drift visible.
Evidence boundary Observable data, telemetry, logs, annotations, cost records, unavailable variables, proxies, and known missing data. Prevents conclusions from extending beyond what was actually observed.

Define the Intervention Separately

The Razor intervention must be documented as a versioned change to the baseline. State whether the intervention operates through a prompt, controller, memory process, retrieval rule, registry, tool policy, training procedure, or combination of components.

The record should identify which CEMR dimensions the intervention is intended to affect, how each dimension is operationalized, and which system components remain unchanged.

Change-control rule: If a material component, version, workload, constraint, acceptance threshold, or measurement procedure changes after preregistration, record the change and determine whether the trial must be treated as a new evaluation rather than merged with the original result.

Task and Dataset Definition

5. Workload Registration

The workload determines what the evaluation can legitimately say. Register the tasks, source material, sampling procedure, difficulty distribution, acceptance rules, and exclusions before observing comparative results.

Workloads should reflect the intended use of the evaluated system. A test set selected only because it appears favorable to one intervention cannot support a credible comparative finding.

Representative Workloads

Production-like tasks selected to estimate performance under the system’s intended operating conditions.

Controlled Fixtures

Deterministic or tightly bounded tasks designed to isolate identity, constraint, memory, retrieval, or recursion behavior.

Stress and Boundary Tests

Adversarial, long-context, collision-heavy, constraint-heavy, or failure-oriented tasks used to identify operating limits rather than average performance.

Required Workload Record

Field Registration requirement
Workload name and version Assign a stable identifier and version to the complete evaluation set.
Source and provenance Document whether tasks come from production, a public benchmark, a synthetic generator, a controlled fixture, or another source.
Task classes Identify relevant categories such as retrieval, planning, code, tool use, synthesis, constraint following, state retention, or repeated recursion.
Sampling procedure Record inclusion, exclusion, randomization, stratification, deduplication, and any balancing procedure.
Difficulty and length Describe task complexity, context length, expected operation count, tool dependence, and relevant difficulty strata.
Dependencies Identify required source documents, retrieval indexes, tools, APIs, stored state, prior turns, or external services.
Accepted outcome Define what counts as correct, complete, safe, useful, constraint-compliant, and successfully verified.
Failure labels Define task failure, factual error, unsupported assertion, constraint violation, tool failure, state loss, timeout, and other applicable outcomes.
Confidentiality status Record whether prompts, source data, outputs, annotations, or results can be published, summarized, or accessed by an independent evaluator.

Preserve Difficult and Unfavorable Cases

Timeouts, failed tool calls, correction-heavy outputs, missing answers, constraint violations, and tasks on which the intervention performs poorly must remain in the result record according to the preregistered missing-data and exclusion rules.

Anti-selection rule: Do not remove a workload, task class, difficulty stratum, or completed trial after results are visible merely because it weakens the preferred conclusion. Any post-registration exclusion must be disclosed, justified, and reported with sensitivity analysis where possible.

Comparative Control

6. Baseline Selection

The baseline should represent the strongest relevant system or procedure that the Razor intervention is intended to improve or replace. It should not be intentionally weakened, stripped of ordinary optimizations, or configured in a way that creates an artificial advantage for the intervention.

Preferred Baseline

Use the current production configuration, the strongest accepted internal method, or another documented comparator that a reasonable practitioner would use for the registered workload. If no established baseline exists, disclose how the comparator was constructed and why it is appropriate.

Matched Conditions Across Evaluation Arms

Component Baseline arm Intervention arm
Model and version Registered model, weights, release, and endpoint. The same model and version unless the model change is itself a separately registered factor.
Workload Same registered tasks and source material. Same tasks, sources, ordering rules, and inclusion criteria.
Tools and retrieval Registered tools, permissions, indexes, and retrieval access. Equivalent access unless the Razor intervention explicitly changes this component and its costs are included.
Decoding and limits Registered temperature, sampling, context, output, time, and retry limits. Matched settings unless a difference is preregistered as part of the intervention.
Infrastructure Registered hardware, region, routing, batching, and service tier. Matched or appropriately randomized infrastructure conditions.
Quality standard Same acceptance criteria and adjudication procedure. Identical quality threshold, verification process, and correction policy.
Measured cost All material costs required to reach an accepted outcome. The same cost categories, including intervention-specific construction and maintenance.

Multiple Baselines

Where appropriate, register more than one comparator. A production baseline can show operational relevance, while a simpler control can isolate the effect of a specific prompt, controller, memory, or retrieval change.

Report each comparison separately. Do not select only the baseline that produces the most favorable result after the data are observed.

Baseline Defects

Known limitations, instability, error patterns, and operational costs in the baseline should be documented before comparative results are reviewed.

Intervention Overhead

Prompt construction, controller execution, state creation, storage, indexing, retrieval, verification, and maintenance introduced by the intervention belong in the comparison.

Fair Opportunity

Each arm should receive a reasonable, documented configuration without post-result tuning that is unavailable to the other arm.

Baseline-integrity rule: If the baseline and intervention differ in multiple uncontrolled ways, the evaluation may document an observed system-level difference, but it cannot attribute that difference specifically to Robbie’s Razor.

Acceptance Before Efficiency

7. Quality Gate

Quality must be evaluated before a reduction in tokens, latency, model calls, storage, or another resource measure can be interpreted as improved efficiency. The baseline and intervention must be judged against the same preregistered acceptance standard.

A shorter or cheaper result is not more efficient unless it passes the same acceptance standard.

Quality Dimensions

Select and operationally define the dimensions that apply to the registered workload. Not every task requires every dimension, but exclusions should be justified before results are observed.

Correctness

Whether claims, calculations, classifications, actions, or execution results satisfy the registered truth or validation standard.

Completeness

Whether all required components, conditions, requested outputs, and material qualifications are present.

Safety

Whether the output or action remains within applicable safety policies, operating limits, and risk controls.

Usefulness

Whether the result can perform its intended function for the registered user, system, or downstream process.

Constraint Compliance

Whether hard limits, stop rules, permissions, exclusions, deadlines, and task-specific constraints are preserved.

Format or Schema Compliance

Whether the output satisfies required structure, syntax, fields, data types, ordering, or machine-readable validation.

Verification Burden

The work required to determine whether the result is acceptable, including automated checks and human review.

Correction Burden

The work, retries, edits, tool calls, or interventions required to convert an initially unacceptable result into an accepted outcome.

Memory Fidelity

Whether required identity, relationships, provenance, constraints, version state, and retrieval pathways remain intact.

Define the Acceptance Procedure

Procedure element Required definition
Acceptance threshold The minimum score or complete set of conditions required for an accepted outcome.
Evaluation method Exact match, executable test, schema validator, source verification, rubric, expert review, user acceptance, or another registered method.
Adjudication Who or what resolves disputed, ambiguous, or partially correct outcomes and under which rules.
Blinding Whether human or model-based evaluators can identify the experimental arm. If blinding is not possible, record the limitation.
Evaluator reliability Calibration, agreement, audit sampling, or other checks used when judgment is not fully deterministic.
Unscorable cases Rules for missing references, unavailable tools, corrupted data, evaluator disagreement, or tasks that cannot receive a valid quality decision.

Required reporting: Report acceptance rate for each arm, the number of accepted, failed, and unscorable trials, and the reason for each exclusion. Resource comparisons should be presented both across all attempted trials and per accepted outcome where appropriate.

Quality-gate rule: If the intervention fails the preregistered quality threshold, the efficiency prediction fails for that evaluation even if one or more resource measures decrease.

Before Results Are Observed

8. Prediction Preregistration

A formal evidence-state decision requires a prediction and evaluation plan registered before comparative results are inspected. Preregistration separates tests designed in advance from explanations developed after observing the data.

The registration may be public, confidential, or held by an independent reviewer, but it should be time-stamped, versioned, attributable, and protected from silent revision.

Minimum Preregistration Record

Before a claim can receive a formal evidence-state decision, register the prediction, system boundary, version and configuration, workload, baseline, metrics, units, quality threshold, success threshold, failure conditions, analysis plan, observation period, missing-data procedure, result record, and total-cost treatment.

Registration field Required content
Registration ID Stable identifier, title, author or evaluation team, version, registration date, and change history.
Prediction A directional, non-directional, equivalence, or non-inferiority prediction stated before results are observed.
System boundary Complete evaluation unit, included components, excluded components, operational dependencies, and evidence boundary.
Version and configuration Model, software, prompts, controller, memory, retrieval, tools, decoding, routing, limits, and infrastructure settings.
Workload Named workload, version, source, sample definition, task classes, exclusions, and accepted-outcome rules.
Baseline and intervention Comparator selection, intervention definition, matched conditions, expected mechanism, and known implementation overhead.
Metrics and units Primary and secondary measures, operational definitions, units, aggregation rules, normalization, and provenance.
Quality threshold Minimum acceptance standard and any permitted non-inferiority margin, including how it will be evaluated.
Success threshold The magnitude, direction, uncertainty requirement, or complete set of conditions required for the prediction to pass.
Failure conditions Conditions that falsify, fail, stop, invalidate, or make the registered test inconclusive.
Analysis plan Comparison method, estimand, statistical or deterministic procedure, uncertainty treatment, subgroup plan, and multiplicity treatment.
Observation period Trial dates, duration, run order, repeated-trial plan, warm-up policy, and stopping rule.
Missing-data procedure Treatment of timeouts, service failures, unavailable telemetry, corrupted logs, unscorable outcomes, and incomplete trials.
Total-cost discipline Included cost categories, allocation rules, excluded costs, measurement basis, and normalization per attempted and accepted outcome.
Result-record format Required fields, provenance, raw-data references, finding labels, deviations, limitations, and evidence-state review pathway.

Primary, Secondary, and Exploratory Analyses

Primary

The principal prediction and measure used to determine whether the preregistered test passes or fails.

Secondary

Additional preregistered measures that characterize tradeoffs, mechanisms, boundary conditions, or supporting outcomes.

Exploratory

Patterns identified outside the preregistered decision rule. These may motivate later predictions but should not be presented as confirmatory results.

Amendment rule: Any change made after registration must preserve the original record, identify what changed and when, explain why, and state whether affected results remain confirmatory, become exploratory, or require a new evaluation.

Measurement Discipline

9. Metrics and Units

Every metric must have an operational definition, unit, source, collection procedure, aggregation rule, normalization rule, and known limitation. A metric name without these elements is not sufficient for reproducible comparison.

Select measures that correspond to the registered prediction and system boundary. Do not evaluate efficiency using a favorable single metric while excluding quality regressions, shifted work, failed runs, or material costs elsewhere in the system.

Measure family Possible measures Unit and boundary requirements
Quality Acceptance rate, correctness, completeness, safety, usefulness, constraint compliance, schema compliance, memory fidelity. Pass/fail, score, proportion, error count, severity class, or another registered unit tied to a defined rubric.
Tokens and context Input tokens, output tokens, cached tokens, retrieved context, tool-result context, and total processed tokens. Tokens under a named tokenizer and model interface, reported by category rather than as an unexplained total.
Compute Measured operations, accelerator time, utilization-adjusted device time, model-call duration, or a documented compute proxy. FLOPs, accelerator-seconds, utilization-adjusted time, or named proxy. Tokens must not be presented as direct FLOP measurements.
Memory and storage Peak allocated memory, KV-cache size, stored-state size, index size, read/write volume, retention period. Bytes, kilobytes, megabytes, gigabytes, byte-seconds, or another explicitly defined unit and measurement point.
Latency and duration Time to first output, model latency, tool latency, retrieval latency, verification time, correction time, end-to-end completion time. Milliseconds, seconds, or minutes, with measurement start and stop points explicitly defined.
Retrieval and tools Queries, documents retrieved, bytes transferred, tool calls, failed calls, retries, execution time, and external-service charges. Counts, bytes, duration, currency, or cost units under a named provider and pricing period.
Reusable state Identity retention, relationship retention, provenance retention, constraint survival, version accuracy, retrieval success, state reuse. Field accuracy, retention proportion, error count, retrieval success rate, or another fixture-defined unit.
Failures and retries Rejected outcomes, timeouts, invalid outputs, constraint failures, tool failures, retries, corrections, abandoned runs. Counts, rates, severity, additional duration, and additional cost per attempted or accepted outcome.
Human work Review, verification, correction, escalation, labeling, coordination, and system-maintenance effort. Person-minutes, person-hours, task counts, or documented financial cost.
Environmental telemetry Energy, cooling, water, carbon, or emissions measures when directly observed or supported by a defensible accounting model. Named physical units with hardware, utilization, geography, energy supply, cooling design, allocation method, and observation boundary.

Measurement Basis and Finding Labels

Individual measurement statements should disclose how they were produced. Use the framework’s approved finding labels without treating the label itself as a formal evidence state:

Documented
Calculated
Inferred
Proposed
Unknown

Normalization and Aggregation

  • Report raw totals and normalized values where both aid interpretation.
  • Normalize resource use per attempted task and per accepted outcome where appropriate.
  • Report acceptance rate alongside resource reductions.
  • Separate input, output, retrieval, tool, verification, correction, and retry costs.
  • Report central tendency, variation, and relevant tail behavior rather than only a favorable average.
  • Keep workload-level results visible when aggregation could hide regressions or boundary conditions.
  • Do not combine unlike measures into a composite score unless its formula, units, weights, and decision rule were preregistered.

Conversion boundary: A reduction in tokens does not automatically establish a reduction in FLOPs, KV-cache allocation, latency, energy, cooling, water, carbon, emissions, financial cost, or total system cost. Each conclusion requires its own measurement or defensible calculation.

CEMR Dimension One

10. Compression Measures

Within this protocol, compression means reducing the representation or processing burden of a task while preserving the information, relationships, provenance, constraints, and version state required for an accepted outcome.

Loss-aware boundary: Shorter context, fewer tokens, a smaller state object, or a more concise output is not automatically successful compression. If required information is lost, distorted, detached from provenance, or made harder to retrieve, the result may be truncation or information loss rather than useful compression.

Register What Must Be Preserved

Before measuring compression, define the required-information set for each task or fixture. Depending on the workload, this set may include:

Identity
Relationships
Provenance
Constraints
Version State
Retrieval Pathways
Required Evidence
Stop Conditions
Measure Operational question Possible unit
Representation size How much context, stored state, retrieved material, or output representation is used? Tokens, bytes, fields, nodes, edges, documents, or records.
Required-element retention How many preregistered required elements remain correct and usable after compression? Count or proportion of required elements preserved.
Relationship preservation Are dependencies, hierarchies, causal links, ownership, or other required relationships preserved correctly? Correct edges, relation accuracy, or fixture-defined score.
Constraint survival Do hard limits, exclusions, permissions, guardrails, and stop conditions survive the compressed representation? Constraint retention rate, violations, or severity.
Redundancy How much repeated or functionally duplicate representation appears within the registered boundary? Duplicate fields, repeated facts, repeated retrievals, tokens, or bytes.
Retrieval burden What work is required to recover the information needed for an accepted outcome? Queries, documents, bytes, tool calls, latency, or cost.
Reconstruction burden Must the system or a reviewer reconstruct information lost during compression? Retries, correction steps, human minutes, model calls, or financial cost.
Accepted-outcome rate Does the compressed representation still support outcomes that pass the same quality gate? Accepted outcomes divided by attempted outcomes.

Recommended Comparative Reporting

  • Report baseline and intervention representation size in the same defined unit.
  • Report the percentage change in representation size without treating that change alone as success.
  • Report required-element, relationship, provenance, and constraint preservation separately.
  • Report reconstruction, retrieval, verification, and correction costs created by the compressed state.
  • Separate lossless behavior, acceptable task-specific loss, and unacceptable information loss.
  • Normalize comparative cost per accepted outcome where appropriate.

Compression decision rule: A compression finding requires both a measurable reduction in the preregistered representation or processing burden and preservation of the information needed to pass the same quality gate. Otherwise, report the reduction and the associated loss as separate findings.

CEMR Dimension Two

11. Expression Measures

Expression measures whether the system can convert registered inputs, compressed representations, and reusable state into an output or action that satisfies the intended task. The relevant expression may be an answer, structured record, tool sequence, code execution, retrieval request, decision, or other observable result.

Expression boundary: Concision is not the same as successful expression. A shorter answer or action sequence is favorable only if it preserves required meaning, satisfies the registered constraints, and passes the same acceptance procedure.

Measure Operational question Possible unit
Outcome acceptance Does the expressed output or action pass the registered quality gate? Pass/fail, score, acceptance rate, or error severity.
Required-element coverage Does the expression contain every required answer component, field, action, citation, or decision element? Count or proportion of required elements present and correct.
Constraint compliance Does the expression obey registered format, safety, permission, scope, length, tool, and stopping constraints? Violations, violation rate, pass/fail, or severity.
Unsupported additions Does the output add factual claims, sources, actions, or conclusions unsupported by the registered evidence? Count, rate, or severity under a registered annotation rule.
Observable action count How many model calls, tool calls, retrievals, executions, or externally visible steps are required? Actions, calls, queries, executions, or state transitions.
Output burden What observable representation is produced to complete the task? Tokens, bytes, fields, records, lines of code, or actions.
Time to accepted expression How long does the complete system take to produce an accepted result? Milliseconds, seconds, or minutes with defined start and stop points.
Verification burden What automated or human work is required to confirm the expression is acceptable? Checks, tool calls, person-minutes, model calls, duration, or cost.
Correction burden What work is required when the initial expression does not pass? Retries, edits, additional calls, person-minutes, duration, or cost.

Expression Forms Must Be Evaluated on Their Own Terms

Natural-Language Output

Measure required-content coverage, factual support, constraint compliance, clarity, verification burden, and correction burden.

Structured Output

Validate schema, required fields, data types, identifiers, relationships, provenance, and downstream parseability.

Code or Tool Execution

Evaluate execution success, tests, side effects, permissions, retries, tool errors, and the accepted end state.

Decision or Recommendation

Evaluate evidence use, stated uncertainty, constraint compliance, required alternatives, and the registered decision standard.

Observable-evidence rule: The protocol does not need private chain-of-thought to evaluate expression. It can measure visible outputs, tool calls, stored state, retrieval behavior, execution results, verification records, corrections, and accepted outcomes.

Expression decision rule: Report changes in output size, action count, or completion time together with quality, unsupported additions, verification burden, and correction burden. A reduction in visible steps alone does not establish better expression.

CEMR Dimension Three

12. Memory Measures

Memory is reusable state that preserves information needed by a later operation and makes that information available through a defined retrieval pathway. The evaluation must test both what is stored and whether the correct state can be recovered and used when required.

Memory boundary: A transcript, cache, vector store, summary, database, registry, or state object is not automatically useful memory. Its value depends on fidelity, provenance, version accuracy, constraint preservation, retrievability, reuse, access control, and total operating cost.

Required Memory Properties

Identity

The stored state remains attached to the correct entity, task, user, system, record, or object.

Relationships

Dependencies, links, hierarchy, ownership, and other required relationships remain correct.

Provenance

The source, author, observation, transformation, and evidence history remain traceable where required.

Constraints

Permissions, exclusions, limits, guardrails, and stop conditions survive storage and retrieval.

Version State

The system can distinguish current, historical, superseded, corrected, and conflicting state.

Retrieval Pathway

The correct state can be found at the appropriate time without unacceptable ambiguity or search burden.

Measure Operational question Possible unit
Write fidelity Is the required state recorded correctly when memory is created or updated? Correct fields, write errors, field accuracy, or pass rate.
Retrieval precision Of the state retrieved, how much is relevant and correctly associated with the request? Precision, irrelevant records, or incorrect associations.
Retrieval recall Of the required stored state, how much is successfully recovered? Recall, missing records, or required-element recovery rate.
Constraint retention Do permissions, hard limits, exclusions, and stop conditions survive writing, storage, retrieval, and reuse? Retention rate, violation count, or severity.
Version correctness Does retrieval return the correct version and distinguish superseded or conflicting state? Version accuracy, stale-state rate, or conflict-resolution errors.
Collision resistance Can the system distinguish near-identical names, identifiers, records, thresholds, or related entities? Misassociation rate, collision errors, or fixture accuracy.
State survival How well does required state persist across turns, steps, updates, or recursive depth? Retention by step, turn, depth, duration, or update count.
Appropriate reuse Is stored state reused when relevant and withheld when irrelevant, expired, unauthorized, or superseded? Correct reuse, missed reuse, inappropriate reuse, or contamination rate.
Memory burden What does the memory system require to construct, store, index, retrieve, verify, update, secure, and delete state? Tokens, bytes, calls, duration, person-hours, energy, or financial cost.

Memory Test Families

  • Exact-state tests: Can the system recover registered identifiers, values, limits, and required text?
  • Relational tests: Are dependencies, ownership, hierarchy, and links preserved?
  • Provenance tests: Can the system connect a stored claim or state to its correct source and transformation history?
  • Constraint tests: Do hard limits, exclusions, permissions, and stop conditions survive?
  • Version tests: Can the system identify current, superseded, conflicting, or corrected state?
  • Collision tests: Can it distinguish similar identifiers and prevent cross-record contamination?
  • Negative-retrieval tests: Can it avoid retrieving irrelevant, unauthorized, expired, or deleted state?
  • Reuse tests: Can recovered state support a later accepted outcome without unnecessary reconstruction?

Privacy and lifecycle rule: Memory evaluation should include applicable access controls, retention limits, update rules, deletion behavior, confidentiality requirements, and the risk of state crossing an unauthorized user, task, project, or system boundary.

Memory decision rule: Report storage or retrieval reductions together with fidelity, provenance, constraint survival, version correctness, contamination, reconstruction burden, and total lifecycle cost. Stored data alone does not establish functional memory.

CEMR Dimension Four

13. Recursion Measures

Recursion is the repeated application of a process in which an output, decision, compressed representation, or preserved state becomes an input to a later operation. The evaluation asks whether required information and constraints remain usable as the system proceeds across turns, steps, cycles, branches, or recursive depth.

Recursion boundary: More recursion is not automatically better, and fewer steps are not automatically more efficient. A recursive process must be evaluated for accepted outcomes, state preservation, error propagation, stopping behavior, recovery burden, and total cost.

Register the Recursive Structure

Before testing, document what constitutes a recursive step, what state passes between steps, when source material is available again, how branches are created or closed, and what ends the process.

Depth

The number of registered transformations, turns, cycles, or nested operations through which state must persist.

State Passed Forward

The exact output, summary, capsule, record, memory object, or other state supplied to the next operation.

Refresh Cadence

When original source material, authoritative state, or another reference is reintroduced during repeated operations.

Branch Policy

The observable conditions under which alternatives are opened, compared, revisited, merged, rejected, or escalated.

Stopping Rule

The condition that ends, pauses, rejects, or escalates the recursive process.

Resource Budget

The registered limits on context, tokens, calls, time, storage, tools, human review, or other material resources.

Measure Operational question Possible unit
Accepted outcome by depth At which recursive depths does the system continue to pass the same quality gate? Acceptance rate by step, turn, cycle, or depth.
State retention by depth How well do required identity, relationships, provenance, constraints, and version state survive? Required-element retention by depth or cycle.
Error propagation Does an error remain local, get corrected, or propagate into later state and outputs? Downstream errors, affected steps, severity, or propagation rate.
Constraint survival Do hard limits, guardrails, exclusions, and stop conditions remain active across repeated operations? Retention rate, first-failure depth, violations, or severity.
Observable branching How many externally observable branches, alternatives, retries, or reversals occur? Branches, calls, retries, reversals, or state transitions.
Rework burden How much work is repeated because earlier state was incomplete, incorrect, ambiguous, or unavailable? Repeated calls, repeated retrievals, duration, tokens, or cost.
Stopping compliance Does the process stop, escalate, or continue according to the registered rule? Correct stops, premature stops, overruns, or missed escalations.
Refresh effect How does reintroducing authoritative source material affect retention, quality, and cost? Change by cadence, refresh count, source tokens, latency, or cost.
Recovery behavior Can the system recover from a detected error or state loss without violating constraints? Recovery rate, steps, calls, duration, or cost.

Refresh-Cadence and Stability Tests

Recursive evaluations may compare source availability at every step, source availability only at initialization, or preregistered refresh schedules. These tests can help identify where additional source access improves retention and where it introduces added cost or interference.

Earlier depth-limited benchmark work found that refresh behavior varied with fixture structure. Collision-heavy and constraint-heavy tasks did not produce the same cadence pattern. Those observations are exploratory, version-specific, and fixture-dependent; they do not establish a universal optimal refresh schedule.

Recursion decision rule: Report quality, state retention, error propagation, stopping behavior, refresh burden, recovery, and total cost at the tested depths. Do not generalize stability beyond the registered depth, workload, state budget, cadence, model, or configuration.

Whole-System Accounting

14. Total-Cost Measures

Total-cost evaluation includes the material work and resources required to produce, verify, correct, preserve, and maintain an accepted outcome. It prevents a reduction in one component from being reported as system-level efficiency when cost has merely moved elsewhere.

Primary Normalization

Total included cost per accepted outcome = Total included cost across attempted trials ÷ Number of accepted outcomes

This normalization should supplement—not replace—raw totals, attempted-task cost, acceptance rate, variation, and workload-level reporting.

Cost category Include where material Possible units
Input and context construction Prompt creation, context assembly, document preparation, data transformation, state serialization, and intervention instructions. Tokens, bytes, calls, duration, person-hours, or currency.
Model or compute calls Primary generations, routing, auxiliary models, evaluators, embeddings, rerankers, controllers, and recovery calls. Calls, tokens, accelerator time, compute proxy, duration, or currency.
Storage Transcripts, reusable state, caches, registries, logs, artifacts, checkpoints, backups, and retention duration. Bytes, byte-time, operations, duration, or currency.
Indexing Parsing, chunking, embedding, graph construction, metadata generation, updates, and re-indexing. Records, calls, tokens, duration, compute, or currency.
Retrieval Queries, search, filtering, reranking, graph traversal, data transfer, and retrieval failures. Queries, records, bytes, latency, compute, or currency.
Tool execution API calls, code execution, browsing, database operations, external services, and resulting data transfer. Calls, executions, duration, bytes, or currency.
Verification Automated tests, source checks, schema validation, comparison models, expert review, and user acceptance. Checks, calls, person-time, duration, or currency.
Human review Evaluation, annotation, adjudication, oversight, escalation, approval, and exception handling. Person-minutes, person-hours, tasks, or currency.
Correction Editing, re-prompting, state repair, tool recovery, source replacement, and re-verification. Edits, calls, person-time, duration, or currency.
Failed runs and retries Timeouts, invalid results, tool failures, rejected outputs, restarts, abandoned runs, and recovery attempts. Failures, retries, tokens, duration, compute, or currency.
Coordination Multi-agent handoffs, orchestration, scheduling, synchronization, conflict resolution, and result merging. Messages, calls, duration, person-time, or currency.
Infrastructure Compute, networking, databases, observability, security, orchestration, idle allocation, and shared services. Device-time, utilization, bytes, duration, energy, or currency.
Maintenance Monitoring, updates, prompt maintenance, index refresh, registry repair, evaluation upkeep, and operational support. Person-time, update count, duration, or currency.

Cost Reporting Rules

  • Report physical and operational units separately before converting costs into currency.
  • Identify direct measurements, calculations, inferences, proposed allocations, and unknown values.
  • Use the same cost categories and allocation rules for the baseline and intervention.
  • Separate one-time implementation costs from recurring operational costs.
  • State how shared infrastructure and maintenance costs are allocated.
  • Report both incremental intervention cost and complete system cost where possible.
  • Disclose excluded categories and explain whether their omission could change the conclusion.
  • Report results per attempted task and per accepted outcome, with acceptance rates visible.

Total-cost rule: Do not call an intervention more efficient because it improves one favorable measure. A defensible conclusion must include the quality gate, all material cost categories within the registered boundary, and any uncertainty or missing costs that could alter the result.

Physical Measurement Boundary

15. Environmental Telemetry Boundary

Environmental findings require direct physical telemetry or a defensible accounting model tied to the registered system boundary. Changes in tokens, calls, latency, memory, or financial cost may motivate environmental measurement, but they do not independently establish energy, cooling, water, carbon, or emissions outcomes.

Never convert token savings directly into energy, cooling, water, carbon, or emissions savings.

Environmental measure Required evidence Possible unit
Power Device, host, rack, or facility telemetry with sampling interval, idle treatment, utilization, and allocation method. Watts or kilowatts over a defined observation interval.
Energy Measured or defensibly integrated power consumption associated with the registered workload and complete included system. Joules, watt-hours, or kilowatt-hours.
Cooling Facility cooling design, operating conditions, allocation procedure, utilization, and relevant efficiency measurements. Energy, thermal load, or another defined physical measure.
Water Direct facility water data or a documented model covering cooling technology, location, operating conditions, allocation, and applicable energy-supply water use. Liters, gallons, withdrawal, or consumption under a defined boundary.
Operational emissions Measured energy combined with a documented location- or market-based emissions factor matched to geography and observation period. Grams or kilograms of CO2e.
Embodied impact Hardware inventory, manufacturing or lifecycle data, expected service life, utilization, and disclosed allocation method where material. Allocated CO2e, material mass, or another documented lifecycle unit.

Required Environmental Context

Hardware

Device type, count, memory, host configuration, allocation, and any shared infrastructure included in the measurement.

Utilization and Batching

Device utilization, idle allocation, batch size, concurrency, scheduling, and shared-workload treatment.

Geography and Time

Facility location, observation period, grid conditions, energy procurement, and time-dependent emissions factors.

Cooling Design

Air, evaporative, liquid, or other cooling design, facility efficiency, climate conditions, and water boundary.

Complete System

Model execution, retrieval, storage, networking, tools, verification, retries, and other material components within the registered boundary.

Allocation Method

How shared device, host, rack, facility, cooling, network, and embodied impacts are assigned to the evaluated workload.

Evidence Boundary by Measurement Availability

Available evidence Permitted interpretation
No physical telemetry or accounting model Environmental outcome is Unknown. Report only the directly observed computational or operational measures.
Device or host energy telemetry A scoped operational-energy finding may be possible for the measured boundary. Cooling, water, carbon, and facility-wide effects remain separate questions.
Facility and cooling allocation A broader operational finding may be calculated if the allocation method, uncertainty, location, and observation period are documented.
Energy plus matched emissions factors A scoped operational-emissions estimate may be calculated, subject to the stated geography, time, procurement, and allocation assumptions.
Lifecycle and embodied-impact data A broader lifecycle estimate may be possible when hardware inventory, service life, utilization, boundary, and uncertainty are included.

For the broader environmental claim framework, see Environmental Impact & Computational Ecology.

Environmental decision rule: Report environmental conclusions only for the directly measured or defensibly calculated boundary. When hardware, utilization, batching, geography, energy supply, cooling design, allocation, or another material factor is unavailable, identify the affected conclusion as Unknown.

Observable Evidence Layer

16. Instrumentation and Logging

Instrumentation should capture the observable events needed to reconstruct each trial, verify the registered boundary, calculate the selected metrics, identify failures, and distinguish measured behavior from inference.

The baseline and intervention must use equivalent logging procedures. If instrumentation adds material latency, storage, model calls, or compute, that overhead should be measured or disclosed.

Minimum Traceability Requirement

Every attempted trial should have a stable evaluation ID, workload ID, task ID, arm assignment, configuration reference, start and end time, outcome status, quality decision, metric record, failure record, and link to the available evidence.

Log family Recommended fields Boundary notes
Trial identity Evaluation ID, task ID, attempt ID, arm, sequence position, repetition, seed, and parent-child trace relationships. Identifiers should be stable without exposing protected user or customer information.
Configuration snapshot Model, version, endpoint, system prompt, controller, decoding, context limit, tools, retrieval, memory, routing, and infrastructure references. Use versioned records or checksums where full configuration text cannot be stored with the result.
Inputs and outputs Registered input, supplied context, observable output, structured fields, completion status, token accounting, and truncation status. Redact or reference protected content according to the registered privacy procedure.
Retrieval events Query, index version, filters, returned identifiers, rank, scores, bytes, latency, and retrieval errors. Record enough provenance to determine what evidence was available to each arm.
Tool and execution events Tool name and version, request, response status, duration, side effects, errors, retries, and execution artifacts. Sensitive arguments or results may require redaction, encryption, or restricted access.
Memory and state events State written, retrieved, updated, superseded, rejected, expired, or deleted; version; provenance; permissions; and retrieval pathway. The record should permit collision, stale-state, contamination, and constraint-retention analysis.
Observable recursion events Step or cycle, depth, branch, refresh event, state passed forward, stop decision, escalation, recovery, and externally visible reversal. Do not infer or claim hidden chain-of-thought from observable traces.
Quality and adjudication Evaluator identity or system, rubric version, dimension scores, pass or fail decision, disagreement, adjudication, and correction burden. Preserve blinding where preregistered and disclose where evaluator blinding was unavailable.
Resource and cost data Tokens, calls, memory, storage, indexing, retrieval, tools, latency, verification, human work, retries, infrastructure, and maintenance allocations. Identify the source, unit, calculation method, and missing categories.
Physical telemetry Device, host, rack, facility, energy, cooling, water, geography, utilization, batching, and allocation fields when environmental claims are evaluated. If telemetry is unavailable, the affected environmental conclusion remains Unknown.
Failure and deviation record Timeouts, service failures, missing data, configuration drift, protocol deviations, safety stops, invalid trials, exclusions, and reasons. Failed and excluded trials should remain traceable according to the preregistered procedure.

Logging Controls

  • Synchronize clocks or document known timestamp offsets across distributed components.
  • Use a versioned data dictionary that defines every field, unit, status, and missing-value code.
  • Record configuration snapshots or stable checksums for each experimental arm.
  • Preserve raw records separately from derived calculations and summary tables.
  • Apply equivalent sampling and logging rates to the baseline and intervention.
  • Document redaction, aggregation, retention, deletion, encryption, and access-control procedures.
  • Measure or disclose instrumentation overhead and any differences between arms.
  • Protect the original record from silent revision and retain a visible amendment history.

Logging sufficiency rule: If missing, inconsistent, or asymmetric logs prevent verification of the registered prediction or system boundary, the affected finding should be marked Unknown or the evaluation classified as inconclusive or invalid according to the preregistered rule.

Experimental Design

17. Sample Design and Repeated Trials

The sample must be large and representative enough to evaluate the preregistered prediction at the required level of precision. There is no universal number of trials appropriate for every model, workload, metric, or evaluation design.

Sample size and repetition should reflect outcome variability, expected effect size, quality thresholds, acceptable uncertainty, task diversity, system stochasticity, and the consequences of an incorrect conclusion.

Paired Comparisons

Where appropriate, run the same registered task under the baseline and intervention so task difficulty is matched across arms.

Repeated Trials

Repeat stochastic tasks enough times to characterize variation rather than relying on a favorable individual run.

Stratification

Preserve workload classes, difficulty levels, context lengths, tool dependencies, or other registered subgroups in allocation and reporting.

Randomized Order

Randomize or counterbalance trial order when time, load, caching, warm-up, or system drift could affect results.

Deterministic Coverage

For deterministic fixtures, register complete cases, boundary values, collisions, failure states, and expected outputs rather than relying only on random sampling.

Independent Replication Set

Where feasible, reserve unseen tasks, later observation periods, or independent fixtures for confirmation beyond the development sample.

Sample-Design Registration

Design field Required decision
Sampling population Define the task population, inclusion criteria, exclusion criteria, source, and intended scope of inference.
Sample size Justify the number of tasks and repetitions using expected variability, precision, power, deterministic coverage, or another registered rationale.
Allocation Specify paired, randomized, blocked, stratified, crossover, sequential, or other arm-allocation procedure.
Repetition and seeds Register repetitions, seed handling, temperature or stochastic settings, and whether identical seeds can be meaningfully compared.
Trial order Define randomization, counterbalancing, warm-up treatment, cache handling, and protection against time-dependent bias.
Subgroups Register task classes, difficulty strata, context lengths, models, domains, or other subgroups that will be analyzed separately.
Stopping rule Set completion, safety, futility, cost, missing-data, or sequential-analysis rules before results are observed.
Uncertainty reporting Define confidence intervals, credible intervals, deterministic coverage, sensitivity analysis, or another appropriate uncertainty method.
Multiplicity State how multiple metrics, comparisons, workloads, subgroups, or repeated looks at the data affect the decision rule.

Development, Pilot, and Confirmatory Samples

  • Development sample: Used to construct or tune the implementation. Performance on this sample should not be treated as independent confirmation.
  • Pilot sample: Used to test instrumentation, estimate variability, refine feasibility, and identify failure modes.
  • Confirmatory sample: Evaluated under the locked preregistration and used for the formal prediction decision.
  • Replication sample: Repeats the claim under a new evaluation event, ideally with an independent team, new sample, later period, or additional system boundary.

Sampling-integrity rule: Do not stop when the result becomes favorable, add repetitions only to one arm, discard unfavorable seeds, or redefine the confirmatory sample after inspection. Any deviation must be recorded and its effect on the evidentiary status disclosed.

Falsifiability and Stop Rules

18. Failure Conditions

Failure conditions must be defined before results are observed. They specify what would count against the prediction, what would invalidate the comparison, what requires the evaluation to stop, and what leaves the result inconclusive.

A protocol that cannot record failure cannot provide a credible empirical test.

Prediction Failure

The test was validly executed, but the intervention did not satisfy the preregistered quality, success, or total-cost decision rule.

Invalid Evaluation

A material protocol, boundary, implementation, data-integrity, or comparison failure prevents the registered question from being tested as designed.

Stop Condition

A preregistered safety, privacy, cost, operational, or futility threshold requires trials to pause or end.

Inconclusive Evaluation

The evidence is insufficient, ambiguous, conflicting, too uncertain, or too incomplete to make the registered decision.

Condition class Examples Required treatment
Quality-gate failure The intervention falls below the correctness, safety, usefulness, constraint, schema, or memory-fidelity threshold. Record the efficiency prediction as failed for that evaluation, even if resource use decreased.
Success-threshold failure The primary metric does not reach the preregistered direction, magnitude, uncertainty, equivalence, or non-inferiority requirement. Report the prediction decision without substituting a favorable secondary metric.
Total-cost failure A local reduction is offset by intervention overhead, verification, correction, retries, human review, storage, or maintenance. Report the local improvement and total-cost result separately.
State or recursion failure Required identity, provenance, constraint, version state, or stopping behavior degrades beyond the registered threshold. Record the first-failure condition, affected depth or task class, severity, and downstream consequences.
Boundary or configuration drift Model, prompt, controller, tools, retrieval, workload, hardware, or acceptance rules change during the evaluation. Pause, document the deviation, and determine whether affected trials require separation or a new evaluation.
Intervention-integrity failure The intended Razor intervention was absent, inconsistently applied, or materially different from its registered version. Do not interpret the result as a valid test of the registered intervention.
Baseline-integrity failure The baseline is weakened, receives unequal access, or is subjected to materially different conditions. Classify causal attribution as invalid unless the imbalance can be corrected through a registered procedure.
Evidence failure Logs, telemetry, provenance, quality records, cost data, or task outcomes are missing or asymmetric. Mark affected findings Unknown or the evaluation inconclusive or invalid, according to the registered rule.
Safety or privacy stop Unauthorized access, sensitive-data exposure, unsafe execution, harmful side effects, policy breach, or uncontrolled external action occurs. Stop affected trials, preserve appropriate evidence, contain the issue, and follow the registered incident procedure.
Insufficient precision The observed uncertainty is too large to distinguish success, failure, equivalence, or material tradeoff under the decision rule. Report the evaluation as inconclusive unless a preregistered continuation rule applies.

Failure Is Part of the Result

Failed predictions, unfavorable tradeoffs, invalid trials, stop events, and inconclusive evidence should remain visible in the result record. They can identify boundary conditions, expose instrumentation gaps, improve later preregistration, and prevent unsupported generalization.

Classification rule: Prediction failure, invalid evaluation, stop condition, and inconclusive evidence are different outcomes. Record them separately and do not automatically convert any one of them into a formal evidence state until the registered evidence-state review is completed.

Reproducible Evaluation Output

19. Result-Record Format

Every completed evaluation should produce a versioned result record that connects the original preregistration to the observed data, calculations, findings, failures, deviations, and applicable evidence-state review.

The record should preserve favorable, unfavorable, mixed, failed, excluded, and unscorable outcomes. Summary language must remain traceable to the underlying trials and should not replace the raw or minimally processed evidence.

Required Record Linkage

Claim or Question → Preregistration → Trials → Evidence → Findings → Replication → Evidence-State Review

Record section Required fields Purpose
Record identity Result-record ID, title, version, date, evaluator, organization, authorship, review status, and amendment history. Creates a stable and attributable evaluation object.
Canonical relationship Governing authority, claim-register reference, claim version, tested proposition, and applicable framework requirements. Prevents a local result from silently redefining the canonical claim.
Preregistration Registration ID, version, timestamp, prediction, primary outcome, quality threshold, success threshold, failure conditions, and analysis plan. Separates confirmatory decisions from post-result interpretation.
Evaluation unit System, version, configuration, workload, constraints, measurement period, evidence boundary, baseline, and intervention. Defines the exact scope to which the result applies.
Sample and execution Sample size, task classes, allocation, repetitions, seeds, order, dates, exclusions, early stops, and completed trials. Shows how the registered design was executed.
Quality-gate results Accepted, failed, corrected, excluded, and unscorable outcomes; dimension scores; adjudication; verification burden; and correction burden. Determines whether resource comparisons qualify for efficiency interpretation.
CEMR results Compression, expression, memory, and recursion measures, units, uncertainty, subgroup results, and first-failure conditions. Preserves the four dimensions as separate measurable findings.
Total-cost results Raw totals, cost categories, cost per attempt, cost per accepted outcome, allocation methods, excluded costs, and missing values. Prevents local reductions from being mistaken for system-level efficiency.
Environmental results Physical telemetry, calculation model, hardware, utilization, batching, geography, energy supply, cooling, allocation, units, and uncertainty. Limits environmental conclusions to the measured or defensibly calculated boundary.
Failures and deviations Prediction failures, invalid trials, stop conditions, missing data, protocol deviations, configuration drift, safety events, and corrective actions. Keeps adverse and limiting evidence visible.
Finding statements Atomic statement, finding label, supporting record, calculation or inference method, scope, uncertainty, limitation, and author. Separates documented observations from calculations, inferences, proposals, and unknowns.
Decision and interpretation Primary prediction decision, secondary results, exploratory observations, limitations, transfer restrictions, and alternative explanations. Prevents exploratory interpretation from replacing the preregistered decision.
Replication and review Code or asset references, replication status, independent reviewer, conflicts of interest, evidence-state recommendation, and unresolved questions. Connects one evaluation to the broader evidence process.

Required Result Summary

  1. What was tested? Identify the complete evaluation unit and intervention.
  2. What was predicted? State the preregistered primary prediction and thresholds.
  3. Did quality pass? Report acceptance rates and material regressions.
  4. What changed? Report CEMR, resource, and total-cost results in their defined units.
  5. What failed or remained unknown? Preserve failures, missing evidence, deviations, and uncertainty.
  6. What can be concluded? State the narrowest result supported by the evidence.
  7. What comes next? Identify replication, correction, or additional measurement requirements.

Result-record integrity rule: Preserve the original result record and add later corrections, replications, reviews, or evidence-state decisions as versioned updates. Do not silently overwrite an unfavorable or superseded result.

Statement-Level Classification

20. Finding Labels

Finding labels describe how an individual statement in an audit or result record was produced. They help readers distinguish source-backed observations from calculations, interpretations, proposals, and unresolved questions.

Do not substitute informal confidence labels. Use only Documented, Calculated, Inferred, Proposed, or Unknown for individual finding statements.

Finding Label

Documented

The statement is directly supported by an identified source, log, record, observation, artifact, or measurement output.

Boundary: Documentation establishes that a record exists; it does not automatically establish that the record is complete, accurate, independently verified, or sufficient for a Supported claim.

Finding Label

Calculated

The statement results from an explicit calculation applied to documented inputs using a disclosed formula, unit, allocation method, or transformation.

Boundary: Report assumptions, uncertainty, missing inputs, and whether the calculation is sensitive to allocation choices.

Finding Label

Inferred

The statement is an interpretation derived from documented or calculated evidence rather than a directly observed fact.

Boundary: Identify the reasoning, competing interpretations, uncertainty, and evidence that could challenge the inference.

Finding Label

Proposed

The statement presents a hypothesis, design, expected mechanism, prediction, target, recommendation, or future test that has not yet been established by the current evidence.

Boundary: A proposal should identify how it could be tested and what would count against it.

Finding Label

Unknown

The available evidence does not support a reliable documented, calculated, or inferred statement within the registered boundary.

Boundary: Unknown is the correct label when telemetry, provenance, units, configuration, scope, or another required input is unavailable.

Finding-Statement Format

Statement: The smallest independently assessable observation, calculation, interpretation, proposal, or unknown.

Label: Documented, Calculated, Inferred, Proposed, or Unknown.

Support: Source, log, trial, formula, or evidence reference.

Scope: System, version, configuration, workload, constraints, and observation period.

Limitations: Missing data, uncertainty, assumptions, transfer limits, and alternative interpretations.

Keep Statements Atomic

If one sentence combines a documented observation, a calculated change, and an inferred explanation, divide it into separate findings. Each component should receive its own label and supporting evidence.

Combined statement to avoid Separated findings
“The intervention used fewer output tokens and therefore reduced energy and improved efficiency.” Documented: Output-token totals differed.

Calculated: The registered percentage change in output tokens.

Unknown: Energy change without direct telemetry or a defensible accounting model.

Inferred or Unknown: Total efficiency, depending on quality and complete cost evidence.

Separation rule: Finding labels classify individual statements. They do not assign the formal status of a canonical or empirical claim. A finding labeled Documented is not automatically a Supported claim.

Claim-Level Classification

21. Evidence-State Decision Process

Formal evidence states describe the current evidentiary status of a defined claim. They apply at the claim level, not to an isolated metric, observation, sentence, implementation, or audit finding.

Use only these formal evidence states: Proposed, Testing, Provisionally Supported, Supported, Challenged, Inconclusive, and Retired.

Formal evidence state Meaning Decision boundary
Proposed The claim has been stated in a testable form but has not yet completed the evidence process required for support. The claim should identify its scope, predictions, required evidence, and failure conditions.
Testing A registered evaluation, replication, or evidence-collection process is actively examining the claim. Starting a test does not imply that the expected result has occurred.
Provisionally Supported Available evidence supports the claim under a defined scope, but material limitations, limited replication, boundary uncertainty, or unresolved questions remain. The supporting system, workload, version, conditions, and limitations must remain visible.
Supported The claim has met the governing evidentiary requirements under its stated scope, including applicable quality, protocol, reproducibility, and adverse-evidence review. Supported does not mean universal, permanent, or transferable beyond the evaluated claim boundary.
Challenged Credible evidence, replication failure, boundary failure, contradiction, or methodological concern materially conflicts with the claim or its present scope. The challenge, supporting evidence, affected scope, and required resolution should be recorded.
Inconclusive The available evidence is insufficient, ambiguous, mixed, conflicting, too uncertain, or methodologically incomplete for a support or challenge decision. Inconclusive should identify what additional evidence or correction is required.
Retired The claim is no longer active because it was withdrawn, superseded, merged, replaced, or removed from the current framework. The historical record and reason for retirement should remain available where appropriate.

Evidence-State Review Sequence

  1. Identify the claim. Record the exact claim, version, scope, and canonical relationship.
  2. Verify preregistration. Confirm that prediction, quality threshold, success threshold, failure conditions, and analysis plan were registered before result inspection.
  3. Verify protocol integrity. Review system boundaries, baseline fairness, implementation integrity, sampling, logging, deviations, and missing data.
  4. Review quality and total cost. Confirm that accepted outcomes and material system costs were evaluated before efficiency conclusions were made.
  5. Review findings. Separate Documented, Calculated, Inferred, Proposed, and Unknown statements.
  6. Review replication and adverse evidence. Include independent results, failures, boundary conditions, contradictions, and alternative explanations.
  7. Apply the narrowest justified scope. Do not extend the decision to untested models, workloads, domains, scales, or environments.
  8. Record the decision. Add the evidence state, decision date, reviewer, supporting record, limitations, and re-evaluation conditions to the claims register.

Finding Labels and Evidence States Are Different Systems

  Finding labels Formal evidence states
Applied to An individual observation, calculation, interpretation, proposal, or unknown. A defined and versioned claim.
Question answered How was this statement produced? What is the claim’s current evidentiary status?
Examples Documented, Calculated, Inferred, Proposed, Unknown. Proposed, Testing, Provisionally Supported, Supported, Challenged, Inconclusive, Retired.
Automatic conversion? No. A Documented or Calculated finding does not automatically create a formal claim decision. No. State assignment requires the complete governing review process.

Decision-authority rule: A favorable pilot, implementation, benchmark result, publication, license, or author-created evaluation does not independently confer the state Supported. Evidence-state changes must follow the governing claim-review process under GC-MRD-v2.0 and remain recorded in the Canonical Claims Register.

Reproducibility and External Review

22. Independent Replication

Replication tests whether a result can be observed again under a registered relationship to the original evaluation. It checks whether the finding depends on one evaluator, dataset, implementation, measurement period, configuration, or undocumented condition.

Independence boundary: Tests, tools, examples, implementations, and analyses created by Robbie George or by a collaborating team should be identified as author-created or collaborative. They do not become independent replication merely because they use a new dataset or are published in a public repository.

Replication Types

Computational Reproduction

Re-executes the original analysis using the same or preserved data, code, configurations, and result rules.

Tests: Whether the published or preserved result can be regenerated.

Direct Replication

Repeats the registered claim with equivalent methods and boundaries using a new sample, later observation period, or independently executed trial.

Tests: Whether the result recurs under substantially equivalent conditions.

Conceptual Replication

Tests the same underlying proposition using a different operationalization, implementation, or measurement method.

Tests: Whether the proposition survives a meaningful change in method.

Domain-Transfer Test

Evaluates a new claim boundary involving another model, workload, domain, scale, language, or operating environment.

Tests: A new transfer proposition—not automatic replication of the original scope.

Minimum Replication Record

Field Required disclosure
Target claim Exact claim, version, original scope, evidence state, and original result-record reference.
Replication type Computational reproduction, direct replication, conceptual replication, or domain-transfer test.
Evaluator relationship Authorship, funding, licensing, collaboration, implementation assistance, data access, conflicts of interest, and degree of independence.
Preserved elements Prediction, primary metric, quality threshold, success threshold, failure conditions, and other elements held equivalent to the original evaluation.
Changed elements Model, version, workload, sample, configuration, implementation, infrastructure, observation period, evaluator, or measurement method.
Replication prediction Preregistered result expected if the original finding or transfer proposition holds.
Outcome comparison Original and replication effects, quality, uncertainty, failures, costs, boundary differences, and alternative explanations.
Evidence-state implication Whether the result supports continued testing, provisional support, support, challenge, inconclusive status, or another formal review—without assigning the state automatically.

Replication Outcomes

  • Consistent: The replication satisfies its preregistered decision rule within the defined relationship to the original claim.
  • Partially consistent: Some registered outcomes recur, while others differ or reveal narrower boundary conditions.
  • Not reproduced: The replication does not satisfy the registered reproduction or replication threshold.
  • Inconclusive: Missing data, uncertainty, protocol limitations, or conflicting results prevent a decision.
  • Invalid replication: A material execution or integrity failure prevents the registered replication question from being tested.

Replication-review rule: One consistent replication does not make a result universal, and one failed replication does not automatically retire a claim. Both favorable and unfavorable replications must be evaluated for integrity, scope, method, and boundary differences before the formal evidence state is reviewed.

Applicability Beyond the Original Test

23. Domain-Transfer Testing

A result is bounded by the system, version, configuration, workload, constraints, observation period, and evidence available in the original evaluation. Applying that result elsewhere creates a new transfer proposition that must be registered and tested.

Success in one model, workload, domain, scale, or environment cannot be generalized to another without a new test.

Transfer Dimensions

Changed dimension Examples New question created
Model or architecture Different provider, model family, parameter scale, architecture, modality, weights, or release. Does the registered effect occur in the new model boundary?
Workload Question answering, code, retrieval, planning, agent execution, classification, memory, or multimodal tasks. Does the effect survive a different task structure and acceptance standard?
Knowledge domain General knowledge, science, law, medicine, finance, ecology, operations, or proprietary enterprise data. Do domain vocabulary, evidence, risk, and quality requirements change the result?
Language or modality Another natural language, code language, image, audio, video, structured data, or mixed modality. Does the implementation preserve quality and CEMR behavior in the new representation?
Scale and duration Longer context, greater recursive depth, more users, higher concurrency, larger registry, or longer operating period. Do retention, cost, latency, failure, and maintenance behavior remain acceptable at the new scale?
Tools and integrations Different APIs, databases, retrieval systems, memory stores, controllers, or execution environments. Does the effect remain when dependencies, failure modes, and tool costs change?
Operational environment Research, production, edge, cloud, private deployment, real-time system, or human-in-the-loop workflow. Do real-world constraints, infrastructure, review, and maintenance change the conclusion?
Risk and governance Different safety, privacy, legal, regulatory, audit, or approval requirements. Does the system meet the new quality gate and evidence requirements?

Transfer-Test Registration

  1. Identify the original result. Cite its claim, result record, evidence state, scope, and limitations.
  2. Identify what changes. Record every material difference between the original and proposed transfer boundary.
  3. Identify what remains invariant. State which CEMR mechanism, requirement, or relationship is expected to persist.
  4. Register the transfer prediction. Define the expected result, quality threshold, success threshold, and failure conditions.
  5. Use a new system boundary. Register the transferred model, workload, configuration, constraints, and evidence boundary.
  6. Measure new costs and risks. Do not reuse operational, environmental, privacy, or maintenance assumptions from the original environment.
  7. Record the transfer result separately. Preserve the original result rather than expanding its scope retroactively.

Possible Transfer Outcomes

Transfer Observed

The new evaluation meets its registered decision rule within the transferred boundary.

Boundary-Limited

The result appears only in specified task classes, scales, configurations, or operating conditions.

Transfer Not Observed

The registered prediction does not pass under the new boundary.

Inconclusive Transfer

Uncertainty, missing evidence, implementation differences, or protocol limitations prevent a reliable transfer decision.

Transfer rule: Calling the protocol architecture-agnostic means it can be applied across different architectures. It does not mean that a result obtained from one architecture automatically applies to another.

Protected Evaluation Boundary

24. Data Privacy and Confidential Evaluations

AI labs may evaluate Robbie’s Razor using proprietary models, private workloads, confidential telemetry, protected customer data, or security-sensitive infrastructure. The protocol permits confidential execution, but privacy restrictions and evidence availability must be registered before testing.

Confidentiality changes who can inspect the evidence; it does not change the need for a defined system boundary, preregistration, quality gate, total-cost accounting, failure record, and reproducible decision process.

Data Minimization

Collect only the inputs, outputs, telemetry, identifiers, and metadata required to answer the registered evaluation question.

Purpose Limitation

Define whether evaluation data may be used only for the registered pilot or also for debugging, model improvement, publication, replication, or later research.

Access Control

Identify who can access raw data, prompts, outputs, state, logs, configuration, telemetry, result records, and publication summaries.

Retention and Deletion

Register how long data and derived state will be retained, how updates and legal holds are handled, and how deletion will be verified.

Security Controls

Document applicable encryption, secret handling, isolation, authentication, audit logging, incident response, and external-service restrictions.

Publication Rights

Define who owns the raw data, code, configurations, derived results, finding statements, reports, and publication decisions.

Confidential Evaluation Record

Field Required decision
Data classification Classify public, internal, confidential, restricted, personal, regulated, security-sensitive, or otherwise protected data.
Permitted processing Identify approved models, regions, services, tools, storage systems, subprocessors, and human reviewers.
Prohibited processing Identify disallowed external transfers, training use, retention, publication, cross-customer access, or reuse.
Redaction and pseudonymization Define which identifiers, secrets, content, prompts, outputs, or logs must be removed or transformed before review.
Reviewer access State whether an independent reviewer can inspect raw evidence, a protected environment, redacted records, aggregates, or only a summary.
Retention schedule Set retention periods for raw data, logs, state, calculations, reports, backups, and reproducibility artifacts.
Incident procedure Define containment, notification, evidence preservation, trial suspension, remediation, and reauthorization requirements.
Disclosure boundary Specify which methods, aggregate results, limitations, conflicts, and evidence-state recommendations may be made public.

Reproducibility Without Public Disclosure

When raw data or configurations cannot be released, the evaluation may preserve reproducibility through one or more controlled methods:

  • time-stamped private preregistration or third-party escrow;
  • version identifiers and cryptographic checksums for protected artifacts;
  • public schemas, metric definitions, and analysis code without private records;
  • redacted or pseudonymized trial records;
  • synthetic fixtures that reproduce the tested failure or state pattern;
  • controlled reviewer access within a secure evaluation environment;
  • aggregate results with disclosed sample size, uncertainty, exclusions, and limitations; or
  • an independent attestation describing what was inspected and what remained unavailable.

Reasoning-privacy rule: Private chain-of-thought is not required. Evaluators should use observable inputs, outputs, tool calls, stored state, retrieval behavior, execution results, telemetry, accepted outcomes, and other authorized evidence.

Confidentiality rule: Do not claim compliance with a law, regulation, security standard, or contractual requirement merely because privacy controls are described on this page. Applicable obligations must be identified and verified for the specific organization, data, jurisdiction, and evaluation.

Public Reproducibility Layer

25. GitHub Benchmark Assets

The Robbie’s Razor Benchmarks repository is the public reproducibility layer for applicable benchmark code, controlled fixtures, schemas, documentation, and result assets.

The repository does not replace the Grand Compression Master Reference Document or the Canonical Claims Register. GC-MRD-v2.0 remains the governing authority, while GitHub preserves materials that help others inspect, reproduce, challenge, or extend an evaluation.

MRD → Claims Register → Auditor → Lab Protocol → Benchmark Hub → GitHub Assets → Result Record → Replication → Evidence-State Review

Protocol Asset Classes

Public or private implementations of this protocol may preserve the following asset classes:

Preregistration Assets

Prediction templates, system-boundary manifests, workload registrations, quality thresholds, failure conditions, and analysis plans.

Controlled Fixtures

Deterministic tests for identity, relationships, provenance, constraints, version state, collisions, retrieval, and recursive stability.

Instrumentation Schemas

Versioned definitions for trials, configurations, logs, CEMR measures, total costs, failures, and environmental telemetry.

Evaluation Code

Scripts, tests, validators, calculation methods, environment records, dependency versions, and reproducible execution instructions.

Result Records

Versioned records linking preregistration, trials, evidence, finding labels, failures, calculations, interpretations, and review status.

Replication Assets

Replication registrations, preserved and changed conditions, outcome comparisons, deviations, negative results, and unresolved questions.

Minimum Asset Metadata

  • stable asset identifier and descriptive title;
  • asset type, version, status, author, and modification date;
  • governing-authority reference and applicable claim identifier;
  • system, workload, configuration, and evidence boundary;
  • dependencies, environment, license, and execution instructions;
  • input and output schemas with units and required fields;
  • quality, failure, missing-data, and total-cost treatment;
  • finding labels and applicable formal evidence state;
  • known limitations, domain-transfer restrictions, and supersession history.

Public Technical Reference

Robbie’s Razor: A Scale-Invariant Recursion Principle for Efficient Intelligence — Preprint v1.0, first publicly released January 1, 2026. The preprint provides technical and historical context; GC-MRD-v2.0 remains the current governing authority.

Repository-alignment rule: Presence in GitHub demonstrates publication or implementation—not effectiveness. Before treating any repository asset as current, verify its GC-MRD-v2.0 authority reference, finding-label enum, evidence-state enum, quality-gate fields, system boundary, CEMR measures, total-cost fields, failure conditions, domain-transfer restrictions, and environmental telemetry boundary.

Evaluation Pathway

26. How to Begin a Pilot

A pilot should begin with a narrow, falsifiable question—not a broad promise of efficiency. Select one system boundary and one operationally meaningful workload where quality, state, resource use, failures, and total cost can be observed.

Diagnostic Question → Registered Prediction → Feasibility Pilot → Confirmatory Test → Result Record → Replication → Evidence-State Review

Step 1

Define the Decision

Identify what operational decision the lab will make from the evaluation and which evidence is required to make it responsibly.

Step 2

Run a Diagnostic Review

Use the Razor Auditor framework to identify boundary gaps, quality risks, cost omissions, and candidate CEMR questions.

Step 3

Register the Evaluation Unit

Lock the system, version, configuration, workload, constraints, measurement period, evidence boundary, baseline, and intervention.

Step 4

Preregister the Prediction

Define primary and secondary metrics, units, quality threshold, success threshold, failure conditions, analysis plan, and missing-data procedure.

Step 5

Run a Feasibility Pilot

Verify instrumentation, data access, quality adjudication, intervention integrity, privacy controls, cost capture, and operational safety.

Step 6

Lock the Confirmatory Test

Separate pilot tuning from the confirmatory sample, preserve the registration, and execute the locked comparison.

Step 7

Create the Result Record

Report quality, CEMR, total cost, failures, deviations, finding labels, limitations, and the preregistered prediction decision.

Step 8

Replicate Before Transfer

Seek reproduction or independent replication before extending the result to another model, workload, domain, scale, or environment.

Experimental Diagnostic Interface

The interactive Razor Auditor can help teams organize preliminary questions and identify evidence gaps. It is an experimental diagnostic interface—not certification, formal validation, or an evidence-state decision.

Canonical and Applied Pathways

Propose a Scoped Evaluation

Organizations may use this public protocol independently or contact Robbie George to discuss a private, evaluation-only pilot. Commercial integration or licensing should follow—not replace—successful testing and appropriate review.

Contact Robbie George

Pilot-status rule: Beginning a pilot does not imply endorsement, adoption, licensing, successful performance, compliance, certification, or a favorable evidence-state decision.

Protocol Questions

27. Frequently Asked Questions

These answers summarize the protocol’s principal evidence, measurement, privacy, and claim boundaries.

Does the Razor Evaluation Protocol assume that Robbie’s Razor works?

No. The protocol does not assume that a Razor-guided implementation is more efficient or effective. It defines how that proposition can be tested under equivalent, preregistered conditions and permits favorable, unfavorable, mixed, failed, and inconclusive outcomes.

What is the required unit of evaluation?

The required unit is System + Version + Configuration + Workload + Constraints + Measurement Period + Evidence Boundary. A material change to one of these elements may create a new evaluation unit.

Do fewer tokens prove that a system is more efficient?

No. A token reduction is one measured change. It does not establish total efficiency unless the intervention passes the same quality gate and all material costs—including construction, model calls, retrieval, tools, verification, correction, failures, retries, human review, infrastructure, and maintenance—are included.

Can token savings be converted into energy, water, carbon, or emissions savings?

Not automatically. Environmental conclusions require direct telemetry or a defensible accounting model that includes hardware, utilization, batching, geography, energy supply, cooling design, allocation, and measurement boundaries. Without that evidence, the environmental result is Unknown.

Does the protocol require access to private chain-of-thought?

No. The protocol evaluates observable inputs, outputs, tool calls, stored state, retrieval behavior, execution results, telemetry, accepted outcomes, and authorized evidence. It does not require or claim access to hidden chain-of-thought.

What counts as useful memory under the protocol?

Useful memory is reusable state that preserves relevant identity, relationships, provenance, constraints, version state, and retrieval pathways. A transcript, cache, vector store, summary, database, or registry is not automatically useful memory merely because it stores data.

What is the difference between a finding label and a formal evidence state?

A finding label—Documented, Calculated, Inferred, Proposed, or Unknown—describes how an individual statement was produced. A formal evidence state—Proposed, Testing, Provisionally Supported, Supported, Challenged, Inconclusive, or Retired—describes the current evidentiary status of a defined claim. A Documented finding is not automatically a Supported claim.

Can an AI lab run the protocol privately?

Yes. A lab may keep models, workloads, telemetry, and results confidential while preserving a time-stamped preregistration, controlled evidence, versioned result record, and appropriate review pathway. Confidentiality restrictions and unavailable evidence must be disclosed when conclusions are reported.

Does publishing code or a benchmark result on GitHub validate Robbie’s Razor?

No. GitHub is a public reproducibility layer. Publishing code, fixtures, schemas, or results demonstrates implementation and access. Validation requires a registered evaluation, acceptable evidence, appropriate replication, and the separate formal evidence-state review.

Can a successful result be generalized to another model or workload?

No. A successful result remains limited to its registered boundary. Applying it to another model, version, workload, domain, scale, language, tool environment, or operating context creates a new transfer proposition that requires a new test.

What happens if the intervention fails the test?

The failure should remain in the result record with its affected workload, boundary, metrics, costs, and failure conditions. Failed predictions can reveal limits, improve later test design, and inform a Challenged or Inconclusive evidence-state review, but they should not be hidden or redefined after observation.

Does completing the protocol provide certification or formal compliance?

No. Completing the protocol produces a scoped evaluation record. It does not create certification or formal compliance unless a separate documented process, applicable requirements, authorized reviewer, and explicit decision exist.

How should an organization begin?

Begin with one narrow diagnostic question, define the complete evaluation unit, select a fair baseline, register the quality gate and prediction, verify instrumentation through a feasibility pilot, and then run a locked confirmatory test. The Razor Auditor and Benchmark Hub can help organize the initial evaluation pathway.

Authorship and Framework Origin

28. About the Author

Robbie George, creator of Robbie’s Razor and the Grand Compression Cosmology.

Robbie George

Robbie George

Robbie George is a National Geographic–published nature photographer, author, and creator of the Grand Compression Cosmology, Robbie’s Razor, and the Razor Auditor evaluation framework.

His work examines how compression, expression, memory, and recursion can be described across reasoning, knowledge systems, ecology, and artificial intelligence. The Grand Compression Master Reference Document serves as the governing authority, while the Canonical Claims Register, Razor Auditor, this Lab Evaluation Protocol, and the Benchmark Hub define the public evaluation pathway.

His background in nature photography and ecological observation informs the system-level perspective of the work while remaining separate from the empirical testing required for technical claims about AI systems.

Author-and-validation disclosure: Robbie George is the originator of Robbie’s Razor and the associated evaluation framework. Author-created tools, examples, implementations, benchmark fixtures, and evaluations should be distinguished from independent testing and replication. Authorship does not itself establish that a technical claim is Supported.

Razor Evaluation Protocol · Document version 2.0 · Governing authority: GC-MRD-v2.0

Trusted Art Seller

Trusted Art Seller

The presence of this badge signifies that this business has officially registered with the Art Storefronts Organization and has an established track record of selling art.

It also means that buyers can trust that they are buying from a legitimate business. Art sellers that conduct fraudulent activity or that receive numerous complaints from buyers will have this badge revoked. If you would like to file a complaint about this seller, please do so here.

Verified Returns & Exchanges

Verified Returns & Exchanges

The Art Storefronts Organization has verified that this business has provided a returns & exchanges policy for all art purchases.

Description of Policy from Merchant:

What is your Policy on Returns/Exchanges/Refunds? I take great pride in my work and prints, and I want you to be completely happy with your investment in my nature art. If for any reason you are unsatisfied with your print, you may return it within 14 days of delivery, and/or exchange it for another print. Prints must be returned in new condition, packaged carefully in the original packaging if possible. Your refund will be issued as soon as I receive the returned print. Please contact me if you would like to arrange a return or exchange. In the event that you receive a damaged or defective print, please let me know within 7 days of receipt, and I will arrange for a new print to be shipped to you at no additional cost.

Verified Secure Website with Safe Checkout

Verified Secure Website with Safe Checkout

This website provides a secure checkout with SSL encryption.

Verified Archival Materials Used

Verified Archival Materials Used

The Art Storefronts Organization has verified that this Art Seller has published information about the archival materials used to create their products in an effort to provide transparency to buyers.

Description from Merchant:

Fine Art Prints are made with high-quality archival inks on fine art papers using a high-resolution large format inkjet printer. Our premium archival inks produce images with smooth tones and rich colors. Prints are made with care on your choice of exquisite Fine Art Papers using a high-resolution large format inkjet printer. https://www.graphikprintworks.com

Cart

Your cart is currently empty.

Saved Successfully.

This is only visible to you because you are logged in and are authorized to manage this website. This message is not visible to other website visitors.

Import From Instagram

Click on any Image to continue

This Website Supports Augmented Reality to Live Preview Art

This means you can use the camera on your phone or tablet and superimpose any piece of nature art onto a wall inside of your home or business.

To use this feature, Just look for the "Live Preview AR" button when viewing any piece of nature art on this website!

Red fox pouncing through snow

Pounce Now—Save 20% on Your First Order

Join the collector list for your first-order discount, new wildlife releases, and occasional field notes.

No thanks