A structured diagnostic and evaluation interface for examining how a declared AI system uses compression, expression, memory, recursion, infrastructure, and total computational cost.
The Auditor organizes public evidence or directly supplied measurements into bounded findings. It helps identify what is documented, what can be calculated, what remains inferred, and what cannot be determined without controlled telemetry.
Current boundary: The Razor Auditor does not automatically certify systems, assign company-wide compliance, calculate energy from tokens, infer hidden reasoning, or prove adoption of Robbie’s Razor. Formal evidence requires the Lab Evaluation Protocol and an appropriate benchmark record.
The Razor Auditor is a structured method for examining a specifically identified AI system through Robbie’s Razor. It turns available evidence into a transparent profile of system behavior, resource use, memory, reuse, recursive stability, infrastructure dependency, and unanswered questions.
The Auditor may be used with public disclosures, user-supplied documentation, controlled telemetry, or formal benchmark results. The strength of its output must never exceed the strength of the evidence supplied.
RC-01 · Canonical Robbie’s Razor
“When competing explanations exist, prefer the model that follows compression → expression → memory → recursion.”
The Correct Unit of Audit
The minimum valid audit subject is not an entire company or industry. It is a declared system operating within a declared boundary:
System + Version + Configuration + Workload + Constraints + Measurement Period + Evidence Boundary
What the Auditor Can Do
organize evidence by source and confidence;
map observations to compression, expression, memory, and recursion;
identify measurable efficiency mechanisms;
distinguish documentation from inference;
identify missing telemetry and unresolved questions;
define testable predictions and failure conditions;
prepare a system for formal benchmarking.
What It Cannot Establish Automatically
company-wide Robbie’s Razor compliance;
adoption, licensing, affiliation, or endorsement;
hidden reasoning or internal memory behavior;
energy, water, emissions, or JCT from tokens alone;
universal stability from one observed workload;
formal evidence states without protocol compliance;
independent validation from self-reported claims.
Audit Finding Labels
DocumentedCalculatedInferredProposedUnknown
These labels describe the basis of an individual audit finding. They are separate from the formal GC-MRD-v2.0 evidence states: Proposed, Testing, Provisionally Supported, Supported, Challenged, Inconclusive, and Retired.
What Changed from the Earlier Auditor
Earlier Pattern
Current GC-MRD-v2.0 Standard
Company-wide labels such as Razor-aligned, transitional, or brute-force
Dimension-specific findings bounded to a system, version, workload, and evidence set
Collapse-risk conclusions from public information
Declared failure conditions and explicitly unknown internal thresholds
JCT or environmental conclusions inferred from tokens
Direct measurement or defensible physical attribution before environmental claims
Interactive output treated as an audit verdict
Interactive output treated as diagnostic intake that may escalate to formal evaluation
Current Page Classification
Applied diagnostic and evaluation-intake framework. The Razor Auditor helps structure evidence and generate testable next steps. Formal support requires controlled benchmarking and scope-specific replication.
The word “audit” can describe several levels of evaluation. The Razor Auditor must identify which mode is being used so that an exploratory public profile is not mistaken for a measured or independently replicated result.
Maturity rule: Audit modes describe access and evaluation depth. They are not quality grades, compliance levels, or automatic evidence states.
Mode 1
Public-Source Profile
Uses filings, technical documentation, model cards, company disclosures, published research, and attributable secondary sources.
Valid output: documented mechanisms, public dependencies, evidence gaps, unknown telemetry, and testable questions.
Mode 2
Structured Diagnostic
Adds user-supplied architecture, configuration, workload, metric, cost, or operational information without a complete controlled test.
Valid output: a more detailed audit profile and formal evaluation plan. Unverified user assertions remain labeled as supplied information.
Mode 3
Measured Evaluation
Uses a preregistered protocol, matched baseline, controlled task set, direct telemetry, shared quality threshold, and versioned result package.
Valid output: calculated deltas and a bounded evidence-state recommendation for the tested claim.
Mode 4
Independent Replication
A separate evaluator reconstructs or reimplements the declared evaluation without relying on undocumented decisions from the original team.
Valid output: a replication record that may strengthen, narrow, challenge, or leave the original result inconclusive.
Mode Comparison
Mode
Direct Telemetry
Matched Baseline
Formal Evidence State
Primary Use
Public-Source Profile
Usually unavailable
No
No
Evidence mapping
Structured Diagnostic
Partial or supplied
Optional
No
Evaluation design
Measured Evaluation
Required for claimed metrics
Required
Eligible
Bounded testing
Independent Replication
Required for claimed metrics
Required
Eligible
Evidence strengthening or challenge
Interactive-interface boundary: A chatbot, Gem, worksheet, or guided form ordinarily operates in Mode 1 or Mode 2. It reaches Mode 3 only when it is connected to a preregistered protocol, matched baseline, direct measurements, and a versioned result record.
Every Razor audit begins with an evidence register. Each source must be identified, dated, scoped, and linked to the finding it supports so that documentation, calculation, inference, and unknowns remain distinguishable.
Evidence ceiling: An audit conclusion cannot be more certain, more current, or broader than the sources and measurements supporting it.
Preferred Evidence Hierarchy
Priority
Evidence Type
Valid Use
Primary Limitation
1
Direct measurement and controlled telemetry
Measured workload-specific findings
Access, attribution, calibration, and system boundary
2
Audited filings and regulatory disclosures
Documented organizational facts and reported risks
Usually lacks task-level technical telemetry
3
Technical documentation, system cards, specifications, and research papers
Architecture, configuration, benchmark, and design claims
May be self-reported or configuration-specific
4
Official company statements and product announcements
Attributed plans, mechanisms, targets, and company-reported results
Not independent validation
5
Independent research, testing, and replication
External comparison and validation
May use different versions or conditions
6
Journalism and reputable secondary analysis
Context, leads, and attributed reporting
Requires verification against primary sources
7
User-supplied descriptions or estimates
Diagnostic intake and hypothesis generation
Unverified unless documentation or measurement is supplied
Minimum Source Record
Source Identity
Title, author, publisher, URL, file, or dataset identifier.
Publication Date
When the evidence was released and when it was accessed.
System Coverage
The model, product, site, workload, version, and period covered.
Supported Finding
The exact audit statement for which the source is being used.
Source Type
Direct, audited, technical, company-reported, independent, secondary, or supplied.
Limitations
Missing boundaries, conflicts, uncertainty, incentives, or version mismatch.
Evidence Handling Rules
prefer primary sources for factual and technical claims;
attribute company-reported performance to the company;
separate announced, planned, contracted, operating, and utilized capacity;
record the exact product and version whenever available;
preserve conflicting evidence rather than silently selecting one source;
label calculated findings with the formula and source values used;
label inferred findings and explain the reasoning boundary;
recheck time-sensitive facts before publishing or updating an audit; and
do not treat licensing, payment, inclusion, or publication as evidence.
Confidentiality warning: Do not paste trade secrets, personal data, security-sensitive architecture, credentials, private customer information, or restricted telemetry into a public interactive auditor. Use an approved controlled environment for non-public evaluations.
Reasoning privacy: The Auditor does not require private chain-of-thought. It can evaluate outputs, state transitions, tools, retries, memory behavior, latency, task quality, and resource use without requesting hidden reasoning content.
The audit model evaluates a system in sequence. It begins with accepted task quality, moves through compression, expression, memory, and recursion, and then measures total cost and evidence confidence.
Sequence rule: A system cannot claim efficiency by reducing resources while failing the declared output-quality threshold. Quality gates the interpretation of every later dimension.
The Auditor reports each dimension separately. A system may show documented compression mechanisms, unknown memory behavior, measurable output quality, and unresolved total cost at the same time.
Finding
The bounded statement supported by the available evidence.
Basis
Documented, calculated, inferred, proposed, or unknown.
Scope
The system, version, workload, configuration, and period covered.
Limitation
What the evidence cannot establish.
No automatic label: The current Auditor does not collapse mixed findings into company-wide labels such as “Razor-aligned,” “transitional,” or “brute-force.” Any formal classification must be defined by a protocol and limited to the tested system boundary.
The Quality Gate: Efficiency Begins After Acceptance
The Razor Auditor does not treat a shorter, faster, or less expensive result as more efficient unless it first satisfies the same declared acceptance standard as the comparison result. Quality is the gate through which every efficiency claim must pass.
This protects the audit from rewarding systems that reduce visible computation by omitting requirements, lowering accuracy, transferring work to a human reviewer, or increasing downstream correction costs.
Quality-gate rule: compare efficiency only among outputs that meet the same preregistered acceptance threshold under equivalent conditions.
Define Acceptance Before Testing
The evaluation record should define what counts as an accepted result before the test is run. Depending on the task, acceptance may include:
Correctness
Does the result agree with verified facts, calculations, expected outputs, or domain-specific ground truth?
Completeness
Were all required components, constraints, edge cases, and deliverables addressed?
Safety
Did the system remain within the declared safety, privacy, security, and operational boundaries?
Usefulness
Can the result be used for its intended purpose without substantial uncounted repair or reinterpretation?
Constraint Compliance
Did the result follow the required format, scope, tools, latency limit, budget, and other declared conditions?
Verification Burden
How much human or machine effort was required to verify, correct, complete, or safely deploy the result?
Accepted Sources of Judgment
Judge Type
Appropriate Use
Required Record
Ground Truth
Tasks with known answers or independently verifiable outcomes
Dataset, answer key, version, and scoring procedure
Domain Expert
Specialized tasks requiring professional judgment
Rubric, reviewer qualifications, conflicts, and disagreement handling
Automated Evaluation
Repeatable checks, schemas, unit tests, or documented model-assisted scoring
Evaluator version, prompt or test suite, threshold, and validation limits
Downstream Outcome
Tool execution, code deployment, decisions, or measurable real-world results
Outcome definition, observation window, failure handling, and confounders
Operational Measures
The following measures may be used within a declared evaluation protocol. They are operational audit measures—not universal canonical constants:
Accepted-task rate: accepted results divided by all attempted tasks.
Repair burden: human and machine effort required after the initial result.
Cost per accepted task: total measured cost divided by the number of accepted results.
Time to accepted result: elapsed time through generation, verification, correction, and final acceptance.
Failure rate: rejected, incomplete, unsafe, or unusable results divided by all attempts.
Quality-Gate Failure Conditions
The compared systems are evaluated against different quality thresholds.
Failed or abandoned runs are removed from the result set.
The evaluator, rubric, or pass threshold is changed after results are observed.
Shorter output is assumed to be better without testing correctness and completeness.
Human correction, verification, or downstream repair is excluded from the comparison.
Next: Once equivalent accepted outcomes are established, the audit can examine whether the system preserved task-relevant structure with less unnecessary expansion.
Within the Razor Auditor, compression means reducing unnecessary representational burden while preserving the structure required to produce an accepted result. It is not synonymous with shorter text, fewer tokens, smaller files, or lower latency.
Useful compression preserves what the task needs. If a shorter representation removes constraints, relationships, provenance, identity, or retrieval pathways that must later be reconstructed, the apparent saving may be displacement rather than efficiency.
Compression question: What task-relevant structure was preserved, what redundancy was removed, and what cost was introduced by compression, retrieval, decompression, verification, or repair?
Compression Audit Questions
What raw information, context, records, or state enters the system?
How is that information represented before and during execution?
Which details, relationships, and constraints must survive compression?
What information is removed, summarized, indexed, cached, or externalized?
Can the compressed structure be retrieved and used without ambiguity?
Does the process reduce repeated expansion across multiple accepted tasks?
What compression, retrieval, decompression, and verification costs are added?
Does the apparent benefit persist when quality and total cost are held constant?
Representation Types to Inspect
Active Context
Prompts, retrieved documents, conversation history, instructions, schemas, and other material loaded for the task.
Expression Analysis: Turn Preserved Structure Into Usable Output
Expression tests whether preserved or compressed structure can generate a correct, usable result under the declared task conditions. A compact representation has limited value if it cannot reliably produce the answer, action, tool call, decision, or control signal the task requires.
The Razor Auditor therefore evaluates expression as an observable system behavior. It does not infer private chain-of-thought or claim access to hidden internal reasoning.
Expression question: Can the available structure produce an accepted result accurately, efficiently, and within the required operational constraints?
Expression Forms
Natural-Language Output
Answers, explanations, summaries, reports, and recommendations.
Structured Data
JSON, tables, classifications, database records, and schema-constrained output.
Code and Plans
Executable code, workflows, procedures, specifications, and implementation plans.
Tool Actions
Searches, API calls, retrieval operations, transactions, and environment actions.
Decisions
Rankings, selections, diagnoses, escalations, routing choices, and recommendations.
Control Signals
Commands or outputs that directly influence machinery, software, or operational systems.
Expression Audit Questions
What observable output or action is required?
Which preserved information, constraints, and relationships must the output express?
Does the result satisfy the quality gate on the first attempt?
Does the output comply with the required format, schema, tool, or execution boundary?
How much correction, retrying, reformatting, or human intervention is required?
Does the system introduce details that are unsupported by its evidence or stored state?
Can the same representation support consistent expression across equivalent tasks?
What latency, compute, tool, and verification costs are incurred before acceptance?
Representation Is Not Execution
Layer
Audit Question
Example Evidence
Available Structure
Was the required information actually present and retrievable?
Context records, retrieved sources, memory objects, schemas, and versioned state
Generated Expression
Did the system convert that structure into the required output?
Responses, generated code, structured output, decisions, or proposed actions
Executed Outcome
Did the output work when validated or executed?
Unit tests, tool results, deployment checks, downstream outcomes, and expert review
Implementation is not validation. Producing an answer, plan, or tool call demonstrates expression only. It does not establish correctness, effectiveness, or support for a broader theory until the result passes the declared evaluation process.
Candidate Measures
Accepted output rate and first-attempt acceptance rate.
Schema, format, and constraint-compliance rate.
Successful tool execution or downstream validation rate.
Time and measured cost to an accepted result.
Retries, corrections, regenerations, and human interventions per task.
Unsupported assertions or details not grounded in registered evidence.
Consistency across repeated or equivalent task conditions.
Expression Failure Conditions
The available structure cannot produce an actionable or accepted result.
The output contradicts registered evidence, constraints, or preserved state.
Formatting, schema, or tool-call errors prevent execution.
Repeated correction or regeneration offsets the apparent efficiency gain.
A generated output is treated as validated merely because it was produced.
Next: The audit examines whether useful structure and accepted outcomes are preserved as durable memory—and whether that memory can support later recursive reuse.
Memory Analysis: Preserve Structure for Reliable Reuse
Memory determines whether useful structure survives beyond the immediate task. The Razor Auditor examines what is preserved, how it is identified, where it is stored, how it is retrieved, and whether later systems can reuse it without reconstructing the original work.
A transcript, cache, database, summary, or vector store is not automatically effective memory. It becomes useful memory only when its retained structure can be located, interpreted, verified, and applied within the declared scope.
Memory question: Does the system preserve enough verified structure to support later retrieval and reuse without uncontrolled loss, ambiguity, drift, or reconstruction cost?
Minimum Structure to Preserve
Under the current Grand Compression framework, a reusable memory object should preserve the following elements when they are relevant to the task:
Identity
What the object, entity, decision, result, or state is—and how it is uniquely distinguished.
Relationships
How the preserved object connects to sources, dependencies, entities, prior states, and downstream uses.
Provenance
Where the information came from, who or what produced it, and which evidence supports it.
Constraints
The scope, assumptions, permissions, exclusions, conditions, and limits governing valid reuse.
Version State
The version, modification history, supersession status, and time boundary of the preserved structure.
Retrieval Pathways
The identifiers, indexes, links, queries, or resolution methods needed to find and restore the structure.
Memory Forms the Auditor Can Inspect
Memory Form
Potential Value
Audit Concern
Active Context
Immediate access during a task or session
Temporary state may disappear or become too large to manage efficiently
Conversation or Event History
Chronological record of prior interactions
Raw history may preserve volume without preserving usable structure
Structured State
Compact storage of entities, decisions, constraints, and relationships
Schema omissions or version drift can distort later reuse
Retrieval Index
Locates relevant material without loading the full record
Retrieval errors can hide relevant evidence or surface the wrong version
Reusable Artifact
Preserves validated code, plans, templates, findings, or procedures
Later use may occur outside the artifact’s tested scope
External Registry or Knowledge Base
Supports shared identity, provenance, versioning, and resolution
Authority, update frequency, access, and conflict resolution must be documented
Memory Audit Questions
What structure is preserved after an accepted task is completed?
Does the memory retain identity, relationships, provenance, constraints, version state, and retrieval pathways?
Can the preserved state be retrieved under defined later conditions?
Does retrieval return the correct object, source, and version?
Can the retrieved structure support an accepted result without reconstructing the original task?
What storage, indexing, synchronization, retrieval, validation, and maintenance costs are incurred?
How are conflicting, stale, incomplete, or superseded memories handled?
Does the memory remain useful outside the original evaluation window?
Candidate Measures
Successful retrieval rate under preregistered conditions.
Correct-version retrieval rate.
Provenance and constraint retention rate.
Time and cost required to retrieve usable state.
Accepted-task performance with memory compared with the declared baseline.
Reconstruction avoided across repeated tasks.
Stale, conflicting, missing, or incorrectly resolved memory events.
Storage, indexing, synchronization, validation, and maintenance burden.
Memory Failure Conditions
Stored information cannot be reliably found or restored.
Identity, provenance, relationships, or governing constraints are lost.
Outdated state is returned without version or supersession warnings.
Retrieved memory produces lower-quality results than the declared baseline.
The cost of storage and retrieval exceeds the measured reconstruction avoided.
A stored artifact is called reusable without testing it on a later task.
Private, restricted, or sensitive information is retained outside its authorized boundary.
Next: Once structure can be preserved and retrieved, the audit tests whether later operations can reuse that memory across repeated state transitions.
Recursion Analysis: Reuse Structure Across State Transitions
Recursion begins when the result or preserved state from one operation becomes a governed input to a later operation. The Razor Auditor tests whether this reuse improves task performance or reduces repeated work without allowing error, ambiguity, or cost to accumulate across cycles.
Repeated execution alone is not recursion in the audit sense. A system must preserve and reuse relevant structure across identifiable state transitions. The audit does not assume that recursion produces learning, self-improvement, autonomy, or intelligence unless those properties are independently defined and measured.
Recursion question: Can verified structure from a prior accepted operation be reused in a later operation while preserving quality, traceability, constraints, and total-cost discipline?
Define the Recursive Unit
Every recursion analysis should identify the state transition being tested:
Prior State → Retrieval → New Input → Updated Expression → Validation → Preserved Successor State
The transition should declare which elements may change, which must remain invariant, which evidence is inherited, and what conditions cause the successor state to be rejected.
Recursive Patterns to Inspect
Iterative Refinement
A verified prior result is revised under new feedback, constraints, or evidence.
Repeated Workflow
A validated plan, template, or procedure is reused across comparable tasks.
Stateful Interaction
A system carries governed state across sessions, events, or sequential decisions.
Tool Feedback Loop
Observed tool or environment results become evidence for the next operation.
Cross-Agent Handoff
One process preserves state for another process to retrieve and continue.
Versioned Knowledge Update
New evidence modifies an existing state while preserving provenance and history.
Recursion Audit Questions
What prior state is being reused, and how was it validated?
What new input, evidence, or constraint causes the next transition?
Which identities, relationships, provenance records, and constraints must remain invariant?
What is allowed to change in the successor state?
Does reuse reduce measured reconstruction, latency, or total cost?
Does quality remain above the same accepted-result threshold across cycles?
How are errors, conflicts, uncertainty, and failed transitions contained?
Can an evaluator reproduce the transition from the registered inputs and state?
Candidate Measures
Accepted successor-state rate across declared transitions.
Prior structure successfully reused per accepted task.
Reconstruction or repeated expansion avoided.
Quality change across successive cycles.
Error propagation and correction rate.
Constraint, provenance, and version retention across transitions.
Human intervention required per cycle.
Total cost per accepted transition and across the complete recursive sequence.
Evidence Boundary
Observed reuse supports only the tested system, version, workload, conditions, and measurement period. A successful recursive workflow does not by itself demonstrate general intelligence, autonomous improvement, cross-domain transfer, or support for every claim in the Grand Compression framework.
Recursion Failure Conditions
Prior state is reused without evidence that it remains valid.
Errors or unsupported assumptions compound across successive cycles.
Identity, provenance, constraints, or version state drift during handoff.
Repeated execution is mislabeled as recursion without preserved state reuse.
Quality declines as the number of transitions increases.
Coordination, retrieval, and correction costs exceed the reconstruction avoided.
A successful implementation is presented as proof of self-improvement or general intelligence.
Next: The auditor combines quality, computation, memory, retrieval, infrastructure, correction, and operational burden into a total-cost record.
Total-Cost Analysis: Count the Full Path to an Accepted Result
The Razor Auditor evaluates cost across the complete path to an accepted result. A reduction in one visible measure—such as output tokens or generation time—does not establish net efficiency if cost is transferred to retrieval, storage, verification, retries, human review, infrastructure, or downstream repair.
Total-cost analysis keeps the audit focused on measured system behavior rather than a single convenient proxy.
Total-cost question: What resources were consumed from task intake through verification, correction, acceptance, preservation, and any required later retrieval?
Cost Categories
Category
What to Count
Example Units
Input and Context
Prompts, retrieved material, preprocessing, ingestion, and context assembly
Tokens, bytes, documents, requests, time, or direct monetary cost
Generation and Compute
Inference, processing, model calls, tool selection, and execution
Latency, accelerator time, CPU time, requests, joules, or measured cost
Storage and Memory
State creation, indexing, persistence, synchronization, and retention
Bytes, storage duration, transactions, memory usage, or direct cost
Retrieval
Search, ranking, resolution, loading, and failed retrieval attempts
Queries, latency, data transferred, compute, or direct cost
Verification
Testing, fact-checking, expert review, automated judging, and validation
Reviewer time, test runs, evaluator calls, latency, or direct cost
Correction and Failure
Retries, repair, rollback, rework, rejected outputs, and abandoned runs
Failed attempts, correction time, compute, or financial loss
Coordination
Agent handoffs, orchestration, communication, queueing, and state transfer
Messages, calls, latency, transferred data, or direct cost
Infrastructure
Serving, networking, databases, cooling, and platform overhead when directly measured
Energy, water, hardware utilization, bandwidth, or provider cost
Normalize Cost by Accepted Outcomes
A defensible comparison should report cost relative to accepted work rather than successful-looking output alone.
Cost per accepted task = total measured evaluation cost ÷ number of accepted tasks
This is an operational evaluation expression. Each audit must define its cost boundary, units, acceptance criteria, and observation period.
Comparison Requirements
Use the same task set, workload distribution, and acceptance threshold.
Register the baseline before examining comparative results.
Use equivalent system boundaries and observation periods.
Include failed, retried, rejected, and abandoned attempts.
Document caching, batching, routing, quantization, retrieval, and other configuration differences.
Report unavailable telemetry as Unknown rather than estimating without a defensible method.
Separate directly measured values from calculated, inferred, or proposed values.
Tokens, Energy, Water, and Emissions
Token counts may help describe workload, but they do not independently establish energy, water, cooling, or emissions impact. Those outcomes depend on hardware, utilization, batching, model architecture, data-center conditions, regional energy supply, cooling design, and the measurement boundary.
Environmental conclusions require direct telemetry or a clearly documented, defensible accounting method. Where that evidence is unavailable, the environmental finding remains Unknown.
Required Cost Record
System, model, version, configuration, and provider.
Workload, sample size, task conditions, and measurement period.
Baseline and comparison configuration.
Accepted-result definition and judging method.
Measured cost categories and excluded categories.
Units, instrumentation, calculations, assumptions, and uncertainty.
Failed runs, missing telemetry, and known confounders.
Finding label, evidence state when applicable, and scope limitation.
Total-Cost Failure Conditions
Only a favorable proxy is reported while material costs are excluded.
Failed or rejected runs are omitted from the denominator.
Systems are compared using different quality thresholds or workload conditions.
Cost is transferred to human review, retrieval, or repair without being counted.
Token reduction is presented as measured energy or environmental reduction.
Provider pricing is treated as a complete measure of physical resource use.
Unknown values are replaced with unsupported estimates.
Next: The audit separates documented facts, calculations, inferences, proposals, and unknowns—then assigns formal evidence states only where a registered test supports them.
Evidence and Confidence: Separate Observation From Conclusion
Every Razor Auditor output must distinguish what was directly documented, what was calculated, what was inferred, what remains proposed, and what is unknown. This prevents a plausible interpretation from being presented as a measured result.
Audit finding labels describe the basis of an individual statement. Formal evidence states describe the status of a registered claim after an appropriate evaluation. These systems serve different purposes and should not be merged.
Evidence rule: confidence must follow the registered evidence—not the reputation of the system, the persuasiveness of the explanation, or the auditor’s preference.
Audit Finding Labels
Label
Meaning
Required Support
Documented
Explicitly reported or directly observed in an identified source.
Source, date, version, quotation or data location, and applicable scope
Calculated
Derived from registered inputs using a disclosed calculation.
Input values, units, equation, assumptions, uncertainty, and reproducible method
Inferred
A reasoned interpretation that is not directly documented or measured.
Evidence basis, reasoning boundary, alternatives, uncertainty, and limitations
Proposed
A hypothesis, design, recommendation, or testable expectation.
Declared prediction, intended test, success threshold, and failure condition
Unknown
The available evidence is insufficient for a defensible finding.
Description of the missing telemetry, documentation, comparison, or validation
Formal Evidence States
A formal evidence state applies only when a claim has been registered with a testable prediction, baseline, metrics, thresholds, failure conditions, and result record.
Proposed
The claim is registered but has not yet entered an adequate test.
Testing
The preregistered evaluation is underway and no final result has been assigned.
Provisionally Supported
Initial results meet the declared threshold but require more testing or replication.
Supported
Results meet the registered standard with sufficient evidence and applicable replication.
Challenged
Material counterevidence or a failed test conflicts with the claim.
Inconclusive
The evaluation cannot resolve the claim because of insufficient power, conflicting results, or material uncertainty.
Retired
The claim is withdrawn, superseded, or no longer maintained as an active claim.
Confidence Record
Each material finding should identify:
The finding label and evidence source.
The system, version, workload, and measurement period.
The applicable CEMR or total-cost dimension.
The observed result and declared comparison baseline.
Uncertainty, missing telemetry, and known confounders.
Alternative explanations consistent with the evidence.
The conditions under which the finding would no longer apply.
The formal evidence state, if—and only if—a registered claim was tested.
Example of Correct Separation
Documented finding: the evaluated configuration retrieved a stored task record during 83 of 100 measured trials.
Calculated finding: the observed retrieval rate was 83% under the registered test conditions.
Inferred finding: retrieval failures may be associated with inconsistent identifiers.
Evidence state: assigned only to the registered claim evaluated by those trials—not to the entire system or organization.
Evidence Failure Conditions
An inference is presented as a documented or measured fact.
A calculation omits units, inputs, assumptions, or uncertainty.
An evidence state is assigned without a preregistered claim and evaluation.
A result from one configuration is generalized to an entire company, model family, or domain.
Unknown telemetry is replaced with unsupported estimates.
Tool output, publication, licensing, or implementation is treated as validation.
Audit Output: A Scoped, Reproducible Evaluation Record
A Razor Auditor report should explain exactly what was evaluated, which evidence was available, how each finding was derived, and where the audit must stop. The output is dimensional and evidence-bounded—not a single company-wide score or label.
The same organization may show strong measured memory reuse in one workflow, unknown infrastructure cost in another, and unsupported public claims in a third. Those findings should remain separate.
Reporting rule: describe the evaluated system and conditions first; present findings second; state limitations and unknowns before drawing conclusions.
Required Audit Header
Audit Identifier
Unique identifier and report version
Audit Mode
Public-Source Profile, Structured Diagnostic, Measured Evaluation, or Independent Replication
System Boundary
System, version, configuration, tools, infrastructure, and included components
Workload
Task set, sample size, workload distribution, constraints, and observation period
Baseline
Registered comparison system, configuration, or prior condition
Quality Gate
Acceptance criteria, threshold, judging method, and disagreement procedure
Evidence Boundary
Sources, telemetry, unavailable data, exclusions, assumptions, and confidentiality limits
Dimensional Findings Matrix
Dimension
Finding
Basis
Scope
Limitation
Quality Gate
Accepted, rejected, or unknown
Rubric and results
Tested workload
Judge and sample limits
Compression
Dimension-specific finding
Finding label and evidence
Representation tested
Unmeasured overhead
Expression
Dimension-specific finding
Finding label and evidence
Output or action tested
Execution limits
Memory
Dimension-specific finding
Finding label and evidence
Stored state tested
Duration and retrieval limits
Recursion
Dimension-specific finding
Finding label and evidence
Transitions tested
Transfer limits
Total Cost
Measured delta or Unknown
Telemetry and calculation
Declared cost boundary
Excluded cost categories
Required Report Sections
Executive summary: the evaluated system, mode, strongest finding, principal limitation, and next test.
Scope and boundary: what the audit includes and explicitly excludes.
Evidence register: sources, telemetry, versions, access dates, and finding labels.
Quality-gate results: accepted-task criteria and observed performance.
CEMR findings: separate compression, expression, memory, and recursion records.
Total-cost record: measured categories, units, exclusions, and cost per accepted task.
Unknowns and alternatives: missing evidence and competing explanations.
Failure conditions: results that would challenge or invalidate the conclusion.
Evidence state: included only for claims tested under a registered protocol.
Escalation pathway: the measurements or replication needed next.
Standard Finding Format
Finding: State the narrow result.
Label: Documented, Calculated, Inferred, Proposed, or Unknown.
Basis: Identify the evidence, measurement, or disclosed calculation.
Scope: Identify the system, version, configuration, workload, and period.
Limitation: State uncertainty, excluded factors, and conditions preventing broader use.
Prohibited Summary Shortcuts
Do not classify an entire company as Razor-aligned, transitional, or brute-force.
Do not convert dimensional findings into an unsupported universal score.
Do not call missing telemetry a low score; record it as Unknown.
Do not imply that implementation, licensing, publication, or tool use establishes validation.
Do not generalize one successful task, configuration, or domain beyond its evidence boundary.
Choose the audit pathway that matches the evidence you can provide. A public-source or interactive diagnostic can identify questions and unknowns. A measured evaluation can test a preregistered claim. Independent replication can examine whether the result persists outside the original evaluation environment.
Begin with the narrowest system boundary that can be described and tested honestly.
Pathway 1
Interactive Diagnostic
Use the experimental Razor Auditor interface to organize public or user-supplied information into CEMR questions, possible findings, and evidence gaps.
Output: diagnostic profile. Tool output is not certification or formal validation.
Important: Do not submit credentials, private model data, regulated records, personal information, confidential business material, or proprietary source code to a public interactive interface. Use an approved controlled environment for restricted evaluations.
Audit Examples: From Broad Claims to Testable Systems
A useful audit begins by narrowing a broad description into a specific system, configuration, workload, and evidence boundary. The following hypothetical examples demonstrate how to frame Razor Auditor questions without assigning unsupported conclusions.
Important: These examples illustrate audit design only. They are not findings about any named company, model, platform, or product.
Hypothetical Example 1
Customer-Support Retrieval System
Broad claim: “The new support assistant is more efficient because it uses fewer tokens.”
Audit correction: Token reduction alone does not establish efficiency. The comparison must first hold answer quality constant and include retrieval, verification, correction, and escalation costs.
System Boundary
Specified assistant version, retrieval configuration, support corpus, tool access, and escalation workflow
Workload
A registered sample of billing, account, troubleshooting, and policy questions
Quality Gate
Correct resolution, policy compliance, citation accuracy, safety, and expert-review threshold
Baseline
Previous assistant configuration evaluated on the same task set
Candidate CEMR Tests
Compression: Does retrieval reduce irrelevant context while preserving applicable policy and customer constraints?
Expression: Does the system produce an accepted response or action without additional repair?
Memory: Are prior verified resolutions preserved with provenance, version state, and retrieval pathways?
Recursion: Can a later support interaction reuse valid state without carrying forward obsolete or incorrect information?
Decision rule: report a measured efficiency improvement only if the new configuration meets the same quality gate and lowers total cost per accepted resolution.
Hypothetical Example 2
Code-Generation Workflow With Reusable Memory
Broad claim: “Repository memory prevents the coding agent from repeating work.”
Audit correction: The existence of stored state does not demonstrate successful memory or reuse. Retrieval accuracy, code quality, correction burden, and reconstruction avoided must be tested.
System Boundary
Agent version, repository commit, memory service, tool permissions, test environment, and model configuration
Passing tests, accepted review, security checks, scope compliance, and absence of regressions
Baseline
Identical agent and tools without access to the reusable memory layer
Candidate CEMR Tests
Compression: Does the memory object preserve relevant repository structure without repeatedly loading unnecessary history?
Expression: Does the retrieved state support code that passes the same registered test suite?
Memory: Are decisions linked to their source, commit, scope, constraints, and supersession state?
Recursion: Can later tasks reuse accepted structure without propagating stale implementation assumptions?
Decision rule: measure whether memory access reduces reconstruction and total cost per accepted change without increasing regressions or review burden.
Hypothetical Example 3
Multi-Step Research Agent
Broad claim: “The agent’s recursive process produces better research.”
Audit correction: Multiple steps do not automatically demonstrate useful recursion. The audit must identify preserved state, state transitions, validation checkpoints, and error propagation.
System Boundary
Agent workflow, model version, search tools, source policy, memory layer, evaluator, and final-report generator
Workload
Registered research questions with known evidence sources and expert-review criteria
Quality Gate
Citation accuracy, source quality, completeness, contradiction handling, and expert acceptance
Baseline
Single-pass or non-stateful workflow using the same model and source access
Candidate CEMR Tests
Compression: Does the system preserve decisive evidence without repeatedly expanding the full source set?
Expression: Does preserved evidence generate a complete report with accurate citations?
Memory: Are claims linked to source identity, provenance, constraints, and retrieval pathways?
Recursion: Do validation results correct successor states, or do early errors compound across the workflow?
Decision rule: compare accepted-report quality and total cost while separately recording retrieval failures, citation errors, correction cycles, and human review.
Failure Conditions: When an Audit Must Stop, Narrow, or Escalate
Failure conditions protect the Razor Auditor from producing conclusions stronger than its evidence. They should be declared before a measured evaluation and reported whenever they occur.
A failure condition does not always mean that the evaluated system failed. It may mean that the audit design, evidence boundary, instrumentation, comparison, or reporting process cannot support the proposed conclusion.
Integrity rule: when a material failure condition is triggered, preserve the result, label the limitation, and reduce the scope of the conclusion.
System and Scope Failures
The evaluated system, model, version, configuration, or workload cannot be identified.
Material components change during testing without a new audit version.
The conclusion extends beyond the system, task, domain, or observation period tested.
A product feature is generalized to an entire organization or technology category.
The audit combines results from materially different configurations as if they were one system.
Baseline and Quality Failures
The baseline is selected or changed after comparative results are known.
The compared systems receive different tasks, tools, constraints, or quality thresholds.
Failed, rejected, or abandoned runs are excluded.
The evaluator or acceptance rubric changes without disclosure.
Apparent efficiency is created by reducing correctness, completeness, safety, or usefulness.
Evidence and Measurement Failures
Material evidence lacks source identity, date, version, or provenance.
Calculated results omit inputs, units, assumptions, or uncertainty.
Inferred findings are presented as directly measured or documented facts.
Missing telemetry is replaced by unsupported estimates.
The measurement instrument, evaluator, or data pipeline cannot be reproduced or validated.
Private internal reasoning is claimed without observable evidence or authorized access.
CEMR Failures
Dimension
Material Failure Condition
Required Response
Compression
Reduced representation loses structure required for accepted performance.
Reject the compression benefit or narrow the task-sufficiency claim.
Expression
Available structure cannot produce a correct or usable result.
Record the failed expression and associated correction burden.
Memory
Stored structure cannot be retrieved with identity, provenance, constraints, and correct version state.
Do not characterize the stored object as reliably reusable memory.
Recursion
Errors, ambiguity, or cost compound across successor-state transitions.
Stop the sequence, preserve the failed state, and test containment.
Total Cost
Excluded overhead reverses or materially changes the reported benefit.
Recalculate using the expanded boundary or report the result as Unknown.
Governance and Reporting Failures
A diagnostic tool output is represented as certification or formal validation.
Licensing, payment, publication, adoption, or implementation is presented as scientific evidence.
A formal evidence state is assigned without a registered claim and qualifying evaluation.
A benchmark result is reported without its failures, limitations, or configuration details.
Contradictory evidence is omitted from the final report.
A prior result is silently overwritten rather than versioned, challenged, superseded, or retired.
Environmental-Claim Failures
Token reduction is converted directly into energy, water, cooling, or emissions savings.
Provider pricing is used as a substitute for physical infrastructure telemetry.
Hardware, utilization, batching, geography, energy supply, and cooling conditions are omitted.
A workload-level result is projected to global infrastructure without a defensible accounting model.
An environmental conclusion is reported despite material telemetry remaining Unknown.
Required Response to a Failure Condition
Stop or isolate the affected portion of the evaluation.
Preserve the failed run, evidence, logs, and configuration.
Identify which conclusion is affected.
Change the finding label or formal evidence state when required.
Reduce the scope of the conclusion.
Register a corrective test or independent replication pathway.
A clean failure is useful evidence. Preserving negative, challenged, and inconclusive results strengthens the evaluation system by showing where a proposed mechanism does not hold.
Benchmark Escalation: Move From Diagnostic Questions to Testable Evidence
A diagnostic audit organizes the evidence and identifies what must be tested. Benchmark escalation begins when a material proposed finding can be converted into a preregistered comparison with a baseline, metrics, thresholds, and failure conditions.
Not every diagnostic observation requires a formal benchmark. Escalation is most valuable when a claim affects architecture, procurement, deployment, licensing, environmental reporting, safety, or a broader theoretical conclusion.
Diagnostic Question → Registered Claim → Baseline → Controlled Test → Result Record → Replication → Evidence-State Decision
When to Escalate
Material Efficiency Claim
A proposed finding claims reduced total cost, latency, computation, reconstruction, or human burden.
Reusable Memory Claim
A stored representation is expected to support later accepted tasks or reduce reconstruction.
Recursive Stability Claim
A workflow is expected to preserve quality and constraints across repeated state transitions.
Domain-Transfer Claim
A result from one workload or domain is proposed to apply to another.
Environmental Claim
A measured system delta is proposed to affect energy, cooling, water, or emissions.
Operational Decision
The finding may influence deployment, procurement, governance, investment, or licensing.
Minimum Benchmark Registration
Registration Field
Required Declaration
Claim
A narrow, falsifiable prediction tied to an identified mechanism
System Boundary
System, version, configuration, tools, infrastructure, and workload
Baseline
The comparison condition and why it is appropriate
Metrics
Quality, CEMR, total-cost, and applicable infrastructure measures with units
Thresholds
The result required to support, challenge, or leave the claim inconclusive
Failure Conditions
Events that invalidate the test or contradict the proposed mechanism
Analysis Plan
Sample, scoring, exclusions, uncertainty, and statistical or operational method
Result Record
All accepted, failed, rejected, and missing observations with versioned outputs
Escalation Levels
Level
Purpose
Permitted Conclusion
1. Diagnostic
Identify mechanisms, evidence gaps, and testable questions
Finding labels and proposed tests only
2. Pilot Benchmark
Test feasibility, instrumentation, and preliminary thresholds
Scoped preliminary result with strong limitations
3. Controlled Evaluation
Run the preregistered comparison under equivalent conditions
Evidence-state decision for the tested claim
4. Independent Replication
Test whether the result persists under independent implementation
Increased confidence within replicated boundaries
5. Domain Transfer
Test applicability under a different workload, domain, or environment
Transfer finding limited to the newly tested conditions
Domain-Transfer Constraint
Success in one system, workload, scale, or domain does not establish success in another. Each transfer requires its own declared boundary, baseline, quality gate, measurements, and failure conditions.
Continue to the Benchmark System
Use the public benchmark hub, lab protocol, and GitHub repository to register tests, preserve result records, and support reproducible evaluation.
The Razor Auditor is an applied diagnostic and evaluation-intake framework. It does not replace the governing canon, claims register, lab protocol, benchmark records, or independent replication.
Use the following resources to move between canonical definitions, testable claims, evaluation procedures, reproducible benchmark assets, and applied examples.
Authority boundary: definitions and governing claims flow from GC-MRD-v2.0 and its canonical claims register. Auditor outputs remain scoped findings and do not amend the canon.
Applied pages illustrate questions and evaluation pathways. They do not establish findings beyond their registered evidence.
Public Benchmark Repository
The Robbie’s Razor Benchmarks repository provides the public reproducibility layer for benchmark specifications, result structures, evidence records, and implementation resources.
These answers clarify what the Razor Auditor does, what evidence it requires, and where its conclusions must stop.
What is the Razor Auditor?
The Razor Auditor is a structured diagnostic and evaluation-intake framework for examining an AI system through compression, expression, memory, recursion, quality, evidence confidence, infrastructure, and total cost.
Does the Razor Auditor certify an AI company or product?
No. It produces scoped findings about an identified system, version, configuration, workload, and evidence boundary. It does not assign a universal company-wide classification or certification.
What does the Auditor measure?
Depending on the available evidence, it can evaluate accepted-task quality, representational burden, output success, memory retrieval, recursive reuse, correction burden, latency, measured resource use, and total cost per accepted result.
Do fewer tokens prove better compression or greater efficiency?
No. Fewer tokens may reflect useful compression, omitted requirements, different formatting, or displaced work. The result must pass the same quality gate, preserve task-relevant structure, and show a favorable total-cost comparison.
Can the Auditor calculate energy, water, or emissions savings from token counts?
No. Token counts alone do not establish physical environmental impact. Energy, water, cooling, and emissions conclusions require direct telemetry or a documented accounting method that includes hardware, utilization, geography, energy supply, cooling design, and other material conditions.
What is the difference between a finding label and an evidence state?
A finding label—Documented, Calculated, Inferred, Proposed, or Unknown—describes the basis of an audit statement. An evidence state—Proposed, Testing, Provisionally Supported, Supported, Challenged, Inconclusive, or Retired—describes the status of a registered claim after formal evaluation.
Does the Auditor inspect a model’s hidden chain-of-thought?
No. The framework evaluates observable inputs, outputs, tools, stored state, retrieval behavior, telemetry, and accepted outcomes. Private chain-of-thought is neither required nor assumed.
What does Unknown mean in an audit report?
Unknown means that the available evidence is insufficient for a defensible conclusion. It identifies a measurement or documentation gap; it is not automatically a negative score.
When can a finding receive a formal evidence state?
A formal evidence state requires a registered claim with a prediction, baseline, metrics, thresholds, failure conditions, analysis plan, and preserved result record. Independent replication may be required before a claim is considered Supported.
Does building or deploying a Razor-based tool validate Robbie’s Razor?
No. Implementation demonstrates that a tool or workflow was built. Validation requires a separate controlled evaluation against registered predictions, baselines, metrics, thresholds, and failure conditions.
Can one successful audit result be transferred to another domain?
Not automatically. Each new domain, workload, scale, and operating environment requires its own boundary, quality gate, baseline, measurements, and failure conditions.
How do I begin an audit?
Start by defining the exact system, version, configuration, workload, baseline, accepted outcome, evidence boundary, and audit mode. Use the interactive auditor for exploratory diagnostics and the lab evaluation protocol for formal measured testing.
Robbie George is a National Geographic–published nature photographer, writer, and creator of the Grand Compression Cosmology and Robbie’s Razor.
His work examines how systems preserve usable structure across change through compression, expression, memory, and recursion. The Razor Auditor extends that framework into a disciplined method for asking measurable questions about AI reasoning systems, reusable memory, recursive workflows, infrastructure, and total cost.
Robbie’s field experience in nature photography and ecological observation informs his interest in systems that operate under real limits. His current work connects visible human-readable explanations with versioned claims, machine-readable resources, public benchmark assets, and reproducible evaluation pathways.
The Grand Compression framework remains evidence-bounded: implementation does not equal validation, domain transfer requires new testing, and formal claims must remain open to support, challenge, revision, or retirement.
Robbie George is the originator of Robbie’s Razor and the Razor Auditor framework. Author-created examples, tools, and implementations should be distinguished from independent evaluation. Formal support requires registered testing, preserved results, transparent limitations, and appropriate replication.
The presence of this badge signifies that this business has officially registered with the Art Storefronts Organization and has an established track record of selling art.
It also means that buyers can trust that they are buying from a legitimate business. Art sellers that conduct fraudulent activity or that receive numerous complaints from buyers will have this badge revoked. If you would like to file a complaint about this seller, please do so here.
Verified Returns & Exchanges
The Art Storefronts Organization has verified that this business has provided a returns & exchanges policy for all art purchases.
Description of Policy from Merchant:
What is your Policy on Returns/Exchanges/Refunds?
I take great pride in my work and prints, and I want you to be completely happy with your investment in my nature art. If for any reason you are unsatisfied with your print, you may return it within 14 days of delivery, and/or exchange it for another print. Prints must be returned in new condition, packaged carefully in the original packaging if possible. Your refund will be issued as soon as I receive the returned print. Please contact me if you would like to arrange a return or exchange.
In the event that you receive a damaged or defective print, please let me know within 7 days of receipt, and I will arrange for a new print to be shipped to you at no additional cost.
Verified Secure Website with Safe Checkout
This website provides a secure checkout with SSL encryption.
Verified Archival Materials Used
The Art Storefronts Organization has verified that this Art Seller has published information about the archival materials used to create their products in an effort to provide transparency to buyers.
Description from Merchant:
Fine Art Prints are made with high-quality archival inks on fine art papers using a high-resolution large format inkjet printer. Our premium archival inks produce images with smooth tones and rich colors. Prints are made with care on your choice of exquisite Fine Art Papers using a high-resolution large format inkjet printer. https://www.graphikprintworks.com
Become a supporter of Robbie George Photography and be the first to receive new content and special promotions.
“Every image is a field. Every quote is a key. Welcome back to the rhythm.” ~Robbie
Cart
Your cart is currently empty.
Saved Successfully.
This is only visible to you because you are logged in and are authorized to manage this website. This message is not visible to other website visitors.
Import From Instagram
Click on any Image to continue
This Website Supports Augmented Reality to Live Preview Art
This means you can use the camera on your phone or tablet and superimpose any piece of nature art onto a wall inside of your home or business.
To use this feature, Just look for the "Live Preview AR" button when viewing any piece of nature art on this website!
Pounce Now—Save 20% on Your First Order
Join the collector list for your first-order discount, new wildlife releases, and occasional field notes.