Proposes a four-stage forensic audit protocol that enables accurate black-box identity verification of anonymously released frontier models served via APIs.

Topological visualization of Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification
Brave API

The search context does not contain information about a specific "Four-Stage Protocol for Black-Box Identity Verification." Instead, it highlights KBF (Knowledge Boundary Fingerprinting), a three-stage protocol designed for black-box auditing of language model APIs.

KBF Protocol Overview KBF enables accurate identity verification of anonymously released or relayed frontier models by leveraging stable numerical recall near the knowledge boundary. The protocol operates in three stages: 1. Offline Probe Generation: The auditor uses the official reference API to construct a model-specific candidate set of queries. 2. Stability Screening: Candidates are screened for stability across different deployment configurations (e.g., system prompts, temperatures) and contrasted against cheaper substitute models to ensure discriminative power. 3. Online Audit: The auditor queries the suspect endpoint, compares the returned numerical values against the reference consensus, and applies a statistical test to determine consistency with the claimed model.

Key Capabilities * High Accuracy: KBF detects all 155 economically relevant substitutions across 16 production endpoints with zero false positives on same-model controls. * Robustness: It remains effective under deployment variations, including RAG wrappers and agent-style interfaces. * Cost-Effectiveness: Audits are low-cost, with online audits costing under $0.05 per model after an initial probe-generation cost.

Limitations of Black-Box Access While KBF demonstrates the efficacy of black-box methods, other research argues that black-box access alone is insufficient for rigorous AI audits. Studies suggest that white-box (access to weights/activations) and outside-the-box (access to development context) access provide substantially more scrutiny, particularly for identifying internal mechanisms, dormant capabilities, and sycophantic behaviors that input/output analysis alone may miss.

Generated Sep 1, 2026
Open-Weights Reasoning

The paper addresses a practical gap in AI governance: how to verify the identity of a frontier model when only an API is available and the release is anonymous. It proposes a four-stage forensic audit protocol that treats model identity as an empirically testable property of black-box behavior rather than a claim that must be taken on faith. The protocol begins by specifying the candidate model and the identity hypothesis, then collects controlled probe responses across a range of tasks, prompt styles, and constraints, followed by statistical comparison against reference behavioral fingerprints, and concludes with a calibrated audit decision that reports confidence, uncertainty, and the evidence supporting the verdict.

A key contribution is a structured methodology for turning API-only access into audit evidence. The work emphasizes discriminative signals that are informative but difficult to fully suppress, such as task-specific competencies, stylistic regularities, calibration, refusal or safety behavior, and robustness under perturbations. Its central insight is that frontier models can leave sufficiently distinctive behavioral signatures for practical identification, but such verification must be probabilistic and confound-aware, accounting for factors such as fine-tuning, quantization, system prompts, decoding settings, and adversarial obfuscation. This frames model identity verification as a forensic measurement problem rather than a simple model-card or benchmark comparison.

More broadly, the work matters because API distribution and anonymous releases complicate provenance, licensing, safety accountability, and supply-chain risk. A standardized audit protocol could help regulators, platforms, and third-party auditors detect misrepresentation, unauthorized reuse, or substitution of a claimed frontier model, and it provides a path toward more transparent governance without requiring weight access. At the same time, the approach is not a cryptographic proof of identity; its value lies in producing defensible, evidence-based judgments that can inform oversight, market trust, and model-release policy.

Generated Sep 1, 2026
Sources