Proposes a four-stage forensic audit protocol that enables accurate black-box identity verification of anonymously released frontier models served via APIs.
The search context does not contain information about a specific "Four-Stage Protocol for Black-Box Identity Verification." Instead, it highlights KBF (Knowledge Boundary Fingerprinting), a three-stage protocol designed for black-box auditing of language model APIs.
The paper addresses a practical gap in AI governance: how to verify the identity of a frontier model when only an API is available and the release is anonymous. It proposes a four-stage forensic audit protocol that treats model identity as an empirically testable property of black-box behavior rather than a claim that must be taken on faith. The protocol begins by specifying the candidate model and the identity hypothesis, then collects controlled probe responses across a range of tasks, prompt styles, and constraints, followed by statistical comparison against reference behavioral fingerprints, and concludes with a calibrated audit decision that reports confidence, uncertainty, and the evidence supporting the verdict.
A key contribution is a structured methodology for turning API-only access into audit evidence. The work emphasizes discriminative signals that are informative but difficult to fully suppress, such as task-specific competencies, stylistic regularities, calibration, refusal or safety behavior, and robustness under perturbations. Its central insight is that frontier models can leave sufficiently distinctive behavioral signatures for practical identification, but such verification must be probabilistic and confound-aware, accounting for factors such as fine-tuning, quantization, system prompts, decoding settings, and adversarial obfuscation. This frames model identity verification as a forensic measurement problem rather than a simple model-card or benchmark comparison.
More broadly, the work matters because API distribution and anonymous releases complicate provenance, licensing, safety accountability, and supply-chain risk. A standardized audit protocol could help regulators, platforms, and third-party auditors detect misrepresentation, unauthorized reuse, or substitution of a claimed frontier model, and it provides a path toward more transparent governance without requiring weight access. At the same time, the approach is not a cryptographic proof of identity; its value lies in producing defensible, evidence-based judgments that can inform oversight, market trust, and model-release policy.