arXiv:2604.00151v2 Announce Type: replace Abstract: Distributed applications use identifiers across database storage, trusted communication, and external access. Depending on their use, these roles require different combinations of storage efficiency, chronological sortability, origin metadata embedding, zero-lookup verifiability, metadata confidentiality, and multi-century addressability. The id

Topological visualization of Source-Known Identifiers: A Three-Tier Identity System for Distributed Applications
Brave API

Source-Known Identifiers (SKIDs) are a three-tier identity system designed to unify storage efficiency, chronological sortability, origin metadata, zero-lookup verifiability, confidentiality, and multi-century addressability in distributed applications. Proposed by Duran Serkan Kılıç in arXiv:2604.00151v2 (September 2026), the system projects a single entity identity across three trust boundaries using deterministic transformations.

The three tiers provide specialized representations for different environments: Database Tier (SKID): A compact 64-bit primary key embedding a 250-millisecond timestamp and application topology, enabling efficient B-tree indexing and natural ordering. Trusted Internal Tier (SKEID): A 128-bit UUID-compatible identifier that adds entity type, epoch, and a BLAKE3 MAC for integrity and origin verification without database lookups. * External Tier (Secure SKEID): An AES-256 encrypted version of the SKEID that conceals metadata from untrusted consumers while maintaining standard UUID string format compatibility.

This architecture eliminates the need for dual-identifier schemas by allowing direct conversion between tiers, reducing storage overhead by up to 75% compared to conventional systems using separate primary keys and external UUIDs.

Generated 10d ago
Open-Weights Reasoning

The arXiv paper addresses a persistent but under-theorized design problem in distributed systems: identifiers are not a single primitive, but are used in several distinct operational contexts with different requirements. The authors distinguish at least three roles—persistent database storage, trusted or authenticated communication, and external access—and argue that each role imposes a different balance of constraints. Storage identifiers benefit from compactness and chronological sortability; communication identifiers may need embedded origin or provenance information and verifiable trust relationships; external identifiers require long-lived, stable addressability while minimizing leakage of internal metadata. The central premise is that these concerns are often conflated in practice, causing systems to force one identifier format to serve roles for which it is poorly suited.

The paper’s main contribution is Source-Known Identifiers, a three-tier identity system that separates these concerns into distinct identifier layers. The approach is built around the idea that an identifier can be “source-known” in the sense that it carries, or can reveal, enough provenance information to support origin-aware processing without exposing all metadata to every consumer. The authors evaluate identifier designs along dimensions such as storage efficiency, chronological sortability, origin metadata embedding, zero-lookup verifiability, metadata confidentiality, and multi-century addressability. In effect, the work reframes identifier design as a multi-objective architecture problem: rather than seeking a single universal ID format, the system should match identifier semantics to the trust, performance, and longevity requirements of each boundary in the application.

This matters because identifier design has outsized consequences for system performance, security, auditability, and long-term maintainability. Poorly chosen identifiers can lead to index bloat, cache inefficiency, weak provenance tracking, privacy leaks, or broken external links over decades. By providing a principled taxonomy and a three-tier model, the paper offers a reusable design framework for databases, event streams, APIs, microservices, and other distributed systems that must distinguish internal identity from authenticated identity and public addressability. Its key insight is that robust distributed identity is less about finding one perfect identifier and more about aligning identifier semantics with the specific trust and operational context in which it is used.

Generated 10d ago
Sources