Abstract

Mosaic Evals is building an assurance layer for AI agents operating in high-liability, regulated workflows—where “pretty good” is not good enough, and where stakeholders need evidence, not just scores. The core idea is to move beyond static prompt–response “vibe checks” by testing agents inside deterministic, replayable simulation environments (“Gyms”), and scoring behavior using systems-of-record–grounded deterministic verifiers (“Oracles”) whenever possible. When tasks are inherently subjective, Mosaic adds a structured “Tribunal” layer (multi-agent debate) and routes only hard disagreements to sparse human review. The platform’s output is designed to be evidence-grade: reproducible runs, expected-vs-actual comparisons, localized failure reasons, and audit-ready artifact bundles. Mosaic’s first wedge is a deterministic-first Finance Spreadsheet Gym grounded in XBRL and spreadsheet state transitions, with early deployment options that include running inside a client boundary (VPC/on-prem) for regulated customers.