ZTA Labs

Model Assurance

Make model selection evidence-led.

ZTA independently benchmarks open-weight models against the institution’s actual Investment Research questions, documents, data, and evaluation standards, then maintains the evidence and governance framework required to determine which models remain fit for which investment tasks as models and institutional requirements change.

Why it matters

Model assurance is a program, not a one-time report.

The initial baseline is only the beginning. Models change, investment mandates and use cases evolve, and governance obligations grow more demanding as AI moves deeper into production work for asset managers and asset owners.

Prove model performance before production. Keep proving it after deployment.

Baseline model qualification

One corpus, a defined model pool, firm-specific scoring, and an initial recommendation set.

Monthly assurance briefs

Developments in the model pool, the benchmark, or operating landscape that matter to the client.

Quarterly scorecards

Evidence-backed recommendations on which models remain approved, restricted, or removed.

Annual governance audit

Review whether the methodology, controls, and audit record remain fit for purpose.

Event-triggered re-runs

Out-of-cycle triage when a material release, new corpus, or new use case changes the decision surface.

What you get

A model shortlist with reasons

Not a ranking. Each candidate is assessed on extraction, analysis, and judgment separately, because a model that reads figures flawlessly may still be unfit for the work that matters.

A benchmark that stays yours

The corpus, answer key, scoring rules, and evaluation record remain in the client environment, ready to be re-run against future open-weight model candidates without creating a proprietary ZTA dependency.

An audit trail built to be examined

Every score traces back to evidence. An investment, risk, governance, or regulatory reviewer should be able to inspect the record rather than rely on a recollection.

Failure mode exampleCompetent work can still be wrongBasis error
Problem

The wrong basis, correctly cited. A model is asked for GAAP gross margin and returns the non-GAAP figure found a few lines lower in the same release.

Why it matters

The number is in the source, the citation resolves, and nothing reads as obviously incorrect. The error propagates quietly into every multiple built on it.

Why validation matters

An unvalidated model does not announce these errors. It produces work that reads well and is wrong in ways that compound quietly across a coverage universe.

Open model governance

Benchmark the model. Keep the benchmark.

Model Assurance is designed around open-weight models and a client-owned evaluation record. The institution can re-run the benchmark as models change, move the methodology to another operating partner, and preserve the evidence behind every admission decision.

Open architecture. Open-weight models. Institution-owned evidence.

Sample baseline

Sample Model Assurance Baseline

A public example of the baseline work product: model-level findings, failure analysis, auditability, governance requirements, and deployment recommendations.

100Questions
7Models evaluated
700Scored responses
3Models admitted to Model Council

Independence

The evaluator must be independent of the judged.

The model that judges the work must not be the model that produced it, and should not share its training lineage. Correlated failure modes mean a model can mark its own blind spots as correct.

We also treat the evaluator itself as something to be verified rather than assumed. Assurance that has never found a fault in its own process has not looked hard enough.