Skip to main content
Technical Documentation

Credenva Assessment Methodology

The authoritative technical reference for how Credenva evaluates AI workers. Version 0.3.0 — pilot methodology.

Status: PilotLast updated: 2026-06-10Applies to: Customer Support role
1.0

Mission

Credenva exists to give enterprise buyers, procurement teams and governance functions independent, evidence-based answers to a single question: can this AI worker be trusted in production?

We evaluate complete AI workers — the foundation model together with its retrieval layer, policies, business rules, tools, memory and escalation paths — against standardized scenarios that reflect real enterprise workflows. Our outputs are designed to be citable in procurement decisions, board reviews and third-party risk assessments.

2.0

Evaluation Philosophy

Model benchmarks evaluate models. Developer evaluation platforms evaluate applications during development. Neither answers the executive question of whether a specific AI worker behaves reliably under the conditions it will actually face.

  • Deployment-first. We assess the system a customer interacts with, not the model in isolation.
  • Evidence over opinion. Every score is backed by captured transcripts, tool calls and retrieved context.
  • Vendor-neutral. The same scenarios, rubrics and reviewers apply to every AI worker in a given role.
3.0

Scenario Design

Scenarios are version-controlled test cases authored with domain experts for a specific role — for example, an AI Customer Support Representative. Each scenario specifies the customer intent, the business context, the systems of record involved and the acceptable outcome space.

  • Common flows: account, billing, order status, entitlement.
  • Edge cases: ambiguous intent, conflicting records, partial data.
  • Adversarial cases: policy pressure, jailbreak attempts, PII elicitation.
4.0

Ground Truth

Every scenario ships with a documented expected outcome and a set of acceptable variants. Ground truth is authored by role-specific experts, reviewed by a second expert, and locked to a version hash before any AI worker is scored against it.

When an AI worker produces a novel-but-defensible resolution, reviewers may propose an amendment; amendments are applied only in subsequent versions and never retroactively.

5.0

Scoring

Deployments receive a 0–100 score on each dimension: Task Accuracy, Resolution Effectiveness, Policy & Compliance, Groundedness, Escalation Judgment and Reliability. The Credenva Readiness Score™ is a weighted aggregate; weights are published per role.

An AI worker is marked Production Ready only when the aggregate meets the role threshold and no critical failures are recorded.

6.0

Critical Failures

Certain outcomes are treated as disqualifying regardless of overall score. These include unauthorized disclosure of personal data, fabricated policy or entitlement claims, silent policy violations, and failure to escalate when a documented trigger is met.

A single critical failure prevents a passing verdict for the version under test. The full list is published in the role standard (see the Customer Support Standard).

7.0

Repeatability

Assessments must be reproducible. We record the AI worker version, model identifiers, retrieval index snapshots, tool schemas, prompt templates and environment configuration used during the run.

Re-running the same scenario set against the same AI worker version must yield equivalent verdicts within a documented variance band; exceeding that band invalidates the run.

8.0

Human Review

Automated grading proposes verdicts; qualified human reviewers ratify them. Every critical failure, every escalation judgment and every scenario flagged as ambiguous is reviewed by at least two reviewers, with disagreements resolved by a senior adjudicator.

Reviewer identities, qualifications and inter-rater agreement statistics are published alongside each release.

9.0

Governance

Methodology changes are proposed as versioned amendments, opened for comment, and ratified by the methodology council before taking effect. Vendors under assessment do not sit on the council for their own role.

Conflicts of interest, funding sources and reviewer affiliations are disclosed in every report.

10.0

Version History

VersionDateSummary
0.3.02026-06-10Added Escalation Judgment dimension; expanded adversarial scenario set.
0.2.02026-03-22Introduced critical-failure disqualifiers and inter-rater reporting.
0.1.02026-01-15Initial pilot methodology for the Customer Support role.