Skip to main content
Methodology

Built to be repeatable, transparent and independent.

The Credenva methodology exists to answer one question with defensible evidence: is this AI worker the right hire for this specific business role?

Design principles

Eight principles behind every assessment.

01

Role-scoped assessment

Each assessment is scoped to the specific business role the AI worker will perform. Verification for one role does not imply readiness for another.

02

Standardized scenarios

Every AI worker faces the same version-controlled scenarios for the role, drawn from real business workflows and reviewed by domain experts.

03

Equivalent conditions

Identical inputs, policies, tools and constraints so results across AI workers are directly comparable.

04

Observable evidence

Scoring is grounded in what the AI worker actually did — prompts, tool calls, actions, outcomes and timing — never marketing narratives.

05

Repeated runs

Scenarios are run multiple times to measure variability, not just a single lucky pass.

06

Critical-failure gates

Designated high-risk categories (policy, safety, privacy) have absolute thresholds that override aggregate scores.

07

Human review sampling

A sampled percentage of every assessment is reviewed by trained human evaluators to validate automated scoring.

08

Independent & vendor-neutral

Credenva does not sell AI workers, models, prompts, integrations or infrastructure — and does not favor any provider or framework.

Scoring model

Three layers of scrutiny per scenario.

  1. LAYER 01

    Deterministic checks

    Rule-based verification of tool calls, permissions, policy adherence and required actions.

  2. LAYER 02

    Structured evaluation

    Rubric-driven scoring of accuracy, resolution, groundedness and escalation judgment across every dimension.

  3. LAYER 03

    Human review

    A trained reviewer independently inspects a sampled percentage of runs to validate automated scores.

What we evaluate

The deployment is the system. Not the model.

Credenva evaluates the complete AI worker — the model, integrations, policies, tools, and workflows that customers and employees actually interact with.

Input
Customer
System under evaluation
Enterprise AI Deployment
Foundation Model
Knowledge Base
Policies
Business Rules
Memory
CRM
External APIs
Authentication
Human Escalation
Output
Credenva Assessment
Why this matters

A foundation model may perform well on a general model benchmark, but a deployment can still fail in production because of how it is integrated, governed, and operated. Credenva evaluates the complete deployment rather than the model in isolation, so the assessment reflects the system a CIO or procurement team would actually rely on.