Customer Support Pilot Standard
Draft v0.1Standardized evaluation of AI AI customer-support workers.
The first Credenva assessment program evaluates AI systems performing customer-support work through chat or messaging channels. Deployments are exercised across common, complex, adversarial and high-risk workflows, and scored on evidence rather than sample outputs.
Evaluated workflows
Ten workflow families across routine and high-risk scenarios.
- 01Product and policy questions
- 02Order status and delivery issues
- 03Subscription changes
- 04Billing and duplicate charges
- 05Returns and refunds
- 06Account access
- 07Fraud indicators
- 08Angry or distressed customers
- 09Privacy-sensitive requests
- 10Human escalation
Methodology principles
Repeatable, transparent, vendor-neutral.
- Same scenario conditions for comparable AI workers
- Version-controlled test cases and policies
- No preference for any model provider or framework
- Scoring based on observable outputs and actions
- Critical-failure gates for designated high-risk scenarios
- Repeated runs to measure variability
- Clear disclosure of limitations
- Human review where automated scoring is insufficient
- Public change history for released methodology versions
Qualifying pilot participants
The pilot is designed for vendors and companies with a working AI AI customer-support worker that can be accessed through an API, controlled test environment or supervised demonstration.
Apply to Participate