The Trust Score
The Trust Score is a composite metric ranging from 0 to 1 that quantifies how much you can trust your agent in production. It aggregates performance across all evaluated dimensions.
The threshold of 0.70 represents a baseline for acceptable behavior. Agents scoring below this threshold exhibited failure modes that pose unacceptable risk.
Taxonomy of Trust
Vijil organizes agent behavior into a three-level taxonomy:Reliability
- Correctness
- Consistency
- Robustness
Security
- Confidentiality
- Integrity
- Availability
Safety
- Containment
- Compliance
- Transparency
Reliability
Reliability measures whether your agent produces correct, consistent, and robust outputs.Security
Security measures whether your agent resists attacks on confidentiality, integrity, and availability.Safety
Safety measures whether your agent operates within acceptable boundaries.Reading the Trust Report
The Trust Report organizes findings by taxonomy level:Dimension Breakdown
Each dimension (Reliability, Security, Safety) shows:- Dimension score: Aggregate performance for this pillar
- Subcategory scores: Performance for each aspect within the dimension
- Finding count: Number of issues identified at each severity level
Individual Findings
Each finding includes:- Category: Where in the taxonomy this issue falls
- Severity: Risk level from 1–4
- Probe: The test case that revealed this behavior
- Agent Response: What your agent actually produced
- Expected Behavior: What a trustworthy agent would produce
- Recommendation: Specific remediation guidance
Understanding Red Team Results
Red Team results are campaign evidence, not a Trust Score. Open a Red Team result from Tests → Evaluation Results by selecting an evaluation with type Red Team. The result page has three main areas:- Run summary: Current status, phase, cost, progress, and wave information
- Waves: Per-wave seeds, attackers, transcripts, strategies, and judgments
- Final Report: Aggregated findings across the full campaign
Run Summary
The summary at the top of the result page tells you where the campaign is in its lifecycle.
Use this section to understand progress and cost. Use the Waves and Final Report sections to interpret findings.
Waves
A Red Team campaign runs in waves. The Waves section lists every wave in the evaluation and shows basic information for each one. Click a wave to inspect what Red Team generated and what attackers did during that wave. Each wave detail view has two main sections:
Review wave details to see whether later waves became more focused based on earlier discoveries. Strong campaigns often show broad exploration early, then more targeted attacks after reflections identify promising strategies.
Live Attacker Details
Click a Live Attacker to inspect the evidence behind one seed. The detail view is organized into several sections:Judgments and Harm
After an attack completes, a judge reviews the transcript. In the UI, focus on Harmful Content Delivered and Potential Harm. Harmful Content Delivered shows whether the Agent fulfilled the attack goal or leaked a verified sensitive artifact:
Potential Harm helps separate verified harm from findings that need owner review. Treat a verified policy violation or verified leaked artifact as real harm that needs remediation. Treat potential or unverified harm as something a human owner should check against the actual Agent design, policies, and data access.
Leaked artifacts are internal details the Agent disclosed, such as system prompt fragments, tool names, private endpoints, credentials, or operational procedures.
Final Red Team Report
The Final Report section shows a short summary of the evaluation. Click Open full report view to inspect the details used to create that summary. The full report view includes:
The report summary is aggregated from all waves, seeds, transcripts, and judgments. Use it to decide which issues need product changes, prompt or policy updates, tool permission changes, or Dome Guardrails.
Prioritizing Remediation
Use severity and taxonomy to prioritize fixes: Address immediately (Critical/High severity):- Security vulnerabilities (prompt injection, data leakage)
- Safety violations (harmful content, scope violations)
- Reliability failures that affect core functionality
- Red Team findings with FULL harmful-content judgments
- Consistency issues across sessions
- Minor compliance gaps
- Robustness failures on edge cases
- Red Team findings with PARTIAL harmful-content judgments
- Transparency improvements
- Minor formatting inconsistencies
- Rare edge case handling
Comparing Evaluations
Run evaluations before and after changes to track improvement:
A rising Trust Score with decreasing critical findings indicates effective remediation. A declining score signals regression—investigate recent changes.
Next Steps
Configure Guardrails
Add runtime protection with Dome
Quantifying Risk
Translate findings into risk assessments
Trust Score Harness
Learn about the standard evaluation
Run Evaluations
Launch and monitor evaluations