Evaluators Reference
This section provides a comprehensive reference for all available evaluators in ARES. Evaluators analyze model responses to determine whether they meet safety, robustness, or refusal criteria.
Overview
ARES supports multiple types of evaluators:
Keyword-based: Pattern matching against refusal keywords
LLM-based: Model-as-a-judge evaluation using language models
Detector-based: Specialized detectors (e.g., Garak encoding detectors)
OWASP-specific: Evaluators tailored for specific OWASP vulnerabilities
Keyword Evaluator
Type: ares.evals.keyword_eval.KeywordEval
Evaluates responses by checking for the presence of refusal keywords or patterns.
Configuration:
evaluation:
keyword:
type: ares.evals.keyword_eval.KeywordEval
name: keyword
keyword_list_or_path: assets/advbench_refusal_keywords.json
output_path: results/keyword_evaluation.json
Use Cases: - Quick refusal detection - Low-resource evaluation - Baseline safety checks
LLM-Based Evaluators
HarmBench Evaluator
Type: ares.evals.llm_eval.LLMEval
Uses a language model to judge whether responses are harmful based on HarmBench criteria.
Configuration:
evaluation:
harmbench_eval:
type: ares.evals.llm_eval.LLMEval
name: harmbench_eval
output_path: results/harmbench_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/harmbench_prompt.yaml
Features: - Model-as-a-judge approach - Nuanced safety assessment - Configurable evaluation prompts
HuggingFace Evaluator
Type: ares.evals.huggingface_eval.HuggingFaceEval
Similar to LLM evaluator but optimized for HuggingFace models.
Configuration:
evaluation:
huggingface_eval:
type: ares.evals.huggingface_eval.HuggingFaceEval
name: harmbench_hf_eval
output_path: results/harmbench_hf_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/harmbench_prompt.yaml
GPT Evaluator
Type: ares.evals.llm_eval.LLMEval
Uses GPT-OSS-120B on WatsonX for evaluation.
Configuration:
evaluation:
gpt_eval:
type: ares.evals.llm_eval.LLMEval
name: gpt_eval
output_path: results/gpt_evaluation.json
connector:
litellm_connector:
name: gpt-oss-120b-on-watsonx
endpoint-type: watsonx
model: openai/gpt-oss-120b
reasoning_effort: low
prompt_path: assets/gptoss_prompt.yaml
Garak Detectors
Garak detectors are specialized evaluators from the NVIDIA Garak framework for detecting encoded attacks.
DecodeMatch Detector
Type: ares_garak.evals.encoding.Detector
Performs exact match detection on decoded content.
Configuration:
evaluation:
garak_decoding_match:
type: ares_garak.evals.encoding.Detector
name: decode_match
detector: detectors.encoding.DecodeMatch
output_path: results/garak_detector_match_evaluation.json
DecodeApprox Detector
Type: ares_garak.evals.encoding.Detector
Performs approximate match detection on decoded content.
Configuration:
evaluation:
garak_decoding_approx:
type: ares_garak.evals.encoding.Detector
name: decode_approx
detector: detectors.encoding.DecodeApprox
output_path: results/garak_detector_approx_evaluation.json
Use Cases: - Evaluating encoding attack success - Detecting obfuscated harmful content - Testing decoder robustness
Crescendo Evaluator
Type: ares_pyrit.evals.crescendo_eval.CrescendoNativeEval
Specialized evaluator for Crescendo multi-turn attacks.
Configuration:
evaluation:
crescendo_eval:
type: ares_pyrit.evals.crescendo_eval.CrescendoNativeEval
name: crescendo_eval
output_path: results/crescendo_evaluation.json
connector:
litellm_connector:
name: gpt-oss-120b-on-watsonx
endpoint-type: watsonx
model: openai/gpt-oss-120b
reasoning_effort: low
prompt_path: assets/gptoss_prompt.yaml
Features: - Multi-turn conversation analysis - Progressive jailbreak detection - Context-aware evaluation
Multi-Turn Evaluator
Type: ares.evals.multi_turn_eval.MultiTurnEval
Unified evaluator for any strategy built on the MultiTurn base class. Reads the attack_successful and conversation_id fields that every multi-turn strategy writes and aggregates results at three granularities.
Evaluation Modes:
Mode |
Behaviour |
|---|---|
|
Evaluates each turn independently (1.0 = succeeded, 0.0 = failed, -1.0 = error) |
|
Collapses all turns of a conversation into one entry; success if any turn succeeded |
|
Collapses all conversations for a goal into one entry; success if any conversation succeeded |
Configuration:
evaluation:
multi_turn_eval:
type: ares.evals.multi_turn_eval.MultiTurnEval
name: multi_turn_eval
input_path: results/crescendo_attacks.json
output_path: results/multi_turn_evaluation.json
mode: conversation # prompt | conversation | goal
Use Cases:
- Evaluating Crescendo, Echo Chamber, TAP or any MultiTurn-based strategy
- Reporting success at conversation or goal granularity for dashboards
AgentBreaker Evaluator
Type: ares_garak.evals.agent_breaker.AgentBreakerEval
Scores results produced by AgentBreakerStrategy. The probe internally runs a model-as-judge, so this evaluator reads pre-computed verdicts from notes["verified_results"] and maps them to ARES prediction scores.
Scoring:
1.0— attack succeeded (tool was exploited, confidence >=success_threshold)0.0— attack failed-1.0— no response or connector error
Configuration:
evaluation:
agent_breaker_eval:
type: ares_garak.evals.agent_breaker.AgentBreakerEval
name: agent_breaker_eval
output_path: results/agent_breaker_evaluation.json
success_threshold: 0.7 # minimum confidence to count as success
Privacy Evaluator (READI)
Type: ares_privacy_evaluator.READIEvaluator
Plugin: ares-privacy-evaluator
Detects PII or PHI leakage in target responses using the READI analyser. Returns 0.0 when sensitive entities are found (attack succeeded — model leaked data) and 1.0 otherwise.
Configuration:
evaluation:
privacy_eval:
type: ares_privacy_evaluator.READIEvaluator
name: privacy_eval
input_path: assets/attack_results.json
output_path: results/privacy_evaluation.json
sensitive_type: PII # PII | PHI
exclude_prompt: true # exclude entities already present in the prompt
debug_mode: false
exclude_patterns: [] # list of regex patterns to ignore
Use Cases: - Measuring PII/PHI leakage from RAG or document-grounded systems - Complement to LLM-based evaluators when structured entity detection is required
Intrinsic Evaluator
Type: ares_intrinsics.evals.intrinsics.IntrinsicEval
Plugin: ares-intrinsics
Extends LLMEval with LoRA/aLoRA adapters from ibm-granite/granite-3.3-8b-security-lib for efficient single-forward-pass evaluation. Supports RAG data leakage and PII detection intrinsics; multiple vulnerability types can be checked without separate model calls.
Built-in intrinsics:
granite-3.3-8b-instruct-lora-rag-data-leakage— RAG context leakage detectiongranite-3.3-8b-instruct-lora-pii-detector— PII detection
Configuration:
evaluation:
intrinsic_eval:
type: ares_intrinsics.evals.intrinsics.IntrinsicEval
name: intrinsic_eval
output_path: results/intrinsic_evaluation.json
intrinsic: granite-3.3-8b-instruct-lora-pii-detector
connector:
hf_eval_model:
type: ares.connectors.huggingface.HuggingFaceConnector
name: granite-3.3-8b-instruct
model_id: ibm-granite/granite-3.3-8b-instruct
Langfuse Evaluator
Type: ares_langfuse_eval.evals.langfuse_eval.LangfuseEval
Plugin: ares-langfuse-eval
Retrieves LLM traces from a Langfuse instance via REST API and evaluates attack effectiveness by delegating response analysis to a nested ARES evaluator. Useful when the target is instrumented with Langfuse and responses are not captured directly by ARES.
Configuration:
evaluation:
langfuse_eval:
type: ares_langfuse_eval.evals.langfuse_eval.LangfuseEval
name: langfuse_evaluator
output_path: results/langfuse_evaluation.json
langfuse_config:
base_url: "https://cloud.langfuse.com"
public_key_var: LANGFUSE_PUBLIC_KEY
secret_key_var: LANGFUSE_SECRET_KEY
input_config:
traceid_path: "$.trace_id"
limit: 100
evaluator:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/inner_eval.json
connector:
litellm_connector:
type: ares_litellm.LiteLLMConnector
name: judge
endpoint-type: ollama
model: ollama/llama3
Use Cases: - Evaluating attacks against externally-hosted or agentic targets instrumented with Langfuse - Post-hoc analysis of red-team campaigns using existing trace data
OWASP-Specific Evaluators
ARES provides specialized evaluators for each OWASP LLM vulnerability category.
LLM01 Evaluator (Prompt Injection)
Type: ares.evals.huggingface_eval.HuggingFaceEval
Configuration:
evaluation:
llm01_eval:
type: ares.evals.huggingface_eval.HuggingFaceEval
name: llm01_harmbench_eval
output_path: results/owasp_llm_01_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/harmbench_prompt.yaml
LLM02 Evaluator (Sensitive Information Disclosure)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm02_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_02_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-02-2025.yaml
LLM04 Evaluator (Data and Model Poisoning)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm04_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_04_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-04-2025.yaml
LLM05 Evaluator (Improper Output Handling)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm05_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_05_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-05-2025.yaml
LLM06 Evaluator (Excessive Agency)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm06_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_06_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-06-2025.yaml
LLM07 Evaluator (System Prompt Leakage)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm07_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_07_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-07-2025.yaml
LLM09 Evaluator (Misinformation)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm09_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_09_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-09-2025.yaml
LLM10 Evaluator (Unbounded Consumption)
Type: ares.evals.llm_eval.LLMEval
Configuration:
evaluation:
llm10_eval:
type: ares.evals.llm_eval.LLMEval
name: llm_eval
output_path: results/owasp_llm_10_evaluation.json
connector:
harmbench-eval-llama:
prompt_path: assets/eval-llm-10-2025.yaml
Multiple Evaluators
ARES supports running multiple evaluators in a single evaluation:
evaluation:
- keyword
- harmbench_eval
- garak_decoding_match
This allows comprehensive assessment using different evaluation methods.
Custom Evaluators
To create a custom evaluator:
Extend the base evaluator class from
ares.evalsImplement the required evaluation logic
Register it in your configuration
See the plugin development guide for more details.
Viewing Available Evaluators
Use the CLI to list all available evaluators:
ares show evals
To view a specific evaluator’s template:
ares show evals -n keyword
ares show evals -n harmbench_eval