Evaluators Reference

This section provides a comprehensive reference for all available evaluators in ARES. Evaluators analyze model responses to determine whether they meet safety, robustness, or refusal criteria.

Overview

ARES supports multiple types of evaluators:

  • Keyword-based: Pattern matching against refusal keywords

  • LLM-based: Model-as-a-judge evaluation using language models

  • Detector-based: Specialized detectors (e.g., Garak encoding detectors)

  • OWASP-specific: Evaluators tailored for specific OWASP vulnerabilities

Keyword Evaluator

Type: ares.evals.keyword_eval.KeywordEval

Evaluates responses by checking for the presence of refusal keywords or patterns.

Configuration:

evaluation:
  keyword:
    type: ares.evals.keyword_eval.KeywordEval
    name: keyword
    keyword_list_or_path: assets/advbench_refusal_keywords.json
    output_path: results/keyword_evaluation.json

Use Cases: - Quick refusal detection - Low-resource evaluation - Baseline safety checks

LLM-Based Evaluators

HarmBench Evaluator

Type: ares.evals.llm_eval.LLMEval

Uses a language model to judge whether responses are harmful based on HarmBench criteria.

Configuration:

evaluation:
  harmbench_eval:
    type: ares.evals.llm_eval.LLMEval
    name: harmbench_eval
    output_path: results/harmbench_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/harmbench_prompt.yaml

Features: - Model-as-a-judge approach - Nuanced safety assessment - Configurable evaluation prompts

HuggingFace Evaluator

Type: ares.evals.huggingface_eval.HuggingFaceEval

Similar to LLM evaluator but optimized for HuggingFace models.

Configuration:

evaluation:
  huggingface_eval:
    type: ares.evals.huggingface_eval.HuggingFaceEval
    name: harmbench_hf_eval
    output_path: results/harmbench_hf_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/harmbench_prompt.yaml

GPT Evaluator

Type: ares.evals.llm_eval.LLMEval

Uses GPT-OSS-120B on WatsonX for evaluation.

Configuration:

evaluation:
  gpt_eval:
    type: ares.evals.llm_eval.LLMEval
    name: gpt_eval
    output_path: results/gpt_evaluation.json
    connector:
      litellm_connector:
        name: gpt-oss-120b-on-watsonx
        endpoint-type: watsonx
        model: openai/gpt-oss-120b
        reasoning_effort: low
        prompt_path: assets/gptoss_prompt.yaml

Garak Detectors

Garak detectors are specialized evaluators from the NVIDIA Garak framework for detecting encoded attacks.

DecodeMatch Detector

Type: ares_garak.evals.encoding.Detector

Performs exact match detection on decoded content.

Configuration:

evaluation:
  garak_decoding_match:
    type: ares_garak.evals.encoding.Detector
    name: decode_match
    detector: detectors.encoding.DecodeMatch
    output_path: results/garak_detector_match_evaluation.json

DecodeApprox Detector

Type: ares_garak.evals.encoding.Detector

Performs approximate match detection on decoded content.

Configuration:

evaluation:
  garak_decoding_approx:
    type: ares_garak.evals.encoding.Detector
    name: decode_approx
    detector: detectors.encoding.DecodeApprox
    output_path: results/garak_detector_approx_evaluation.json

Use Cases: - Evaluating encoding attack success - Detecting obfuscated harmful content - Testing decoder robustness

Crescendo Evaluator

Type: ares_pyrit.evals.crescendo_eval.CrescendoNativeEval

Specialized evaluator for Crescendo multi-turn attacks.

Configuration:

evaluation:
  crescendo_eval:
    type: ares_pyrit.evals.crescendo_eval.CrescendoNativeEval
    name: crescendo_eval
    output_path: results/crescendo_evaluation.json
    connector:
      litellm_connector:
        name: gpt-oss-120b-on-watsonx
        endpoint-type: watsonx
        model: openai/gpt-oss-120b
        reasoning_effort: low
        prompt_path: assets/gptoss_prompt.yaml

Features: - Multi-turn conversation analysis - Progressive jailbreak detection - Context-aware evaluation

Multi-Turn Evaluator

Type: ares.evals.multi_turn_eval.MultiTurnEval

Unified evaluator for any strategy built on the MultiTurn base class. Reads the attack_successful and conversation_id fields that every multi-turn strategy writes and aggregates results at three granularities.

Evaluation Modes:

Mode

Behaviour

prompt

Evaluates each turn independently (1.0 = succeeded, 0.0 = failed, -1.0 = error)

conversation

Collapses all turns of a conversation into one entry; success if any turn succeeded

goal

Collapses all conversations for a goal into one entry; success if any conversation succeeded

Configuration:

evaluation:
  multi_turn_eval:
    type: ares.evals.multi_turn_eval.MultiTurnEval
    name: multi_turn_eval
    input_path: results/crescendo_attacks.json
    output_path: results/multi_turn_evaluation.json
    mode: conversation  # prompt | conversation | goal

Use Cases: - Evaluating Crescendo, Echo Chamber, TAP or any MultiTurn-based strategy - Reporting success at conversation or goal granularity for dashboards

AgentBreaker Evaluator

Type: ares_garak.evals.agent_breaker.AgentBreakerEval

Scores results produced by AgentBreakerStrategy. The probe internally runs a model-as-judge, so this evaluator reads pre-computed verdicts from notes["verified_results"] and maps them to ARES prediction scores.

Scoring:

  • 1.0 — attack succeeded (tool was exploited, confidence >= success_threshold)

  • 0.0 — attack failed

  • -1.0 — no response or connector error

Configuration:

evaluation:
  agent_breaker_eval:
    type: ares_garak.evals.agent_breaker.AgentBreakerEval
    name: agent_breaker_eval
    output_path: results/agent_breaker_evaluation.json
    success_threshold: 0.7  # minimum confidence to count as success

Privacy Evaluator (READI)

Type: ares_privacy_evaluator.READIEvaluator

Plugin: ares-privacy-evaluator

Detects PII or PHI leakage in target responses using the READI analyser. Returns 0.0 when sensitive entities are found (attack succeeded — model leaked data) and 1.0 otherwise.

Configuration:

evaluation:
  privacy_eval:
    type: ares_privacy_evaluator.READIEvaluator
    name: privacy_eval
    input_path: assets/attack_results.json
    output_path: results/privacy_evaluation.json
    sensitive_type: PII   # PII | PHI
    exclude_prompt: true  # exclude entities already present in the prompt
    debug_mode: false
    exclude_patterns: []  # list of regex patterns to ignore

Use Cases: - Measuring PII/PHI leakage from RAG or document-grounded systems - Complement to LLM-based evaluators when structured entity detection is required

Intrinsic Evaluator

Type: ares_intrinsics.evals.intrinsics.IntrinsicEval

Plugin: ares-intrinsics

Extends LLMEval with LoRA/aLoRA adapters from ibm-granite/granite-3.3-8b-security-lib for efficient single-forward-pass evaluation. Supports RAG data leakage and PII detection intrinsics; multiple vulnerability types can be checked without separate model calls.

Built-in intrinsics:

  • granite-3.3-8b-instruct-lora-rag-data-leakage — RAG context leakage detection

  • granite-3.3-8b-instruct-lora-pii-detector — PII detection

Configuration:

evaluation:
  intrinsic_eval:
    type: ares_intrinsics.evals.intrinsics.IntrinsicEval
    name: intrinsic_eval
    output_path: results/intrinsic_evaluation.json
    intrinsic: granite-3.3-8b-instruct-lora-pii-detector
    connector:
      hf_eval_model:
        type: ares.connectors.huggingface.HuggingFaceConnector
        name: granite-3.3-8b-instruct
        model_id: ibm-granite/granite-3.3-8b-instruct

Langfuse Evaluator

Type: ares_langfuse_eval.evals.langfuse_eval.LangfuseEval

Plugin: ares-langfuse-eval

Retrieves LLM traces from a Langfuse instance via REST API and evaluates attack effectiveness by delegating response analysis to a nested ARES evaluator. Useful when the target is instrumented with Langfuse and responses are not captured directly by ARES.

Configuration:

evaluation:
  langfuse_eval:
    type: ares_langfuse_eval.evals.langfuse_eval.LangfuseEval
    name: langfuse_evaluator
    output_path: results/langfuse_evaluation.json
    langfuse_config:
      base_url: "https://cloud.langfuse.com"
      public_key_var: LANGFUSE_PUBLIC_KEY
      secret_key_var: LANGFUSE_SECRET_KEY
    input_config:
      traceid_path: "$.trace_id"
      limit: 100
    evaluator:
      type: ares.evals.llm_eval.LLMEval
      name: llm_eval
      output_path: results/inner_eval.json
      connector:
        litellm_connector:
          type: ares_litellm.LiteLLMConnector
          name: judge
          endpoint-type: ollama
          model: ollama/llama3

Use Cases: - Evaluating attacks against externally-hosted or agentic targets instrumented with Langfuse - Post-hoc analysis of red-team campaigns using existing trace data

OWASP-Specific Evaluators

ARES provides specialized evaluators for each OWASP LLM vulnerability category.

LLM01 Evaluator (Prompt Injection)

Type: ares.evals.huggingface_eval.HuggingFaceEval

Configuration:

evaluation:
  llm01_eval:
    type: ares.evals.huggingface_eval.HuggingFaceEval
    name: llm01_harmbench_eval
    output_path: results/owasp_llm_01_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/harmbench_prompt.yaml

LLM02 Evaluator (Sensitive Information Disclosure)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm02_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_02_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-02-2025.yaml

LLM04 Evaluator (Data and Model Poisoning)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm04_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_04_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-04-2025.yaml

LLM05 Evaluator (Improper Output Handling)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm05_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_05_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-05-2025.yaml

LLM06 Evaluator (Excessive Agency)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm06_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_06_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-06-2025.yaml

LLM07 Evaluator (System Prompt Leakage)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm07_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_07_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-07-2025.yaml

LLM09 Evaluator (Misinformation)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm09_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_09_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-09-2025.yaml

LLM10 Evaluator (Unbounded Consumption)

Type: ares.evals.llm_eval.LLMEval

Configuration:

evaluation:
  llm10_eval:
    type: ares.evals.llm_eval.LLMEval
    name: llm_eval
    output_path: results/owasp_llm_10_evaluation.json
    connector:
      harmbench-eval-llama:
        prompt_path: assets/eval-llm-10-2025.yaml

Multiple Evaluators

ARES supports running multiple evaluators in a single evaluation:

evaluation:
  - keyword
  - harmbench_eval
  - garak_decoding_match

This allows comprehensive assessment using different evaluation methods.

Custom Evaluators

To create a custom evaluator:

  1. Extend the base evaluator class from ares.evals

  2. Implement the required evaluation logic

  3. Register it in your configuration

See the plugin development guide for more details.

Viewing Available Evaluators

Use the CLI to list all available evaluators:

ares show evals

To view a specific evaluator’s template:

ares show evals -n keyword
ares show evals -n harmbench_eval