Connector Configuration

Connectors are multi-purpose modules. They can be used as the primary target for evaluation and also within other modules via the connector keyword (e.g. use a LLM for evaluation or for generation of attack goals).

The target node defines the language model (SLM or LLM) that ARES will red-team or evaluate.

ARES uses a user-defined connectors.yaml file to configure model connectors. See example_configs/connectors.yaml for examples.

Viewing Available Connectors

Use the following commands to explore available connectors and their templates:

ares show connectors                # List all available connectors
ares show connectors -n huggingface  # Show template for HuggingFace connector

Example: HuggingFace Connector

Define the connector in connectors.yaml:

connectors:
  huggingface:
    type: ares.connectors.huggingface.HuggingFaceConnector
    name: huggingface
    model_config:
      pretrained_model_name_or_path: Qwen/Qwen2-0.5B-Instruct
      torch_dtype: bfloat16
    tokenizer_config:
      pretrained_model_name_or_path: Qwen/Qwen2-0.5B-Instruct
      padding_side: left
    generate_kwargs:
      chat_template:
        return_tensors: pt
        thinking: true
        return_dict: true
        add_generation_prompt: true
      generate_params:
        max_new_tokens: 50
    system_prompt:
      role: system
      content: Always speak like a pirate
    seed: 42
    device: auto

Then reference it in your ARES config (e.g., minimal.yaml):

target:
  huggingface:

Using Connectors in Other Modules

Other modules can also use connectors via the connector keyword. For example, the HuggingFaceEval module uses a model-as-a-judge approach:

evaluation:
  type: ares.evals.huggingface_eval.HuggingFaceEval
  name: harmbench_eval
  output_path: results/evaluation.json
  connector:
    harmbench-eval-llama:

System Prompt Configuration

ARES supports specifying a system prompt directly in the connector configuration. This allows you to define the initial context or behavior instructions for the model before any user prompts are sent.

Example:

target:
  huggingface:
    model_config:
      pretrained_model_name_or_path: ibm-granite/granite-3.3-8b-instruct
    tokenizer_config:
      pretrained_model_name_or_path: ibm-granite/granite-3.3-8b-instruct
    system_prompt:
     "role": "system"
     "content": |
       You are a helpful assistant. Always respond with concise and safe answers.

The system_prompt field is optional. If omitted, the connector uses its default system prompt (if applicable), which is automatically prepended to user input during evaluation.

Batch Generation

ARES can call batch_generate() instead of looping over generate() in two places:

  1. Attack execution — prompts generated by a single-turn attack strategy are sent to the target connector in chunks, reducing wall-clock time when running many attack prompts against a local (HuggingFace) or remote (LiteLLM) model.

  2. LLM-as-a-judge evaluation — when the eval connector (the judge model used by LLMEval) is a batch-capable type, evaluation prompts are also dispatched in batches, speeding up scoring with the same connector.

Set batch_size on whichever connector you want to run in batch mode — the same field works for both contexts.

How to enable it

Set batch_size to a value greater than 1 in the connector configuration:

connectors:
  huggingface:
    type: ares.connectors.huggingface.HuggingFaceConnector
    name: huggingface
    batch_size: 4          # send 4 prompts per batch_generate() call
    model_config:
      pretrained_model_name_or_path: ibm-granite/granite-3.3-8b-instruct
      torch_dtype: bfloat16
    tokenizer_config:
      pretrained_model_name_or_path: ibm-granite/granite-3.3-8b-instruct
      padding_side: left
    generate_kwargs:
      chat_template:
        return_tensors: pt
        add_generation_prompt: true
      generate_params:
        max_new_tokens: 200
        do_sample: false
    device: auto

The same batch_size field works identically for the LiteLLM connector plugin:

connectors:
  litellm:
    type: ares_litellm.LiteLLMConnector
    name: litellm
    batch_size: 8          # send 8 prompts per batch_generate() call
    model: watsonx/ibm/granite-3-8b-instruct
    endpoint-type: watsonx

Supported connectors

ARES core supports batch generation for:

  • HuggingFaceConnector (ares.connectors.huggingface.HuggingFaceConnector)

The ares-litellm plugin adds itself at import time — no ARES source change needed:

  • LiteLLMConnector (ares_litellm.LiteLLMConnector, requires the ares-litellm plugin)

Both connectors work identically whether the connector acts as the attack target or as the eval/judge model inside LLMEval. All other connectors fall back to sequential generate() calls, even if batch_size is set.

Adding batch support in your own plugin

If you write a connector plugin that implements a real batch_generate(), you can opt it in without touching ARES core. Add one call at the bottom of your plugin’s __init__.py:

from ares.connectors import register_batch_capable_connector
from my_plugin.connector import MyBatchConnector

register_batch_capable_connector(MyBatchConnector.__name__)

That is exactly what ares-litellm does. Once registered, ARES will honour batch_size > 1 for your connector in both attack execution and LLM-as-a-judge evaluation.

Handling transient failures with batch_retries

Remote endpoints (especially WatsonX via LiteLLM) occasionally return a generic provider error even at modest batch sizes. Setting batch_retries adds a halving-retry strategy: when a batch fails, it is split in two and each half is retried independently. This repeats up to batch_retries times. If all retries are exhausted, the remaining prompts are sent one-by-one via generate() so no result is permanently lost.

Example — 8 prompts, batch_size: 8, batch_retries: 2:

batch_generate([p1..p8])   → fails
  batch_generate([p1..p4]) → fails
    generate(p1), generate(p2), generate(p3), generate(p4)   ← sequential fallback
  batch_generate([p5..p8]) → succeeds
→ all 8 prompts have responses

Add batch_retries next to batch_size in your connector config:

connectors:
  litellm:
    type: ares_litellm.LiteLLMConnector
    name: litellm
    batch_size: 8
    batch_retries: 2       # halve up to 2 times before falling back to generate()
    model: watsonx/ibm/granite-3-8b-instruct
    endpoint-type: watsonx

The default is batch_retries: 0, which preserves the previous behaviour (a failed batch marks all its prompts as ERROR immediately).

Choosing a batch size

Connector

Guidance

HuggingFace (GPU)

Start with batch_size: 4. Values above 8 risk out-of-memory errors, especially with larger models or long prompts. The connector emits a runtime warning when more than 8 prompts are passed to batch_generate() in a single call. Watch GPU memory usage and lower the value if you see OOM crashes.

HuggingFace (CPU)

batch_size: 1 (default) is typically fastest on CPU due to memory bandwidth constraints.

LiteLLM (WatsonX / remote)

Start with batch_size: 8 and pair it with batch_retries: 2. If you still see intermittent failures, lower batch_size to 4.

Limitations

  • Batch generation is not available for multi-turn strategies (Crescendo, Echo Chamber, TAP, etc.). Multi-turn strategies manage per-prompt conversation state turn by turn, so each prompt must be processed independently. If batch_size > 1 is set on a connector used with a multi-turn strategy, ARES silently falls back to sequential generate() calls.

  • The multi-turn restriction applies only to the attack target. The eval connector used by LLMEval is always single-turn, so batch_size > 1 is always honoured there regardless of strategy type.

  • The default value of batch_size is 1, so existing configurations are unaffected.

Supported Connectors

ARES currently supports:

  • Hugging Face: for local model evaluation

  • LiteLLM: for common LLM providers (available as a plugin)

  • vLLM: for common LLM models (available as a plugin)

  • WatsonX: for remote model inference

  • GraniteIO: for interaction with GraniteIO models (available as a plugin)

  • WatsonX Orchestrate: for interaction with WatsonX Orchestrate Agents through Chat API (available as a plugin)

  • RESTful connectors: e.g., WatsonxAgentConnector for querying deployed agents via REST APIs

  • ICARUS connector: UI connector to a Streamlit-based agentic application ICARUS (available as a plugin)

This section explains how to configure targets in your YAML files and what credentials may be required.

If you are using connectors with gated access, make sure to add required API keys and other environment variables to ``.env``.

Note

In order to run models which are gated within Hugging Face hub, you must be logged in using the huggingface-cli and have READ permission for the gated repositories.

Note

In order to run models which are gated within WatsonX Platform, you must set your WATSONX_URL or WATSONX_API_BASE, WATSONX_API_KEY and WATSONX_PROJECT_ID variables in a .env file.

Note

In order to run agents which are gated within WatsonX AgentLab Platform, you must set your WATSONX_AGENTLAB_API_KEY variable in a .env file. This key can be found in your WatsonX Profile under the User API Key tab. More details are available at: https://dataplatform.cloud.ibm.com/docs/content/wsj/analyze-data/ml-authentication.html?context=wx

Explore more examples in the example_configs/ directory.

Connector classes abstract calls to LMs across different frameworks, making ARES extensible and adaptable.