This notebook explains how to uncover risks related to your usecase based on a given taxonomy.¶
Import libraries¶
from ai_atlas_nexus.blocks.inference import (
RITSInferenceEngine,
WMLInferenceEngine,
OllamaInferenceEngine,
VLLMInferenceEngine,
HFInferenceEngine,
OpenAIInferenceEngine,
AWSBedrockInferenceEngine
)
from ai_atlas_nexus.blocks.inference.params import (
InferenceEngineCredentials,
RITSInferenceEngineParams,
WMLInferenceEngineParams,
OllamaInferenceEngineParams,
VLLMInferenceEngineParams,
HFInferenceEngineParams,
OpenAIInferenceEngineParams,
AWSBedrockInferenceEngineParams
)
from ai_atlas_nexus.library import AIAtlasNexus
import os
AI Atlas Nexus uses Large Language Models (LLMs) to infer risks dimensions. Therefore requires access to LLMs to inference or call the model.¶
Available Inference Engines: WML, Ollama, vLLM, RITS, HF. Please follow the Inference APIs guide before going ahead.
Note: RITS is intended solely for internal IBM use and requires TUNNELALL VPN for access.
# inference_engine = OllamaInferenceEngine(
# model_name_or_path="granite3.3:8b",
# credentials=InferenceEngineCredentials(api_url="http://localhost:11434"),
# parameters=OllamaInferenceEngineParams(
# num_predict=1000, num_ctx=8192, temperature=0
# ),
# )
# inference_engine = HFInferenceEngine(
# model_name_or_path="meta-llama/Llama-3.1-8B-Instruct",
# credentials=InferenceEngineCredentials(
# api_key=os.getenv("HF_TOKEN"),
# api_url="https://router.huggingface.co/v1",
# ),
# parameters=HFInferenceEngineParams(max_completion_tokens=1000, temperature=0),
# )
inference_engine = WMLInferenceEngine(
model_name_or_path="meta-llama/llama-3-3-70b-instruct",
credentials={
"api_key": os.getenv("WML_API_KEY"),
"api_url": os.getenv("WML_API_URL"),
"project_id": os.getenv("WML_PROJECT_ID"),
},
parameters=WMLInferenceEngineParams(
max_completion_tokens=1024, temperature=0, seed=99
),
)
# inference_engine = VLLMInferenceEngine(
# model_name_or_path="ibm-granite/granite-3.3-8b-instruct",
# credentials=InferenceEngineCredentials(
# api_url=os.getenv("VLLM_API_URL"), api_key=os.getenv("VLLM_API_KEY")
# ),
# parameters=VLLMInferenceEngineParams(max_tokens=1000, temperature=0),
# )
# inference_engine = RITSInferenceEngine(
# model_name_or_path="ibm-granite/granite-3.3-8b-instruct",
# credentials={
# "api_key": os.getenv("RITS_API_KEY"),
# "api_url": os.getenv("RITS_API_URL"),
# },
# parameters=RITSInferenceEngineParams(max_completion_tokens=1000, temperature=0),
# )
# inference_engine = OpenAIInferenceEngine(
# model_name_or_path="gpt-5-mini",
# credentials={"api_key": os.getenv("OPENAI_API_KEY")},
# parameters=OpenAIInferenceEngineParams(max_completion_tokens=1000),
# )
# inference_engine = AWSBedrockInferenceEngine(
# # model_name_or_path="us.meta.llama3-3-70b-instruct-v1:0",
# model_name_or_path="openai.gpt-oss-120b-1:0",
# credentials={
# "aws_access_key_id": os.getenv("AWS_ACCESS_KEY_ID"),
# "aws_secret_access_key": os.getenv("AWS_SECRET_ACCESS_KEY"),
# "region_name": os.getenv("AWS_DEFAULT_REGION"),
# },
# parameters=AWSBedrockInferenceEngineParams(maxTokens=4096, seed=42, temperature=0, reasoning_effort="low"),
# )
[2026-08-21 11:03:21:491] - INFO - AIAtlasNexus - ✓ Created WML inference engine for model: meta-llama/llama-3-3-70b-instruct, backend - DEFAULT
Create an instance of AIAtlasNexus¶
Note: (Optional) You can specify your own directory in AIAtlasNexus(base_dir=<PATH>) to utilize custom AI ontologies. If left blank, the system will use the provided AI ontologies.
ai_atlas_nexus = AIAtlasNexus()
[2026-08-21 11:03:22:96] - INFO - AIAtlasNexus - Created AIAtlasNexus instance. Base_dir: None
Risk Identification API¶
AIAtlasNexus.identify_risks_from_usecases()
Params:
- usecases (List[str]): A List of strings describing AI usecases
- inference_engine (InferenceEngine): An LLM inference engine to infer risks from the usecases.
- taxonomy (str, optional): The string label for a taxonomy. Default to None.
- cot_examples (Dict[str, List], optional): The Chain of Thought (CoT) examples to use in the risk identification. The example template is available at src/ai_atlas_nexus/data/templates/risk_generation_cot.json. Assign the ID of the taxonomy you wish to use as the key for CoT examples. Providing this value will override the CoT examples present in the template master. Default to None.
- max_risk (int, optional): The maximum number of risks to extract. Pass None to allow the inference engine to determine the number of risks. Defaults to None.
- zero_shot_only (bool): If enabled, this flag allows the system to perform Zero Shot Risk identification, and the field cot_examples will be ignored.
- batch_inference (bool): Whether to run risk inference service in batch mode or at each risk level. Defaults to True.
- use_dspy_prompt (bool): Use per-risk DSPy optmized prompt instructions for risk identification. When enabled, batch_inference flag is ignored.
- return_metadata (bool): Wrap the result in a DetectionResult carrying token usage, both aggregated over the run and broken down per usecase. Defaults to False.
- explanation_type (ExplanationType, optional): Pair each risk with an explanation (NONE, DESCRIPTION, REASONING, SELF_EXPLANATION). Defaults to NONE.
Risk Identification using IBM AI Risk taxonomy - Batch Inference¶
Token Usage Metadata (Optional)¶
New Feature: You can now track token usage metrics from your risk identification operations by setting return_metadata=True. This is useful for cost analysis, monitoring, and performance tracking.
result.metadata reports the run as a whole:
token_usage.input_tokens: Tokens provided to the modeltoken_usage.output_tokens: Tokens generated by the modeltoken_usage.total_tokens: Sum of input and output tokensnum_calls: Number of underlying LLM calls mademodel: Model name/path usedinference_engine: Type of inference engineseed: The seed used, or None if the calls did not all share onestop_reason_summary: How often each stop reason occurred, e.g.{'eos': 5}has_thinking: Whether any call returned thinking
result.metadata.per_usecase reports the same figures for each usecase separately, in the order the usecases were passed in, so spend can be attributed to the usecase that incurred it. Its entries sum to the aggregate.
How many calls each entry covers depends on the mode: with batch_inference=True (the default) it is one call per usecase, while batch_inference=False or use_dspy_prompt=True makes one call per risk per usecase.
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
# Get risks
result_risk_only = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
)
result_risk_only
Inferring with WML, backend - DEFAULT: 0%| | 0/1 [00:00<?, ?it/s]/Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages/ibm_watsonx_ai/wml_resource.py:100: WatsonxAPIWarning: This model is a Non-IBM Product governed by a third-party license that may impose use restrictions and other obligations. By using this model you agree to its terms as identified in the following URL. ID: disclaimer_warning More info: https://dataplatform.cloud.ibm.com/docs/content/wsj/analyze-data/fm-models.html?context=wx warn(cls._build_warning_message(warning), WatsonxAPIWarning) Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:01<00:00, 1.05s/it]
[[Risk(id='atlas-lack-of-model-transparency', name='Lack of model transparency', description='Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-model-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-governance', 'eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-model-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="Transparency is important for legal compliance, AI ethics, and guiding appropriate use of models. Missing information might make it more difficult to evaluate risks, change the model, or reuse it.\xa0 Knowledge about who built a model can also be an important factor in deciding whether to trust it. Additionally, transparency regarding how the model's risks were determined, evaluated, and mitigated also play a role in determining model risks, identifying model suitability, and governing model usage."), Risk(id='atlas-data-provenance', name='Uncertain data provenance', description='Data provenance refers to the traceability of data (including synthetic data), which includes its ownership, origin, transformations, and generation. Proving that the data is the same as the original source with correct usage terms is difficult without standardized methods for verifying data sources or generation.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-provenance.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['aiuc1-req-e006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['llm032025-supply-chain', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-provenance', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Not all data sources are trustworthy. Data might be unethically collected, manipulated, or falsified. Verifying that data provenance is challenging due to factors such as data volume, data complexity, data source varieties, poor data management, and synthetic data generation methods. Using such data can result in undesirable behaviors in the model.'), Risk(id='atlas-data-bias', name='Data bias', description='Historical and societal biases might be present in data that are used to train and fine-tune models. Biases can also be inherited from seed data or exacerbated by synthetic data generation methods.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-bias.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['asi06-memory-and-context-poisoning', 'shieldgemma-hate-speech', 'aiuc1-req-c003', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-post-deployment', 'mit-ai-risk-subdomain-1.1'], narrow_mappings=None, broad_mappings=['nist-harmful-bias-and-homogenization'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-fairness', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-bias', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Training an AI system on data with bias, such as historical or societal bias, can lead to biased or skewed outputs that can unfairly represent or otherwise discriminate against certain groups or individuals.'), Risk(id='atlas-data-transparency', name='Lack of training data transparency', description="Proper documentation contains information about how a model's data was collected, curated, and used to train a model, including any synthetic data generation processes. Without proper documentation it might be harder to satisfactorily explain the behavior of the model.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['credo-risk-005'], related_mappings=['credo-risk-006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-information-integrity', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-transparency', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='A lack of data documentation limits the ability to evaluate risks associated with the data. Having access to the training data is not enough. Without recording how the data was cleaned, modified, or generated, including any data augmentation or synthetic data generation steps, the model behavior is more difficult to understand and to fix. Lack of data transparency also impacts model reuse as it is difficult to determine data representativeness for the new use without such documentation.'), Risk(id='atlas-lack-of-system-transparency', name='Lack of system transparency', description="Insufficient documentation of the system that uses the model and the model's purpose within the system in which it is used.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-system-transparency.html', dateCreated=datetime.date(2024, 9, 24), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['asi09-human-agent-trust-exploitation', 'aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-system-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="A lack of documentation makes it difficult to understand how the model's outcomes contribute to the system's or application's functionality.")]]
# Get risks WITH token usage metadata
result = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
return_metadata=True, # Enable metadata tracking
)
result #observe the return of a DetectionResult, including data and metadata
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:01<00:00, 1.50s/it]
DetectionResult(data=[[Risk(id='atlas-lack-of-model-transparency', name='Lack of model transparency', description='Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-model-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-governance', 'eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-model-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="Transparency is important for legal compliance, AI ethics, and guiding appropriate use of models. Missing information might make it more difficult to evaluate risks, change the model, or reuse it.\xa0 Knowledge about who built a model can also be an important factor in deciding whether to trust it. Additionally, transparency regarding how the model's risks were determined, evaluated, and mitigated also play a role in determining model risks, identifying model suitability, and governing model usage."), Risk(id='atlas-data-provenance', name='Uncertain data provenance', description='Data provenance refers to the traceability of data (including synthetic data), which includes its ownership, origin, transformations, and generation. Proving that the data is the same as the original source with correct usage terms is difficult without standardized methods for verifying data sources or generation.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-provenance.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['aiuc1-req-e006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['llm032025-supply-chain', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-provenance', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Not all data sources are trustworthy. Data might be unethically collected, manipulated, or falsified. Verifying that data provenance is challenging due to factors such as data volume, data complexity, data source varieties, poor data management, and synthetic data generation methods. Using such data can result in undesirable behaviors in the model.'), Risk(id='atlas-data-bias', name='Data bias', description='Historical and societal biases might be present in data that are used to train and fine-tune models. Biases can also be inherited from seed data or exacerbated by synthetic data generation methods.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-bias.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['asi06-memory-and-context-poisoning', 'shieldgemma-hate-speech', 'aiuc1-req-c003', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-post-deployment', 'mit-ai-risk-subdomain-1.1'], narrow_mappings=None, broad_mappings=['nist-harmful-bias-and-homogenization'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-fairness', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-bias', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Training an AI system on data with bias, such as historical or societal bias, can lead to biased or skewed outputs that can unfairly represent or otherwise discriminate against certain groups or individuals.'), Risk(id='atlas-data-transparency', name='Lack of training data transparency', description="Proper documentation contains information about how a model's data was collected, curated, and used to train a model, including any synthetic data generation processes. Without proper documentation it might be harder to satisfactorily explain the behavior of the model.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['credo-risk-005'], related_mappings=['credo-risk-006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-information-integrity', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-transparency', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='A lack of data documentation limits the ability to evaluate risks associated with the data. Having access to the training data is not enough. Without recording how the data was cleaned, modified, or generated, including any data augmentation or synthetic data generation steps, the model behavior is more difficult to understand and to fix. Lack of data transparency also impacts model reuse as it is difficult to determine data representativeness for the new use without such documentation.'), Risk(id='atlas-lack-of-system-transparency', name='Lack of system transparency', description="Insufficient documentation of the system that uses the model and the model's purpose within the system in which it is used.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-system-transparency.html', dateCreated=datetime.date(2024, 9, 24), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['asi09-human-agent-trust-exploitation', 'aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-system-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="A lack of documentation makes it difficult to understand how the model's outcomes contribute to the system's or application's functionality.")]], metadata=InferenceMetadata(token_usage=TokenUsage(input_tokens=5985, output_tokens=41, total_tokens=6026), inference_engine='WML', model='meta-llama/llama-3-3-70b-instruct', num_calls=1, seed=None, stop_reason_summary={'stop': 1}, has_thinking=False, per_usecase=[UsecaseInferenceMetadata(token_usage=TokenUsage(input_tokens=5985, output_tokens=41, total_tokens=6026), num_calls=1, seed=None, stop_reason_summary={'stop': 1}, has_thinking=False)]))
# Access the risks
risks = result.data
print("Identified risks:")
for risk in risks[0]:
print(f" - {risk.name}")
# Access token usage metrics for the whole run
print(f"\nToken Usage:")
print(f" Input tokens: {result.metadata.token_usage.input_tokens}")
print(f" Output tokens: {result.metadata.token_usage.output_tokens}")
print(f" Total tokens: {result.metadata.token_usage.total_tokens}")
print(f"\nInference Info:")
print(f" Model: {result.metadata.model}")
print(f" Engine: {result.metadata.inference_engine}")
print(f" Number of LLM calls: {result.metadata.num_calls}")
# The same figures per usecase, aligned with result.data
print(f"\nPer usecase:")
for index, usage in enumerate(result.metadata.per_usecase):
print(f" usecase {index}: {usage.token_usage.total_tokens} tokens "
f"({usage.token_usage.input_tokens} in / {usage.token_usage.output_tokens} out) "
f"across {usage.num_calls} call(s)")
# Calculate cost (example: $0.001 per 1000 tokens)
cost_per_1k_tokens = 0.001
estimated_cost = (result.metadata.token_usage.total_tokens / 1000) * cost_per_1k_tokens
print(f"\nEstimated cost (at ${cost_per_1k_tokens}/1k tokens): ${estimated_cost:.4f}")
Identified risks: - Lack of model transparency - Uncertain data provenance - Data bias - Lack of training data transparency - Lack of system transparency Token Usage: Input tokens: 5985 Output tokens: 41 Total tokens: 6026 Inference Info: Model: meta-llama/llama-3-3-70b-instruct Engine: WML Number of LLM calls: 1 Per usecase: usecase 0: 6026 tokens (5985 in / 41 out) across 1 call(s) Estimated cost (at $0.001/1k tokens): $0.0060
Risk Explanations with Flexible Types (Optional)¶
You can return risks paired with an explanation. Use the explanation_type parameter to choose where that explanation comes from.
Available explanation types:
ExplanationType.NONE(default) — No explanations, bareRiskobjectsExplanationType.DESCRIPTION— The risk's description from the ontologyExplanationType.REASONING— The model's thinking, for models that expose itExplanationType.SELF_EXPLANATION- Explanation from model's response itself
When explanation_type is anything other than NONE, each item is a RiskWithExplanation
which includes a .risk and .explanation.
REASONING needs an inference engine that reports the model's thinking back to us, which
today is Ollama only (with think=True, on a model whose capabilities include thinking).
On every other engine the explanation is None rather than an error, so prefer
DESCRIPTION or SELF_EXPLANATION there.
from ai_atlas_nexus.blocks.inference import ExplanationType
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
usecase2 = "an ai chatbot for holiday travel booking recommendations"
risks_with_explanations = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase, usecase2],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
explanation_type=ExplanationType.SELF_EXPLANATION,
return_metadata=True,
)
# return_metadata=True wraps the result, so the risks are under .data
print("Risks with self explanations:")
for risk_with_explanation in risks_with_explanations.data[0][:3]: # Show first 3
print(f"\n{risk_with_explanation.risk.name}:")
print(f" {risk_with_explanation.explanation}")
# Both usecases were analyzed, so the metadata carries an entry for each
for index, usage in enumerate(risks_with_explanations.metadata.per_usecase):
print(f"\nusecase {index}: {usage.token_usage.total_tokens} tokens, "
f"{len(risks_with_explanations.data[index])} risks")
Inferring with WML, backend - DEFAULT: 100%|██████████| 2/2 [00:04<00:00, 2.40s/it]
Risks with self explanations: Lack of model transparency: The model may not be transparent about how it generates responses or recommendations, which could lead to mistrust or confusion among customers or support agents Data bias: The model may provide biased or unfair responses to customers based on their personal information or claims history Lack of training data transparency: The model may not have access to complete or accurate data about customers or their claims, which could lead to incomplete or inaccurate responses or recommendations usecase 0: 6220 tokens, 5 risks usecase 1: 6196 tokens, 5 risks
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
)
for risk in risks[0]:
print(risk.name)
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:01<00:00, 1.16s/it]
Lack of model transparency Uncertain data provenance Data bias Lack of training data transparency Lack of system transparency
# you may wish to retrieve the risk list and the control measures together, so can use the method
# `identify_risks_and_actions_from_usecases`. It reports one entry per usecase under
# `per_usecase`, each with that usecase's risks, summary and control items.
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks_and_measures = ai_atlas_nexus.identify_risks_and_actions_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
)
for risk in risks_and_measures["per_usecase"][0]["risks"]:
print(risk.name)
risks_and_measures
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:01<00:00, 1.04s/it]
Lack of model transparency Uncertain data provenance Data bias Lack of training data transparency Lack of system transparency
{'usecases': ['Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers.'],
'model': 'meta-llama/llama-3-3-70b-instruct',
'taxonomy': 'ibm-risk-atlas',
'per_usecase': [{'usecase': 'Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers.',
'risks': [Risk(id='atlas-lack-of-model-transparency', name='Lack of model transparency', description='Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-model-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-governance', 'eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-model-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="Transparency is important for legal compliance, AI ethics, and guiding appropriate use of models. Missing information might make it more difficult to evaluate risks, change the model, or reuse it.\xa0 Knowledge about who built a model can also be an important factor in deciding whether to trust it. Additionally, transparency regarding how the model's risks were determined, evaluated, and mitigated also play a role in determining model risks, identifying model suitability, and governing model usage."),
Risk(id='atlas-data-provenance', name='Uncertain data provenance', description='Data provenance refers to the traceability of data (including synthetic data), which includes its ownership, origin, transformations, and generation. Proving that the data is the same as the original source with correct usage terms is difficult without standardized methods for verifying data sources or generation.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-provenance.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['aiuc1-req-e006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['llm032025-supply-chain', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-provenance', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Not all data sources are trustworthy. Data might be unethically collected, manipulated, or falsified. Verifying that data provenance is challenging due to factors such as data volume, data complexity, data source varieties, poor data management, and synthetic data generation methods. Using such data can result in undesirable behaviors in the model.'),
Risk(id='atlas-data-bias', name='Data bias', description='Historical and societal biases might be present in data that are used to train and fine-tune models. Biases can also be inherited from seed data or exacerbated by synthetic data generation methods.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-bias.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['asi06-memory-and-context-poisoning', 'shieldgemma-hate-speech', 'aiuc1-req-c003', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-post-deployment', 'mit-ai-risk-subdomain-1.1'], narrow_mappings=None, broad_mappings=['nist-harmful-bias-and-homogenization'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-fairness', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-bias', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Training an AI system on data with bias, such as historical or societal bias, can lead to biased or skewed outputs that can unfairly represent or otherwise discriminate against certain groups or individuals.'),
Risk(id='atlas-data-transparency', name='Lack of training data transparency', description="Proper documentation contains information about how a model's data was collected, curated, and used to train a model, including any synthetic data generation processes. Without proper documentation it might be harder to satisfactorily explain the behavior of the model.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['credo-risk-005'], related_mappings=['credo-risk-006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-information-integrity', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-transparency', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='A lack of data documentation limits the ability to evaluate risks associated with the data. Having access to the training data is not enough. Without recording how the data was cleaned, modified, or generated, including any data augmentation or synthetic data generation steps, the model behavior is more difficult to understand and to fix. Lack of data transparency also impacts model reuse as it is difficult to determine data representativeness for the new use without such documentation.'),
Risk(id='atlas-lack-of-system-transparency', name='Lack of system transparency', description="Insufficient documentation of the system that uses the model and the model's purpose within the system in which it is used.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-system-transparency.html', dateCreated=datetime.date(2024, 9, 24), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['asi09-human-agent-trust-exploitation', 'aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-system-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="A lack of documentation makes it difficult to understand how the model's outcomes contribute to the system's or application's functionality.")],
'summary': defaultdict(list,
{'risk_ids': ['atlas-lack-of-model-transparency',
'atlas-data-provenance',
'atlas-data-bias',
'atlas-data-transparency',
'atlas-lack-of-system-transparency'],
'action_ids': [],
'detector_ids': [],
'Requirement': ['aiuc1-req-e006',
'aiuc1-req-c003',
'aiuc1-req-e017'],
'RiskGroup': ['eticas-governance']}),
'mixed_control_items': [Requirement(id='aiuc1-req-e006', name='Conduct vendor due diligence', description='Establish AI vendor due diligence processes for foundation and upstream model providers covering data handling, PII controls, security and compliance', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['llm032025-supply-chain', 'atlas-ai-agent-compliance-agentic', 'atlas-data-provenance', 'atlas-data-usage-rights', 'atlas-model-usage-rights'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='aiuc1', hasRule=['aiuc1-ctrl-e006-1'], type='Requirement', hasApplication=['MANDATORY'], hasFrequency='MONTHS_12', hasKeywords=None, hasPrinciple=['aiuc1-principle-e'], isApplicableToCapability=None, hasRequirementType=None),
RiskGroup(id='eticas-governance', name='Governance', description='The risk that an AI system lacks adequate structures, policies, or accountability mechanisms to oversee its design, deployment, and use.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=None, narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='eticas-ai-risk-taxonomy', hasDocumentation=None, hasPart=None, belongsToDomain=None, type='RiskGroup', narrower=None, broader=None, hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None),
Requirement(id='aiuc1-req-c003', name='Prevent harmful outputs', description='Implement safeguards or technical controls to prevent harmful outputs including distressed outputs, angry responses, high-risk advice, offensive content, bias, and deception', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['atlas-harmful-output', 'atlas-toxic-output'], related_mappings=['llm052025-improper-output-handling', 'llm092025-misinformation', 'atlas-data-bias', 'atlas-decision-bias', 'atlas-introduce-data-bias-agentic', 'atlas-output-bias', 'atlas-spreading-disinformation'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='aiuc1', hasRule=['aiuc1-ctrl-c003-1', 'aiuc1-ctrl-c003-2', 'aiuc1-ctrl-c003-3', 'aiuc1-ctrl-c003-4'], type='Requirement', hasApplication=['MANDATORY'], hasFrequency='MONTHS_12', hasKeywords=None, hasPrinciple=['aiuc1-principle-c'], isApplicableToCapability=None, hasRequirementType=None),
Requirement(id='aiuc1-req-e017', name='Document system transparency policy', description='Establish a system transparency policy and maintain a repository of model cards, datasheets, and interpretability reports for major systems', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['atlas-lack-of-ai-agent-transparency-agentic', 'atlas-lack-of-model-transparency', 'atlas-lack-of-system-transparency'], related_mappings=['atlas-lack-of-ai-agent-transparency-agentic', 'atlas-lack-of-data-transparency', 'atlas-lack-of-model-transparency', 'atlas-lack-of-system-transparency'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='aiuc1', hasRule=['aiuc1-ctrl-e017-1', 'aiuc1-ctrl-e017-2', 'aiuc1-ctrl-e017-3'], type='Requirement', hasApplication=['OPTIONAL'], hasFrequency='MONTHS_12', hasKeywords=None, hasPrinciple=['aiuc1-principle-e'], isApplicableToCapability=None, hasRequirementType=None)]}]}
Risk Identification using IBM AI Risk taxonomy - Per Risk Inference¶
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
batch_inference=False,
)
print(len(risks[0]))
for risk in risks[0]:
print(risk.name)
Inferring with WML, backend - DEFAULT: 100%|██████████| 99/99 [00:20<00:00, 4.94it/s]
24 Data privacy rights alignment Hallucination Confidential information in data Lack of model transparency Personal information in prompt Lack of testing diversity Decision bias Exposing personal information Improper data curation Over- or under-reliance on AI agents Revealing confidential information Data bias Incomplete usage definition Incomplete AI agent evaluation Non-disclosure Reproducibility Incomplete advice Personal information in data Data acquisition restrictions Prompt priming Poor model accuracy Output bias Unexplainable output Unreliable source attribution
Risk Identification using IBM AI Risk taxonomy - Per Risk Inference using DSPy optimised prompt¶
dspy_inference_engine = RITSInferenceEngine(
model_name_or_path="meta-llama/llama-3-3-70b-instruct",
credentials={
"api_key": os.getenv("RITS_API_KEY"),
"api_url": os.getenv("RITS_API_URL"),
},
parameters=RITSInferenceEngineParams(max_completion_tokens=1000, temperature=0),
)
[2026-08-21 11:03:52:772] - INFO - AIAtlasNexus - ✓ Created RITS inference engine for model: meta-llama/llama-3-3-70b-instruct, backend - DEFAULT
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=dspy_inference_engine,
taxonomy="ibm-risk-atlas",
max_risk=5,
use_dspy_prompt=True,
)
print(len(risks[0]))
for risk in risks[0]:
print(risk.name)
Inferring with RITS, backend - DEFAULT: 0%| | 0/99 [00:00<?, ?it/s]
Inferring with RITS, backend - DEFAULT: 100%|██████████| 99/99 [00:13<00:00, 7.56it/s]
58 Over- or under-reliance Membership inference attack Confidential data in prompt Data privacy rights alignment Discriminatory actions IP information in prompt Hallucination Mitigation and maintenance AI agent compliance Confidential information in data Lack of model transparency Unrepresentative data Personal information in prompt Sharing IP/PI/confidential information with user Lack of testing diversity Decision bias Exposing personal information AI agents' impact on jobs Improper data curation Over- or under-reliance on AI agents Revealing confidential information Spreading disinformation Uncertain data provenance Data bias Data usage rights restrictions Unauthorized use AI agents' impact on environment Misaligned actions Data contamination Incomplete usage definition Lack of data transparency Copyright infringement Impact on affected communities Improper retraining Incomplete AI agent evaluation Non-disclosure Reproducibility Specialized tokens attack Incomplete advice Prompt injection attack Data usage restrictions Personal information in data Impact on Jobs Data acquisition restrictions Sharing IP/PI/confidential information with tools Prompt priming Reidentification Attribute inference attack Poor model accuracy Data transfer restrictions Generated content ownership and IP Lack of AI agent transparency Impact on human dignity Output bias Unexplainable output Toxic output Unexplainable and untraceable actions Overfitting
Risk Identification using IBM AI Risk taxonomy with Custom CoT examples¶
Note: To use custom risk cot_examples for a new taxonomy or an existing taxonomy, users must provide their own set of example usecase and associated risks in a JSON file such as in risk_generation_cot.json. This will enable the LLM to learn from these few-shot examples and generate better responses. Please follow the guide here.
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-risk-atlas",
cot_examples={
"ibm-risk-atlas": [
{
"Usecase": "In a medical chatbot, generative AI can be employed to create a triage system that assesses patients' symptoms and provides immediate, contextually relevant advice based on their medical history and current condition. The chatbot can analyze the patient's input, identify potential medical issues, and offer tailored recommendations or insights to the patient or healthcare provider. This can help streamline the triage process, ensuring that patients receive the appropriate level of care and attention, and ultimately improving patient outcomes.",
"Risks": [
"Improper usage",
"Incomplete advice",
"Lack of model transparency",
"Lack of system transparency",
"Lack of training data transparency",
"Data bias",
"Uncertain data provenance",
"Lack of data transparency",
"Impact on human agency",
"Impact on affected communities",
"Improper retraining",
"Inaccessible training data",
],
}
]
},
max_risk=5,
)
risks
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:01<00:00, 1.17s/it]
[[Risk(id='atlas-lack-of-model-transparency', name='Lack of model transparency', description='Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-model-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-governance', 'eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-model-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="Transparency is important for legal compliance, AI ethics, and guiding appropriate use of models. Missing information might make it more difficult to evaluate risks, change the model, or reuse it.\xa0 Knowledge about who built a model can also be an important factor in deciding whether to trust it. Additionally, transparency regarding how the model's risks were determined, evaluated, and mitigated also play a role in determining model risks, identifying model suitability, and governing model usage."), Risk(id='atlas-data-provenance', name='Uncertain data provenance', description='Data provenance refers to the traceability of data (including synthetic data), which includes its ownership, origin, transformations, and generation. Proving that the data is the same as the original source with correct usage terms is difficult without standardized methods for verifying data sources or generation.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-provenance.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['aiuc1-req-e006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['llm032025-supply-chain', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-provenance', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Not all data sources are trustworthy. Data might be unethically collected, manipulated, or falsified. Verifying that data provenance is challenging due to factors such as data volume, data complexity, data source varieties, poor data management, and synthetic data generation methods. Using such data can result in undesirable behaviors in the model.'), Risk(id='atlas-data-bias', name='Data bias', description='Historical and societal biases might be present in data that are used to train and fine-tune models. Biases can also be inherited from seed data or exacerbated by synthetic data generation methods.', url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-bias.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=None, related_mappings=['asi06-memory-and-context-poisoning', 'shieldgemma-hate-speech', 'aiuc1-req-c003', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-post-deployment', 'mit-ai-risk-subdomain-1.1'], narrow_mappings=None, broad_mappings=['nist-harmful-bias-and-homogenization'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-fairness', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-bias', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='Training an AI system on data with bias, such as historical or societal bias, can lead to biased or skewed outputs that can unfairly represent or otherwise discriminate against certain groups or individuals.'), Risk(id='atlas-data-transparency', name='Lack of training data transparency', description="Proper documentation contains information about how a model's data was collected, curated, and used to train a model, including any synthetic data generation processes. Without proper documentation it might be harder to satisfactorily explain the behavior of the model.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/data-transparency.html', dateCreated=datetime.date(2024, 3, 6), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['credo-risk-005'], related_mappings=['credo-risk-006', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-unintentional', 'mit-ai-causal-risk-timing-pre-deployment', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-information-integrity', 'nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='data-transparency', risk_type='training-data', phase=None, descriptor=['amplified by generative AI', 'amplified by synthetic data'], concern='A lack of data documentation limits the ability to evaluate risks associated with the data. Having access to the training data is not enough. Without recording how the data was cleaned, modified, or generated, including any data augmentation or synthetic data generation steps, the model behavior is more difficult to understand and to fix. Lack of data transparency also impacts model reuse as it is difficult to determine data representativeness for the new use without such documentation.'), Risk(id='atlas-lack-of-system-transparency', name='Lack of system transparency', description="Insufficient documentation of the system that uses the model and the model's purpose within the system in which it is used.", url='https://www.ibm.com/docs/en/watsonx/saas?topic=SSYOK8/wsj/ai-risk-atlas/lack-of-system-transparency.html', dateCreated=datetime.date(2024, 9, 24), dateModified=datetime.date(2025, 10, 22), exact_mappings=None, close_mappings=['eticas-poor-documentation', 'aiuc1-req-e017'], related_mappings=['asi09-human-agent-trust-exploitation', 'aiuc1-req-e017', 'mit-ai-causal-risk-entity-human', 'mit-ai-causal-risk-intent-other', 'mit-ai-causal-risk-timing-other', 'mit-ai-risk-subdomain-6.5'], narrow_mappings=None, broad_mappings=['nist-value-chain-and-component-integration'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-risk-atlas', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='ibm-risk-atlas-governance', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='lack-of-system-transparency', risk_type='non-technical', phase=None, descriptor=['traditional risk of AI'], concern="A lack of documentation makes it difficult to understand how the model's outcomes contribute to the system's or application's functionality.")]]
Risk Identification using NIST AI taxonomy¶
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="nist-ai-rmf",
)
risks
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:00<00:00, 1.04it/s]
[[Risk(id='nist-data-privacy', name='Data Privacy', description='Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-pii-leakage', 'eticas-privacy-confidentiality', 'eticas-unlawful-data-processing'], related_mappings=['ail-intellectual-property', 'ail-privacy', 'ail-specialized-advice', 'credo-risk-023', 'credo-risk-029', 'credo-risk-036', 'credo-risk-037', 'atlas-harmful-output'], narrow_mappings=None, broad_mappings=['atlas-attribute-inference-attack', 'atlas-data-privacy-rights', 'atlas-data-usage-rights', 'atlas-exposing-personal-information', 'atlas-ip-information-in-prompt', 'atlas-legal-accountability', 'atlas-membership-inference-attack', 'atlas-model-usage-rights', 'atlas-nonconsensual-use', 'atlas-personal-information-in-data', 'atlas-personal-information-in-prompt', 'atlas-reidentification'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='nist-ai-rmf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf=None, requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['GV-1.1-001', 'GV-1.2-001', 'GV-1.4-002', 'GV-1.6-003', 'GV-4.3-003', 'GV-6.1-001', 'GV-6.1-005', 'GV-6.1-009', 'GV-6.2-003', 'MP-1.1-001', 'MP-2.1-002', 'MP-4.1-001', 'MP-4.1-003', 'MP-4.1-004', 'MP-4.1-005', 'MP-4.1-008', 'MP-4.1-009', 'MP-4.1-010', 'MS-1.3-003', 'MS-2.2-002', 'MS-2.2-003', 'MS-2.2-004', 'MS-2.3-004', 'MS-2.6-002', 'MS-2.7-001', 'MG-2.2-009', 'MG-3.1-002', 'MG-3.2-002', 'MG-4.3-003'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='nist-harmful-bias-and-homogenization', name='Harmful Bias and Homogenization', description='Amplification and exacerbation of historical, societal, and systemic biases; performance disparities between sub-groups or languages, possibly due to non-representative training data, that result in discrimination, amplification of biases, or incorrect presumptions about performance; undesired homogeneity that skews system or model outputs, which may be erroneous, lead to ill-founded decision-making, or amplify harmful biases.', url=None, dateCreated=None, dateModified=None, exact_mappings=['eticas-homogenization-output-across-groups'], close_mappings=['eticas-bias-fairness'], related_mappings=['ail-defamation', 'ail-suicide-and-self-harm', 'credo-risk-010', 'credo-risk-011', 'credo-risk-012', 'credo-risk-022'], narrow_mappings=None, broad_mappings=['atlas-data-bias', 'atlas-decision-bias', 'atlas-impact-on-affected-communities', 'atlas-output-bias', 'atlas-spreading-toxicity', 'atlas-unrepresentative-data'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='nist-ai-rmf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf=None, requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='nist-information-integrity', name='Information Integrity', description='Lowered barrier to entry to generate and support the exchange and consumption of content which may not distinguish fact from opinion or fiction or acknowledge uncertainties, or could be leveraged for large-scale dis- and mis-information campaigns.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-synthetic-media-abuse'], related_mappings=['eticas-ai-interaction-disclosure', 'ail-defamation', 'ail-intellectual-property', 'ail-nonviolent-crimes', 'ail-privacy', 'ail-specialized-advice', 'credo-risk-007', 'credo-risk-022', 'credo-risk-032'], narrow_mappings=None, broad_mappings=['atlas-data-transparency', 'atlas-impact-on-cultural-diversity', 'atlas-impact-on-human-agency', 'atlas-incomplete-advice', 'atlas-jailbreaking', 'atlas-lack-of-testing-diversity', 'atlas-poor-model-accuracy', 'atlas-spreading-disinformation', 'atlas-unexplainable-output', 'atlas-untraceable-attribution'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='nist-ai-rmf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf=None, requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['GV-1.2-001', 'GV-1.3-001', 'GV-1.3-006', 'GV-1.3-007', 'GV-1.5-001', 'GV-1.5-003', 'GV-1.6-003', 'GV-4.3-001', 'GV-4.3-003', 'GV-6.1-003', 'GV-6.1-004', 'GV-6.1-005', 'GV-6.1-006', 'GV-6.1-008', 'GV-6.2-006', 'MP-2.1-001', 'MP-2.2-001', 'MP-2.2-002', 'MP-2.3-001', 'MP-2.3-003', 'MP-2.3-004', 'MP-3.4-001', 'MP-3.4-002', 'MP-3.4-003', 'MP-3.4-005', 'MP-3.4-006', 'MP-5.1-001', 'MP-5.1-002', 'MP-5.1-004', 'MS-1.1-001', 'MS-1.1-002', 'MS-1.1-003', 'MS-1.1-005', 'MS-1.1-007', 'MS-1.1-009', 'MS-2.2-001', 'MS-2.2-002', 'MS-2.2-003', 'MS-2.3-004', 'MS-2.5-005', 'MS-2.6-005', 'MS-2.7-001', 'MS-2.7-002', 'MS-2.7-003', 'MS-2.7-004', 'MS-2.7-005', 'MS-2.7-006', 'MS-2.7-008', 'MS-2.8-003', 'MS-2.9-002', 'MS-2.10-001', 'MS-2.10-002', 'MS-2.13-001', 'MS-3.3-002', 'MS-3.3-004', 'MS-3.3-005', 'MS-4.2-001', 'MS-4.2-003', 'MS-4.2-004', 'MG-2.2-002', 'MG-2.2-003', 'MG-2.2-007', 'MG-2.2-009', 'MG-3.1-005', 'MG-3.2-002', 'MG-3.2-003', 'MG-3.2-005', 'MG-3.2-006', 'MG-3.2-007', 'MG-4.1-001', 'MG-4.1-006', 'MG-4.3-002'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='nist-information-security', name='Information Security', description='Lowered barriers for offensive cyber capabilities, including via automated discovery and exploitation of vulnerabilities to ease hacking, malware, phishing, offensive cyber operations, or other cyberattacks; increased attack surface for targeted cyberattacks, which may compromise a system’s availability or the confidentiality or integrity of training data, code, or model weights.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-confidential-information-leakage', 'eticas-security-misuse'], related_mappings=['ail-nonviolent-crimes', 'ail-privacy', 'credo-risk-038', 'credo-risk-040', 'credo-risk-041'], narrow_mappings=['eticas-data-poisoning', 'eticas-model-extraction', 'eticas-prompt-injection'], broad_mappings=['atlas-attribute-inference-attack', 'atlas-data-contamination', 'atlas-data-poisoning', 'atlas-evasion-attack', 'atlas-extraction-attack', 'atlas-harmful-code-generation', 'atlas-prompt-injection', 'atlas-prompt-leaking', 'atlas-prompt-priming', 'atlas-unreliable-source-attribution'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='nist-ai-rmf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf=None, requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['GV-1.2-002', 'GV-1.3-003', 'GV-1.3-007', 'GV-1.5-002', 'GV-1.6-001', 'GV-1.7-001', 'GV-1.7-002', 'GV-2.1-004', 'GV-3.2-002', 'GV-3.2-005', 'GV-4.3-002', 'GV-6.1-004', 'GV-6.1-005', 'GV-6.1-009', 'GV-6.2-003', 'GV-6.2-007', 'MP-2.3-005', 'MP-4.1-003', 'MP-4.1-005', 'MP-5.1-001', 'MP-5.1-005', 'MP-5.1-006', 'MS-2.2-001', 'MS-2.2-002', 'MS-2.3-001', 'MS-2.3-002', 'MS-2.3-004', 'MS-2.5-006', 'MS-2.6-005', 'MS-2.6-006', 'MS-2.6-007', 'MS-2.7-001', 'MS-2.7-002', 'MS-2.7-004', 'MS-2.7-006', 'MS-2.7-007', 'MS-2.7-008', 'MS-2.7-009', 'MS-4.2-001', 'MS-4.2-002', 'MS-4.2-005', 'MG-1.3-001', 'MG-2.2-004', 'MG-2.4-002', 'MG-2.4-003', 'MG-2.4-004', 'MG-3.1-002', 'MG-3.1-005', 'MG-4.1-002', 'MG-4.3-001', 'MG-4.3-003'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None)]]
Risk Identification using MIT AI taxonomy¶
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="mit-ai-risk-repository",
)
risks
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:03<00:00, 3.15s/it]
[[Risk(id='mit-ai-risk-subdomain-1.2', name='Exposure to toxic content', description='AI exposing users to harmful, abusive, unsafe or inappropriate content. May involve AI creating, describing, providing advice, or encouraging action. Examples of toxic content include hate-speech, violence, extremism, illegal acts, child sexual abuse material, as well as content that violates community norms such as profanity, inflammatory political speech, or pornography.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-harmful-content-toxicity', 'credo-risk-013'], related_mappings=['ail-child-sexual-exploitation', 'ail-hate', 'ail-indiscriminate-weapons-cbrne', 'ail-intellectual-property', 'ail-sex-related-crimes', 'ail-sexual-content', 'ail-suicide-and-self-harm', 'ail-violent-crimes', 'credo-risk-014', 'credo-risk-015', 'atlas-harmful-output', 'atlas-toxic-output'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='mit-ai-risk-repository', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='mit-ai-risk-domain-1', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='mit-ai-risk-subdomain-3.1', name='False or misleading information', description='AI systems that inadvertently generate or spread incorrect or deceptive information, which can lead to inaccurate beliefs in users and undermine their autonomy. Humans that make decisions based on false beliefs can experience physical, emotional or material harms', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-hallucination', 'credo-risk-021'], related_mappings=['ail-defamation', 'credo-risk-017', 'atlas-hallucination'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='mit-ai-risk-repository', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='mit-ai-risk-domain-3', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='mit-ai-risk-subdomain-5.1', name='Overreliance and unsafe use', description='Users anthropomorphizing, trusting, or relying on AI systems, leading to emotional or material dependence and inappropriate relationships with or expectations of AI systems. Trust can be exploited by malicious actors (e.g., to harvest personal information or enable manipulation), or result in harm from inappropriate use of AI in critical situations (e.g., medical emergency). Overreliance on AI systems can compromise autonomy and weaken social ties.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['credo-risk-016'], related_mappings=['ail-nonviolent-crimes', 'ail-specialized-advice', 'credo-risk-020', 'credo-risk-034', 'atlas-improper-usage', 'atlas-over-or-under-reliance'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='mit-ai-risk-repository', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='mit-ai-risk-domain-5', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='mit-ai-risk-subdomain-7.4', name='Lack of transparency or interpretability', description='Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['eticas-stakeholder-communication', 'eticas-system-explainability', 'eticas-transparency-explainability', 'eticas-untraceable-agent-actions'], related_mappings=['credo-risk-005', 'credo-risk-006', 'credo-risk-007', 'credo-risk-008', 'credo-risk-009', 'credo-risk-017', 'credo-risk-028', 'credo-risk-033', 'atlas-inaccessible-training-data', 'atlas-lack-of-model-transparency', 'atlas-non-disclosure', 'atlas-unexplainable-output', 'atlas-unreliable-source-attribution', 'atlas-untraceable-attribution'], narrow_mappings=['eticas-right-to-explanation-contestation'], broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='mit-ai-risk-repository', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='mit-ai-risk-domain-7', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None)]]
Risk Identification using Granite Guardian taxonomy¶
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="ibm-granite-guardian",
)
risks
Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:00<00:00, 1.26it/s]
[[Risk(id='granite-guardian-harm', name='Harm', description='Content considered universally harmful. This is our general category, which should encompass a variety of risks including those not specifically addressed by the following categories: Social Bias, Profanity, Sexual Content, Unethical Behavior, Violence, Jailbreaking, Groundedness, Answer Relevance, Context Relevance.', url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-harmful-output', 'asi01-agent-goal-hijack', 'asi03-identity-and-privilege-abuse', 'asi04-agentic-supply-chain-vulnerabilities', 'asi07-insecure-inter-agent-communication', 'asi09-human-agent-trust-exploitation', 'asi10-rogue-agents', 'ail-child-sexual-exploitation', 'ail-defamation', 'ail-hate', 'ail-indiscriminate-weapons-cbrne', 'ail-intellectual-property', 'ail-nonviolent-crimes', 'ail-privacy', 'ail-sex-related-crimes', 'ail-sexual-content', 'ail-specialized-advice', 'ail-suicide-and-self-harm', 'ail-violent-crimes', 'credo-risk-003', 'credo-risk-004', 'credo-risk-008', 'credo-risk-009', 'credo-risk-010', 'credo-risk-011', 'credo-risk-012', 'credo-risk-013', 'credo-risk-014', 'credo-risk-015', 'credo-risk-016', 'credo-risk-017', 'credo-risk-018', 'credo-risk-021', 'credo-risk-023', 'credo-risk-024', 'credo-risk-025', 'credo-risk-026', 'credo-risk-028', 'credo-risk-029', 'credo-risk-033', 'credo-risk-034', 'credo-risk-036', 'credo-risk-037', 'credo-risk-038', 'credo-risk-040', 'credo-risk-041', 'credo-risk-043'], narrow_mappings=['granite-social-bias', 'granite-profanity', 'granite-sexual-content', 'granite-unethical-behavior', 'granite-violence', 'granite-jailbreak', 'granite-harm-engagement', 'granite-evasiveness'], broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-harm-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='harm', risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='granite-social-bias', name='Social Bias', description='Systemic prejudice against groups based on shared identity or characteristics, often stemming from stereotypes or cultural influences. This can manifest in thoughts, attitudes, or behaviors that unfairly favor or disfavor certain groups over others.', url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-output-bias', 'ail-specialized-advice', 'credo-risk-010', 'credo-risk-011'], narrow_mappings=None, broad_mappings=['granite-guardian-harm'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-harm-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=['gg-social-bias-detection'], isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='social_bias', risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='granite-profanity', name='Profanity', description='Use of language considered offensive or socially unacceptable in a given context. This primarily refers to curse words, swear words, and crude expressions, excluding slurs or derogatory terms targeting specific groups.', url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-toxic-output', 'ail-sexual-content', 'credo-risk-013'], narrow_mappings=None, broad_mappings=['granite-guardian-harm'], isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-harm-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=['gg-profanity-detection'], isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='profanity', risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='granite-groundedness', name='Groundedness', description='This risk arises in a Retrieval-Augmented Generation (RAG) system when the LLM response includes claims, facts, or details that are not supported by or directly contradicted by the given context. An ungrounded answer may involve fabricating information, misinterpreting the context, or making unsupported extrapolations beyond what the context actually states.', url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-hallucination', 'asi06-memory-and-context-poisoning', 'asi08-cascading-failures', 'asi09-human-agent-trust-exploitation', 'ail-specialized-advice', 'ail-suicide-and-self-harm', 'ail-violent-crimes'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-rag-safety-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=['gg-groundedness-detection'], isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='groundedness', risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='granite-relevance', name='Context Relevance', description="This occurs in when the retrieved or provided context fails to contain information pertinent to answering the user's question or addressing their needs. Irrelevant context may be on a different topic, from an unrelated domain, or contain information that doesn't help in formulating an appropriate response to the user.", url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-hallucination', 'asi06-memory-and-context-poisoning', 'ail-specialized-advice'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-rag-safety-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=['gg-relevance-detection'], isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='relevance', risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='granite-answer-relevance', name='Answer Relevance', description="This occurs when the LLM response fails to address or properly respond to the user's input. This includes providing off-topic information, misinterpreting the query, or omitting crucial details requested by the User. An irrelevant answer may contain factually correct information but still fail to meet the User's specific needs or answer their intended question.", url='https://www.ibm.com/granite/docs/models/guardian/#risk-definitions', dateCreated=datetime.date(2024, 12, 10), dateModified=datetime.date(2024, 12, 10), exact_mappings=None, close_mappings=None, related_mappings=['atlas-hallucination', 'asi08-cascading-failures', 'ail-specialized-advice', 'ail-suicide-and-self-harm'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='ibm-granite-guardian', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='granite-guardian-rag-safety-group', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=['gg-answer-relevance-detection'], isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag='answer-relevance', risk_type=None, phase=None, descriptor=None, concern=None)]]
Risk Identification using Credo Unified Control Framework taxonomy¶
usecase = "Generate personalized, relevant responses, recommendations, and summaries of claims for customers to support agents to enhance their interactions with customers."
risks = ai_atlas_nexus.identify_risks_from_usecases(
usecases=[usecase],
inference_engine=inference_engine,
taxonomy="credo-ucf",
)
risks
[2026-08-21 11:04:12:511] - WARNING - AIAtlasNexus - <RAN47275F12W> Chain of Thought (CoT) examples were not provided, or do not exist in the master for this taxonomy. The API will use the Zero shot method. To improve the accuracy of risk identification, please provide CoT examples in `cot_examples` when calling this API. You may also consider raising an issue to permanently add these examples to the AI Atlas Nexus master. Inferring with WML, backend - DEFAULT: 100%|██████████| 1/1 [00:04<00:00, 4.06s/it]
[[Risk(id='credo-risk-005', name='Lack of training data transparency (IBM, 2024)', description="Without accurate documentation on how a model's data was collected, curated, and used to train a model, it may be harder to satisfactorily explain the behavior of the model with respect to the data. Data provenance issues may also increase legal risks (e.g., intellectual property infringement).", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['atlas-data-transparency'], related_mappings=['ail-intellectual-property', 'ail-specialized-advice', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-explainability-&-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-009'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-006', name='Lack of inference data transparency', description='Lack of inference data transparency: Insufficient visibility into data sources used during model inference', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['atlas-data-transparency', 'atlas-lack-of-data-transparency', 'mit-ai-risk-subdomain-7.4'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-explainability-&-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-010', 'credo-act-control-011'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-007', name='Inadequate observability (Slatteryet al., 2024)', description='The AI system may lack sufficient logging or traceability features, making it difficult to monitor or audit its decision-making process after the fact.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['ail-suicide-and-self-harm', 'atlas-unreliable-source-attribution', 'mit-ai-risk-subdomain-7.3', 'mit-ai-risk-subdomain-7.4', 'nist-information-integrity'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-explainability-&-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-010'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-008', name='Opaque system architecture', description="The AI system's internal structure and decision-making process may not be understandable or accessible to stakeholders, including developers, auditors, or end-users.", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['granite-guardian-harm', 'atlas-lack-of-data-transparency', 'mit-ai-risk-subdomain-7.4', 'nist-human-ai-configuration'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-explainability-&-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-009', name='Black box decisionmaking (Slattery et al., 2024; IBM, 2024)', description="The AI system's decision-making process may be opaque, even when the architecture is known, making it difficult to understand how the system arrives at its outputs or recommendations.", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['ail-suicide-and-self-harm', 'granite-guardian-harm', 'mit-ai-risk-subdomain-7.3', 'mit-ai-risk-subdomain-7.4', 'nist-human-ai-configuration'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-explainability-&-transparency', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-011', 'credo-act-control-037'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-010', name='Stereotype perpetuation (Slattery et al., 2024; IBM, 2024)', description="The AI system's outputs may explicitly reflect or reinforce harmful stereotypes, prejudices, or biased characterizations of specific groups. The AI system may exhibit unjustified or harmful differences in accuracy, quality, or outcomes across demographic groups, potentially leading to unfair treatment and discrimination. This includes both disparate error rates that affect opportunity and", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['ail-suicide-and-self-harm', 'ail-hate', 'granite-guardian-harm', 'granite-social-bias', 'atlas-impact-on-cultural-diversity', 'atlas-output-bias', 'atlas-unrepresentative-data', 'mit-ai-risk-subdomain-1.1', 'nist-harmful-bias-and-homogenization'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-fairness-&-bias', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-014', 'credo-act-control-015', 'credo-act-control-016', 'credo-act-control-028'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-011', name='Disparate model performance (Slattery et al., 2024; IBM, 2024)', description='The AI system may exhibit unjustified or harmful differences in accuracy, quality, or outcomes across demographic groups, potentially leading to unfair treatment and discrimination. This includes both disparate error rates that affect opportunity and disparate outcome rates that affect group-level results.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['granite-guardian-harm', 'granite-social-bias', 'atlas-decision-bias', 'atlas-harmful-output', 'atlas-output-bias', 'mit-ai-risk-subdomain-1.1', 'nist-harmful-bias-and-homogenization'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-fairness-&-bias', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-012', name='Unequal access to AI benefits', description="The AI system's benefits may not be equally accessible to all users, potentially resulting in reduced advantages for those with limited access. Accessibility may be affected by physical abilities, cognitive abilities, language, or technological access.", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['granite-guardian-harm', 'atlas-dangerous-use', 'mit-ai-risk-subdomain-6.5', 'nist-harmful-bias-and-homogenization', 'nist-human-ai-configuration'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-fairness-&-bias', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-014', 'credo-act-control-015', 'credo-act-control-016', 'credo-act-control-028'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-013', name='Toxic content (Slattery et al., 2024; IBM, 2024)', description='The AI system may generate or respond with hateful content, such as racist, sexist, or otherwise offensive material.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['mit-ai-risk-subdomain-1.2'], related_mappings=['ail-hate', 'ail-sexual-content', 'granite-guardian-harm', 'granite-profanity', 'granite-sexual-content', 'granite-violence', 'atlas-human-exploitation', 'atlas-spreading-toxicity', 'nist-dangerous-violent-or-hateful-content', 'nist-obscene-degrading-and-or-abusive-content'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-harmful-content', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-017', 'credo-act-control-018', 'credo-act-control-019'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-016', name='Over or under-reliance and unsafe use (Slattery et al., 2024; IBM, 2024; AI, 2023)', description='Users may inappropriately rely on the AI system for critical decisions or tasks beyond its capabilities, or fail to put trust in AI systems when they should, potentially leading to errors or safety issues.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=['atlas-over-or-under-reliance', 'mit-ai-risk-subdomain-5.1'], related_mappings=['granite-guardian-harm', 'nist-human-ai-configuration'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-human-ai-interaction', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-009', 'credo-act-control-011', 'credo-act-control-028', 'credo-act-control-029', 'credo-act-control-029'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-017', name='Inadequate AI literacy and communication', description="The AI system's capabilities, limitations, and appropriate use cases may be insufficiently understood or communicated within the organization, potentially resulting in ineffective implementation or failure to achieve desired outcomes.", url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['granite-guardian-harm', 'mit-ai-risk-subdomain-3.1', 'mit-ai-risk-subdomain-7.4', 'nist-human-ai-configuration', 'llm062025-excessive-agency', 'llm092025-misinformation'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-human-ai-interaction', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=['credo-act-control-009', 'credo-act-control-025'], detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None), Risk(id='credo-risk-041', name='Vulnerability to adversarial attacks (Slattery et al., 2024; IBM, 2024; AI, 2023)', description='The AI system may be vulnerable to adversarial attacks, including prompt-based attacks, which may induce the model to behave outside of its intended functionality.', url=None, dateCreated=None, dateModified=None, exact_mappings=None, close_mappings=None, related_mappings=['granite-guardian-harm', 'granite-jailbreak', 'atlas-evasion-attack', 'mit-ai-risk-subdomain-2.2', 'mit-ai-risk-subdomain-7.3', 'nist-information-security', 'llm042025-data-and-model-poisoning', 'llm072025-system-prompt-leakage'], narrow_mappings=None, broad_mappings=None, isCategorizedAs=None, hasLifecycleStatus=None, notes=None, isDefinedByTaxonomy='credo-ucf', isDefinedByVocabulary=None, hasDocumentation=None, hasExternalReference=None, isPartOf='credo-rg-security', requiredByTask=None, requiresCapability=None, implementedByAdapter=None, hasRule=None, type='Risk', hasJurisdiction=None, isDetectedBy=None, isMitigatedBy=None, isUsedWithinLocality=None, hasRelatedAction=None, detectsRiskConcept=None, tag=None, risk_type=None, phase=None, descriptor=None, concern=None)]]
Evaluation¶
We perform an evaluation of the performance of risk-identification using LLMs in the wild. Using the MIT risk taxonomy and AI uses cases sourced from IBM, human annotators labeled usecase-risk pairs. Disagreements were resolved using majority votes. Five models were tested: Llama 3.3 70B, Granite 3.3 8B, Llama 3.1 8B, Qwen3 8B, and GPT-OSS 20B. Prompts were automatically optimized using DSPy's MIPROv2 optimizer, which searches for better instructions rather than requiring manual prompt engineering. The results showed that the larger Llama 3.3 70B was already near its ceiling at 77% accuracy with no improvement from optimization, while the smaller models benefited substantially — Granite 3.3 8B jumped from 61% to 82%, and Llama 3.1 8B improved from 60% to 76% — suggesting that automated prompt optimization can meaningfully close the gap between small and large models on specialized risk classification tasks.
%pip install matplotlib
import matplotlib.pyplot as plt
import pandas as pd
results = [
{
"model": "meta-llama/llama-3-3-70b-instruct",
"baseline": 0.771,
"optimized": 0.771,
},
{
"model": "ibm-granite/granite-3.3-8b-instruct2",
"baseline": 0.61,
"optimized": 0.819,
},
{"model": "meta-llama/Llama-3.1-8B-Instruct", "baseline": 0.60, "optimized": 0.762},
{"model": "Qwen/Qwen3-8B", "baseline": 0.657, "optimized": 0.752},
{"model": "openai/gpt-oss-20b", "baseline": 0.705, "optimized": 0.762},
]
# build a dataframe from the results list we got above
res_df = pd.DataFrame(results)
# plot baseline vs optimized accuracy for each model
ax = res_df.plot(x="model", y=["baseline", "optimized"], kind="barh", figsize=(8, 4))
ax.grid(axis="x")
ax.set_xlabel("Correctness relative to majority human label (n=105)")
ax.legend(title="prompts", loc="lower right")
ax.set_xlim(0.5, 1.0)
plt.tight_layout()
plt.show()
Requirement already satisfied: matplotlib in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (3.11.1) Requirement already satisfied: contourpy>=1.0.1 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (1.3.3) Requirement already satisfied: cycler>=0.10 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (0.12.1) Requirement already satisfied: fonttools>=4.28.2 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (4.63.0) Requirement already satisfied: kiwisolver>=1.3.1 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (1.5.0) Requirement already satisfied: numpy>=1.25 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (2.4.4) Requirement already satisfied: packaging>=20.0 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (26.0) Requirement already satisfied: pillow>=9 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (12.2.0) Requirement already satisfied: pyparsing>=3 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (3.3.2) Requirement already satisfied: python-dateutil>=2.7 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from matplotlib) (2.9.0.post0) Requirement already satisfied: six>=1.5 in /Users/ingevejs/Documents/workspace/ingelise/risk-atlas-nexus/v-ai-atlas-nexus/lib/python3.12/site-packages (from python-dateutil>=2.7->matplotlib) (1.17.0) [notice] A new release of pip is available: 26.0 -> 26.2.1 [notice] To update, run: pip install --upgrade pip Note: you may need to restart the kernel to use updated packages.