bdi-mental-states

Pass

Audited by Gen Agent Trust Hub on Sep 14, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest untrusted external RDF context and translate it into agent mental states, creating a surface for potential data-driven manipulation of agent reasoning.
  • Ingestion points: The skill primarily ingests external RDF data as world state configurations (SKILL.md) and uses Logic Augmented Generation (LAG) to generate mental states from context (references/framework-integration.md).
  • Boundary markers: The LAG implementation in references/framework-integration.md employs structured prompt templates with clear delimiters (e.g., ## BDI Ontology, ## Context to Model) to differentiate instructions from untrusted data.
  • Capability inventory: The skill utilizes RDF parsing, SPARQL querying, and the generation of logic-based production rules (SEMAS).
  • Sanitization: The skill implements comprehensive sanitization and validation mechanisms, including the _validate_against_ontology function in references/framework-integration.md and a suite of 27 SPARQL competency queries in references/sparql-competency.md designed to verify the causal and temporal integrity of the agent's mental model.
  • [SAFE]: The skill follows security best practices by emphasizing explainability and provenance tracking. All external URLs point to official standard bodies (W3C, FIPA, ODP), and referenced software libraries (rdflib, fipa_acl) are established packages in the semantic web and agent communities. The code snippets are instructional templates and do not execute arbitrary or dangerous commands.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 14, 2026, 05:02 PM
Security Audit — agent-trust-hub — bdi-mental-states