sample-aiml-security-assessment

AGENTS.md

This file provides guidance to coding agents when working with code in this repository.

What this is

A serverless framework that scans AWS accounts for AI/ML security misconfigurations and produces interactive HTML reports. The full catalog contains 208 checks across seven assessment areas: 94 core checks (40 Amazon Bedrock, 29 Amazon SageMaker AI with SM-29 reserved, 17 Amazon Bedrock AgentCore, and 8 AWS Agent Registry), up to 38 Agentic AI Security checks, 64 optional Responsible AI GRC checks, and 12 optional OWASP Top 10 for LLM checks. Checks are derived from the AWS Well-Architected Generative AI Lens, the Agentic AI Lens, AWS Responsible AI GRC guidance, and the OWASP Top 10 for LLM 2025.

Commands

All Python tooling runs from the repo-local .venv/ — pytest, ruff, cfn-lint, and pip live under .venv/bin/. Do not use system python/python3 or pip install globally; the version pins in tests/requirements.txt and each Lambda’s requirements.txt are only reproducible inside the venv. sam is system-installed (Homebrew).

Everything under tests/ runs as one pytest session. Multiple assessment packages each have a top-level app.py, so each test module loads its package through importlib.util.spec_from_file_location under a distinct module name (bedrock_app, agentcore_app, agent_registry_app, …) rather than a bare import app. Follow that pattern for a new package’s tests; a bare import app after sys.path.insert collides with whichever package was imported first and produces confusing cross-package failures.

The suites that need their own session are the ones that live outside tests/: responsible_ai_grc_tests/ (its own conftest.py) and generate_consolidated_report/test_generate_report.py (runs from its own directory).

# Bootstrap / refresh venv (run once, or after requirements files change)
.venv/bin/pip install -r tests/requirements.txt \
  -r aiml-security-assessment/functions/security/agentcore_assessments/requirements.txt \
  -r aiml-security-assessment/functions/security/agent_registry_assessments/requirements.txt \
  -r aiml-security-assessment/functions/security/bedrock_assessments/requirements.txt \
  -r aiml-security-assessment/functions/security/cleanup_bucket/requirements.txt \
  -r aiml-security-assessment/functions/security/responsible_ai_grc_assessments/requirements.txt \
  -r aiml-security-assessment/functions/security/generate_consolidated_report/requirements.txt \
  -r aiml-security-assessment/functions/security/iam_permission_caching/requirements.txt \
  -r aiml-security-assessment/functions/security/owasp_assessments/requirements.txt \
  -r aiml-security-assessment/functions/security/resolve_regions/requirements.txt \
  -r aiml-security-assessment/functions/security/sagemaker_assessments/requirements.txt
.venv/bin/pip check

# Verify tooling is picked up from the venv (not the system Python)
.venv/bin/python --version         # 3.12.x matches the Lambda runtime and CI
.venv/bin/pip check                 # dependency sanity — no conflicts
which -a python pytest ruff cfn-lint  # the venv paths should win when PATH-activated

# Required env vars for any test run (see tests/conftest.py for the autouse fixture)
export AIML_ASSESSMENT_BUCKET_NAME=test-assessment-bucket
export AWS_DEFAULT_REGION=us-east-1
export AWS_ACCESS_KEY_ID=testing
export AWS_SECRET_ACCESS_KEY=testing

# Core suite — one session covers every package under tests/
.venv/bin/python -m pytest tests/ -v --tb=short

# assessment_history package — keep 100% line and branch coverage
.venv/bin/python -m pytest tests/test_assessment_history_*.py --cov=assessment_history --cov-branch --cov-report=term-missing --cov-fail-under=100

# Responsible AI GRC suite — separate session (lives outside tests/, own conftest)
.venv/bin/python -m pytest aiml-security-assessment/functions/security/responsible_ai_grc_tests/ -v --tb=short

# Report-pipeline tests live next to the report code and run from that dir
(cd aiml-security-assessment/functions/security/generate_consolidated_report \
  && ../../../../.venv/bin/python -m pytest test_generate_report.py -v --tb=short)

# Single test
.venv/bin/python -m pytest tests/test_bedrock_checks.py::test_name -v

# Lint / format (CI gate — Ruff automatically loads the repository's ruff.toml).
# CI (.github/workflows/python-lint.yml) runs ruff over the PR's *changed*
# .py files, not a fixed directory. Match that scope locally — otherwise a
# ruff diff in tests/ or consolidate_html_reports.py passes the security/
# subtree scan and still fails CI (this bit us on PR #53).
changed_py=$(git diff --name-only --diff-filter=ACMR origin/main...HEAD -- '*.py')
.venv/bin/ruff check         $changed_py
.venv/bin/ruff format --check $changed_py
# Fallback / belt-and-braces — check the whole repo before pushing:
.venv/bin/ruff check .
.venv/bin/ruff format --check .

# Template validation
.venv/bin/cfn-lint deployment/*.yaml aiml-security-assessment/template.yaml aiml-security-assessment/template-multi-account.yaml
(cd aiml-security-assessment && sam validate --template template.yaml --lint && sam build --template template.yaml)

# Invoke one assessment Lambda locally (sam build requires Python 3.12 on PATH for the target runtime)
(cd aiml-security-assessment && sam build --template template.yaml && sam local invoke BedrockSecurityAssessmentFunction --event testfile.json)

Prefix every command with .venv/bin/ explicitly (rather than relying on source .venv/bin/activate) so a stale shell activation cannot silently reach the system interpreter — that mistake is why previous runs failed with No module named pytest.

Architecture

Two-phase, two-mode. Phase 1 is CloudFormation deployment of roles + central infra; phase 2 is CodeBuild (buildspec.yml) orchestrating per-account SAM deploys and Step Functions executions. The same code runs in single-account mode (one account, deployed via template.yaml) and multi-account mode (Organizations-wide, template-multi-account.yaml + deployment/2-aiml-security-codebuild.yaml assuming AIMLSecurityMemberRole cross-account).

Service selection: Four Enable*Assessment parameters (Bedrock, SageMaker, AgentCore, AgentRegistry; default true) are substituted into the initial Configure Service Assessments Pass state. Its ServiceSelection map gates direct service branches, OWASP source reads, and report artifact requirements. Both CodeBuild deployment paths forward these switches and use the same selection during collection and consolidation. Deselected services are shown as Not selected, never N/A or clean. Optional GRC/OWASP scans remain independent; see docs/DEVELOPER_GUIDE.md → Service Selection.

Step Functions workflow (aiml-security-assessment/statemachine/assessments.asl.json): Cleanup S3 → IAM Permission Caching (global, once) → Resolve Regions → Map over regions (MaxRegionConcurrency) → Bedrock / SageMaker / AgentCore / AWS Agent Registry plus conditional Responsible AI GRC → conditional OWASP → Generate Consolidated Report. Responsible AI GRC runs only at RegionIndex == 0 when enableResponsibleAIGRC == "true" or enableOWASP == "true". Direct execution input using legacy "enableFinServ": "true" is rejected; the legacy CloudFormation parameter remains supported through CodeBuild alias resolution.

Optional assessment branches must preserve execution input. A Step Functions Task or Pass that produces a result replaces its input with that result when ResultPath is omitted. When inserting a branch after an existing assessment, set ResultPath on result-producing states in success and skipped routes, and on error handlers, so downstream states retain the selection flags, Execution, Region, and other required fields. A Catch-level ResultPath protects only the error path; it does not protect successful transitions. Test each transition through the next Choice and report-generation state. See the AWS ResultPath documentation.

Each direct service Lambda (functions/security/{bedrock,sagemaker,agentcore,agent_registry}_assessments/app.py) probes service availability, reads cached IAM permissions from S3 where needed, creates regional boto3 clients with explicit region_name, runs checks, and writes a region-suffixed CSV. Agent Registry receives the Step Functions Execution object and must derive the shared artifact/cache key from Execution.Name, writing agent_registry_security_report_<execution_id>_<region>.csv. Pass region= to every create_finding() call. responsible_ai_grc_assessments/app.py runs once and writes responsible_ai_grc_security_report_<execution_id>.csv. owasp_assessments/app.py reads the BR/SM/AC regional CSVs plus that unsuffixed Responsible AI GRC CSV on the first region, maps OW-01..OW-10, runs native OW-11/OW-12 checks, and writes region-suffixed OWASP CSVs. Agent Registry rows are intentionally excluded from OWASP source ingestion because the current AR controls do not prove an LLM01–LLM10 control. When OWASP is enabled by itself, FS-* source rows are hidden from the report UI.

Findings are CSV rows with a shared base schema produced by create_finding() in each package’s schema.py: Check_ID, Finding, Finding_Details, Resolution, Reference, Severity, Status, Region. Responsible AI GRC extends this with Compliance_Frameworks; the report layer ignores unknown extra columns but downstream CSV consumers may depend on them. The report layer parses CSVs back into findings, so the base column set and Check_ID prefix are a contract the report depends on.

Conventions that span files

Status / error semantics

Review checklist (run before committing changes to checks or IAM)

When reviewing a diff that touches assessment code or policies, verify, in order:

  1. API names and response semantics — every boto3 op and paginator exists on the service its client was constructed with (check the boto3.client("…") string, not the variable name); validate against botocore service models. Verify the exact field, enum values, resource scope, and successful empty-response behavior behind each Passed/Failed predicate.
  2. IAM and artifact coverage — every new assessment op is granted in both SAM templates on the correct per-Lambda policy, and is absent from the deployment/member roles. For any newly added check, validate the required IAM actions against the AWS Knowledge MCP server (mcp__AWS_Knowledge__aws___search_documentation / aws___read_documentation) rather than relying on memory, and update both SAM templates in the same commit if any grant is missing. Check S3 list/read grants for every new artifact consumer and write grants for its producer. Review top-level deployment policies separately only for deployment-time API or artifact-access changes.
  3. Status semantics — access-denied/region-unsupported → N/A (via is_region_unsupported()/describe_api_error()); tooling conditions → N/A/Informational, never Failed; no Failed + “No action required”; indeterminate ≠ absent in finding text.
  4. Inventory and per-resource errors — no truncating list calls; verify non-compliant resources on later pages; isolate per-resource detail calls; distinguish empty, denied, and partial inventories; avoid duplicate account-global findings across regions.
  5. AG- numbering — no collisions across the Bedrock, AgentCore, and Agent Registry mapping dictionaries and native AG-24..27; the current catalog is AG-01..38, and every mapped source ID must still exist.
  6. OWASP/compliance mapping and report structure — every OWASP_CHECK_MAPPINGS source ID still exists in its scanner, every source CSV the OWASP Lambda reads actually feeds at least one mapping, every emitted OW- ID is documented in docs/SECURITY_CHECKS_OWASP.md, and COMPLIANCE_STANDARDS prefixes in report_template.py match the report/consolidator routing expectations. Values in mappings are lists (one source check can emit multiple OW rows). An N/A source must produce Severity=Informational on the OW row (never inherit High/Medium). For a new standard, check that its registry slug drives the shared sidebar, filter, card, section, scope chip, and source link without breaking the Assessment Scope insertion.
  7. CSV schema drift — base columns remain present everywhere; Responsible AI GRC’s Compliance_Frameworks remains intentional and documented; report parsing continues to tolerate the extra column.
  8. Test coverage — every new check needs compliant/pass, non-compliant/fail (or advisory N/A), no-resource, and access-denied/API-unavailable coverage. Use SDK-shaped responses for both positive and negative cases, including a successful empty response where the API can return one; an empty-inventory-only test does not exercise the check predicate. Shared inventories need later-page, list-error, per-resource detail-error, and two-region tests where applicable. Policy parsers need both accepted-boundary and false-positive/false-negative regression cases. Assert the exact Check_ID emitted on fallback and error paths.
  9. Identifier hygiene — no real AWS account IDs, organization IDs, ARNs, or resource identifiers appear in tests, fixtures, examples, snapshots, or generated reports; all newly introduced values are clearly synthetic.
  10. Documentation completeness — for every catalog change, review AGENTS.md, README.md, docs/DEVELOPER_GUIDE.md, docs/SECURITY_CHECKS.md, applicable docs/SECURITY_CHECKS_*.md catalogs, docs/TROUBLESHOOTING.md, CHANGELOG.md, and any relevant scope, migration, severity, sample-report, or diagram documentation. Compare each catalog claim and remediation with the actual API evidence and check predicate. Check count locations, headings/TOC/deep links, service and compliance-standard routing, report artifacts, deployment parameters, and user-facing troubleshooting. Update every affected document in the same change. Responsible AI GRC provenance: when changing responsible_ai_grc_assessments/app.py or docs/SECURITY_CHECKS_RESPONSIBLE_AI_GRC.md, run .venv/bin/python generate_provenance.py --check; if stale, run .venv/bin/python generate_provenance.py and commit the regenerated responsible_ai_grc_assessments/provenance.json. Do not edit that generated file manually.
  11. Changelog and deployment impact — every releasable behavior or deployment change is recorded under CHANGELOG.md → Unreleased; the deployment-impact entry matches the actual changed templates and required update order.
  12. Gates — ruff check and ruff format --check clean over the PR’s changed .py files, cfn-lint clean on edited templates, and each relevant pytest session passes. Tests assert exact finding text, so reword carefully.
  13. Optional assessment end to end — trace each enable parameter through both deployment modes, CodeBuild and direct SAM execution, every Step Functions success/skip/catch transition, the CSV key and schema, report Lambda S3 permissions, selected-artifact validation, and rendered single- and multi-account reports. Verify that a declared CloudFormation parameter is actually referenced by the state machine, not merely forwarded by CodeBuild. Test enabled and disabled paths with downstream state input intact; generate a report from a representative synthetic CSV artifact and verify that a required artifact missing or unreadable is visible as incomplete rather than a clean result.

Docs to consult