This repository includes instructions for developers using AI coding agents:
.venv/ toolchain, separate pytest sessions, architecture and
schema contracts, status semantics, pagination requirements, IAM coverage
in both SAM runtime templates, deployment-role separation, mapping checks,
identifier hygiene, and the pre-commit review gates.AGENTS.md so the
instructions have a single source of truth.AI-assisted development should begin by loading the repository-root
AGENTS.md, then use this guide for implementation workflows and architecture
details. Human reviewers should evaluate agent-generated changes against the
same instructions. Update AGENTS.md when repository-wide agent guidance
changes; keep CLAUDE.md as the delegation shim unless a tool requires
additional compatibility syntax.
The AI/ML Security Assessment Framework is a serverless, multi-account security assessment solution for AWS AI/ML workloads. It performs 94 core security checks across Amazon Bedrock, Amazon SageMaker AI, Amazon Bedrock AgentCore, and AWS Agent Registry, plus 38 always-on Agentic AI Security checks, with optional 64-check Responsible AI GRC and 12-check OWASP Top 10 for LLM assessments, generating interactive HTML reports with findings and remediation guidance.
The current deployment is validated only in the standard AWS commercial
partition (aws). Partition-aware implementation details must not be treated
as support for AWS GovCloud (US) (aws-us-gov) or AWS China (aws-cn).
Adding either partition requires service-availability review, partition-safe
Step Functions integrations and generated URLs, and end-to-end validation of
both single-account and multi-account deployment modes.



1-aiml-security-member-roles.yaml)AIMLSecurityMemberRole to all target accounts2-aiml-security-codebuild.yaml)MultiAccountCodeBuildRole with cross-account access permissionsMultiAccountListOverrideAIMLSecurityMemberRole in each target accountenableResponsibleAIGRC and enableOWASP from the deployment parameterssample-aiml-security-assessment/
├── AGENTS.md # Canonical AI coding-agent guidance
├── CLAUDE.md # Compatibility shim that loads AGENTS.md
├── aiml-security-assessment/
│ ├── functions/security/
│ │ ├── bedrock_assessments/ # Bedrock security checks (40)
│ │ ├── sagemaker_assessments/ # SageMaker checks (29; SM-29 reserved)
│ │ ├── agentcore_assessments/ # AgentCore security checks (17)
│ │ ├── agent_registry_assessments/ # AWS Agent Registry checks (8)
│ │ ├── responsible_ai_grc_assessments/ # Optional Responsible AI GRC checks (64)
│ │ ├── owasp_assessments/ # Optional OWASP Top 10 for LLM checks (12)
│ │ ├── responsible_ai_grc_tests/ # Responsible AI GRC-specific unit and coverage tests
│ │ ├── iam_permission_caching/ # AWS IAM permissions cache
│ │ ├── cleanup_bucket/ # Amazon S3 cleanup
│ │ ├── resolve_regions/ # Multi-region resolution Lambda
│ │ └── generate_consolidated_report/ # HTML/CSV report generation
│ ├── statemachine/ # AWS Step Functions definition
│ ├── images/ # SAM application images
│ ├── template.yaml # AWS SAM template (single-account)
│ ├── template-multi-account.yaml # AWS SAM template (multi-account)
│ ├── samconfig.toml # SAM deployment configuration
│ ├── envvars.json # Environment variables for local testing
│ └── testfile.json # Test event file for local invocation
├── assessment_history/ # Changes-since-last-assessment report (CodeBuild post-build)
├── deployment/ # AWS CloudFormation templates
├── docs/ # Documentation
│ ├── DEVELOPER_GUIDE.md # This guide
│ ├── ASSESSMENT_HISTORY.md # Changes since last assessment report
│ ├── SECURITY_CHECKS.md # Security checks reference (core + Agentic)
│ ├── SECURITY_CHECKS_RESPONSIBLE_AI_GRC.md # Responsible AI GRC checks reference
│ ├── SECURITY_CHECKS_OWASP.md # OWASP Top 10 for LLM checks reference
│ ├── SECURITY_CHECKS_RESPONSIBLE_AI_GRC_SEVERITY_METHODOLOGY.md # Severity model
│ ├── SECURITY_CHECKS_RESPONSIBLE_AI_GRC_SEVERITY_REGISTER.md # Per-finding severities
│ ├── TROUBLESHOOTING.md # Troubleshooting guide
│ ├── CLEANUP.md # Resource removal guide
│ ├── diagrams/ # Architecture diagrams
│ └── icons/ # AWS service icons
├── sample-reports/ # Sample assessment reports
│ ├── scripts/ # Screenshot capture and changes-sample scripts
│ ├── *.html # Sample HTML reports
│ └── *.png # Report screenshots
├── tests/ # Unit tests for assessment functions
│ └── requirements.txt # Test dependencies
├── .github/workflows/ # PR lint, test, SAM validate, and security scans
├── buildspec.yml # AWS CodeBuild orchestration
└── consolidate_html_reports.py # Multi-account report consolidation
# buildspec.yml execution flow
1. Get active accounts from AWS Organizations
2. For each account:
- Assume AIMLSecurityMemberRole
- Deploy AI/ML assessment stack through AWS SAM
- Start AWS Step Functions execution
3. Wait for completion and consolidate results
{
"Comment": "AI/ML Assessment Module",
"StartAt": "Cleanup S3 Bucket",
"States": {
"Cleanup S3 Bucket": {
"Type": "Task",
"Next": "IAM Permission Caching"
},
"IAM Permission Caching": {
"Type": "Task",
"Next": "Resolve Target Regions"
},
"Resolve Target Regions": {
"Type": "Task",
"Comment": "Resolves target regions from TARGET_REGIONS env var",
"Next": "Scan Regions"
},
"Scan Regions": {
"Type": "Map",
"ItemsPath": "$.ResolvedRegions.regions",
"MaxConcurrency": ${MaxRegionConcurrency},
"ItemProcessor": {
"ProcessorConfig": {"Mode": "INLINE"},
"StartAt": "Run Security Assessments",
"States": {
"Run Security Assessments": {
"Type": "Parallel",
"Branches": [
{"StartAt": "Bedrock Security Assessment", "States": {...}},
{"StartAt": "Sagemaker Security Assessment", "States": {...}},
{"StartAt": "AgentCore Security Assessment", "States": {...}},
{"StartAt": "AWS Agent Registry Security Assessment", "States": {...}},
{
"StartAt": "Responsible AI GRC Enabled?",
"States": {
"Responsible AI GRC Enabled?": {
"Type": "Choice",
"Comment": "Runs Responsible AI GRC when enableResponsibleAIGRC or enableOWASP is true and RegionIndex is 0"
},
"Responsible AI GRC Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true},
"Responsible AI GRC Assessment Skipped": {"Type": "Pass", "End": true}
}
}
],
"Next": "OWASP Enabled?"
},
"OWASP Enabled?": {
"Type": "Choice",
"Choices": [
{
"Variable": "$.OriginalInput.enableOWASP",
"StringEquals": "true",
"Next": "OWASP Security Assessment"
}
],
"Default": "OWASP Assessment Skipped"
},
"OWASP Security Assessment": {
"Type": "Task",
"Resource": "arn:aws:states:::lambda:invoke",
"End": true
},
"OWASP Assessment Skipped": {
"Type": "Pass",
"End": true
}
}
},
"Next": "Generate Consolidated Report"
},
"Generate Consolidated Report": {
"Type": "Task",
"End": true
}
}
}
EnableBedrockAssessment, EnableSageMakerAssessment,
EnableAgentCoreAssessment, and EnableAgentRegistryAssessment are string
parameters accepting true or false, defaulting to true. Both SAM templates
substitute these values into the initial Configure Service Assessments Pass
state. It writes $.ServiceSelection before cleanup and region resolution; the
regional Choice gates read that map through $.OriginalInput.ServiceSelection.
This keeps direct SAM executions consistent with CodeBuild deployments and
prevents execution input from overriding a deployment’s selected scope.
Both top-level deployment templates expose the switches as ENABLE_BEDROCK,
ENABLE_SAGEMAKER, ENABLE_AGENTCORE, and ENABLE_AGENT_REGISTRY CodeBuild
variables. The root buildspec.yml validates and forwards them to all three SAM
deployment sites (member, management, and single account). Its artifact checks
require CSVs only for enabled services. No change to the member-role StackSet
is needed; this feature adds no API calls or IAM permissions.
The report Lambda validates every selected service’s CSV for every resolved region, but does not require deselected artifacts. Missing selected artifacts still fail report generation. Both report modes label deselected areas Not selected, exclude their rows, and explain reduced lens coverage. An all-disabled selection may produce an HTML report without service CSVs.
Agentic AI rows come only from selected source assessments (Bedrock, AgentCore, and Agent Registry). OWASP receives the same selection map and skips reads of deselected source CSVs, while preserving missing-artifact notices for selected sources. Its native checks and Responsible AI GRC dependency still run when OWASP is enabled. Responsible AI GRC remains independently controlled and can scan services omitted from the direct-service selection. Neither service selection nor a skipped branch removes deployed Lambdas or IAM policies.
The post-build phase independently reapplies default-enabled flags for older CodeBuild projects with no service-selection environment variables. Artifact validation must still reject missing selected-service CSVs in that upgrade path.
OWASP reads only BR/SM/AC direct evidence, never Agent Registry CSVs. For each OW ID whose mapped sources include a deselected service, it emits one N/A/Informational selection-coverage row per regional invocation. Existing rows from remaining sources are preserved; a sole-source control such as OW-07 stays visible as unassessed when Bedrock is off. Missing selected artifacts still use OW-00 and are not conflated with intentional deselection.
GRC intentionally remains independent and can call deselected services’ APIs. It also runs when OWASP alone is enabled. Disable both optional areas to run only the selected direct assessments. GRC remediation must be self-contained rather than refer to a direct-service check that might have been omitted.
Direct-service scores exclude GRC, Agentic AI, and compliance rows. Changing selection changes the score denominator and can raise or lower the pass rate; compare reports with the same scope. Central buckets retain historical CSVs; downstream readers must filter by execution ID, as the consolidation path does.
Catalog totals describe the available controls, not the number executed by every
selection. The default sample reports still illustrate all services enabled.
Regression coverage in tests/test_service_selection.py exercises all 16 direct
service combinations, both deployment paths, artifact requirements, OWASP source
selection, and single-/multi-account reporting.
The framework includes 94 core security checks across Amazon Bedrock, Amazon SageMaker AI, Amazon Bedrock AgentCore, and AWS Agent Registry, plus 38 always-on Agentic AI Security checks, 64 optional Responsible AI GRC checks when EnableResponsibleAIGRCAssessment is enabled, and 12 optional OWASP Top 10 for LLM checks when EnableOWASPAssessment is enabled. For the complete list of checks with descriptions, see the Security Checks Reference.
Each core service assessment AWS Lambda function:
region_name parameterRegion column)AWS Agent Registry is a separate regional assessment Lambda. It creates its
own agent-registry-control client, writes
agent_registry_security_report_<execution_id>_<region>.csv, and runs
independently of Amazon Bedrock AgentCore availability because the services
have separate endpoints.
The Responsible AI GRC assessment Lambda is different. It is deployed in both SAM templates, but Step Functions invokes it only from the first region iteration (RegionIndex == 0) when the execution input includes "enableResponsibleAIGRC": "true" or "enableOWASP": "true". The OWASP path uses FS-* findings as hidden source rows unless the capability was explicitly enabled. It receives the full TargetRegions list and emits findings with Region values so the report can display the same regional filters as the core services.
Compatibility contracts. The
FS-*check IDs are permanent.EnableFinServAssessment/ENABLE_FINSERVare retained permanently as a legacy alias for the primaryEnableResponsibleAIGRCAssessment/ENABLE_RESPONSIBLE_AI_GRC/enableResponsibleAIGRCnames — see Responsible AI GRC alias migration guide — and archived reports/CSVs generated before this rename keep their original filenames and selectors. See Responsible AI GRC — scope, sources, and compatibility for the full list of what changed and what stayed the same.The alias stops at the CloudFormation parameter / CodeBuild environment variable layer — it is not a compatibility contract for the Step Functions execution input.
buildspec.ymlresolvesEnableFinServAssessment/ENABLE_FINSERVinto the effectiveENABLE_RESPONSIBLE_AI_GRCvalue and passes only"enableResponsibleAIGRC"intoStartExecution. The state machine’sResponsible AI GRC Enabled?Choice state does not have a passthrough branch for a legacy"enableFinServ"execution-input key: if"enableFinServ": "true"reachesStartExecutiondirectly (bypassing CodeBuild/buildspec entirely, e.g. a hand-written script or an old runbook), the execution fails immediately with errorLegacyEnableFinServInputRejectedinstead of silently skipping the Responsible AI GRC checks. Use"enableResponsibleAIGRC": "true"(or"enableOWASP": "true") in the execution input instead.
Additional Functions:
TargetRegions parameter for the Map stateOrganization-specific baselines are CloudFormation parameters rather than
hard-coded scanner assumptions. Keep each parameter wired through both direct
SAM templates, both top-level deployment templates, the corresponding CodeBuild
environment variable, and every sam deploy path in buildspec.yml.
| CloudFormation parameter | Lambda environment variable | Check |
|---|---|---|
RequireBedrockZeroDataRetention |
REQUIRE_BEDROCK_ZERO_DATA_RETENTION |
BR-37 |
RequireMarketplaceEndpointCMK |
REQUIRE_MARKETPLACE_ENDPOINT_CMK |
BR-40 |
RequireAgentCoreOnlineEvaluation |
REQUIRE_AGENTCORE_ONLINE_EVALUATION |
AC-17 |
RequireAgentRegistryManualApproval |
REQUIRE_AGENT_REGISTRY_MANUAL_APPROVAL |
AR-03 |
RequireAgentRegistryCMK |
REQUIRE_AGENT_REGISTRY_CMK |
AR-05 |
AgentCoreTokenVaultId |
AGENTCORE_TOKEN_VAULT_ID |
AC-14 |
ApprovedExternalAccountIds |
AIML_APPROVED_EXTERNAL_ACCOUNT_IDS |
SM-30 |
ApprovedOrganizationIds |
AIML_APPROVED_ORG_IDS |
SM-30 |
The top-level deployment templates expose the approved-account and
approved-organization values to CodeBuild as APPROVED_EXTERNAL_ACCOUNT_IDS
and APPROVED_ORGANIZATION_IDS; buildspec.yml then maps them to the SAM
parameters shown above. Add or rename a baseline only when all layers and the
public deployment documentation are updated together.
To add a new AI/ML service (for example, Amazon Comprehend, Amazon Textract):
# Example: Adding Comprehend security assessment
mkdir -p aiml-security-assessment/functions/security/comprehend_assessments
cd aiml-security-assessment/functions/security/comprehend_assessments
# app.py
import boto3
import os
import json
from botocore.config import Config
from botocore.exceptions import ClientError, EndpointConnectionError
from schema import create_finding
boto3_config = Config(retries=dict(max_attempts=10, mode="adaptive"))
def lambda_handler(event, context):
"""Main assessment handler for new service"""
all_findings = []
# Extract target region from Step Functions Map state
region = event.get("Region", os.environ.get("AWS_REGION", "us-east-1"))
# Verify service availability in this region
try:
test_client = boto3.client("comprehend", config=boto3_config, region_name=region)
test_client.list_endpoints(MaxResults=1)
except EndpointConnectionError:
# Service not available — create an Informational N/A finding, write
# the regional CSV artifact, then return its URL. Do not return early
# without an artifact: that makes the assessment area look empty.
return write_unavailable_report(
execution_id=event["Execution"]["Name"],
region=region,
detail=f"Comprehend is not available in {region}.",
)
except ClientError as error:
if is_region_unsupported(error):
return write_unavailable_report(
execution_id=event["Execution"]["Name"],
region=region,
detail=f"Comprehend is not available in {region}.",
)
raise
# Get cached permissions
execution_id = event["Execution"]["Name"]
permission_cache = get_permissions_cache(execution_id)
# Run assessment checks (pass region to each)
findings = check_new_service_security(permission_cache, region=region)
all_findings.append(findings)
# Generate and upload report (include region in S3 key)
csv_content = generate_csv_report(all_findings)
bucket_name = os.environ.get("AIML_ASSESSMENT_BUCKET_NAME")
s3_url = write_to_s3(execution_id, csv_content, bucket_name, region=region)
return {
"statusCode": 200,
"body": {
"message": "New service assessment completed",
"findings": all_findings,
"report_url": s3_url,
},
}
def check_new_service_security(permission_cache, region: str = ""):
"""Implement your security checks here"""
findings = {
"check_name": "New Service Security Check",
"status": "PASS",
"details": "",
"csv_data": [],
}
# Create regional client
client = boto3.client("comprehend", config=boto3_config, region_name=region)
# Your assessment logic here
# Pass region= to all create_finding() calls
return findings
# requirements.txt
boto3==1.43.85
botocore==1.43.85
# schema.py
from enum import Enum
class SeverityEnum(str, Enum):
HIGH = "High"
MEDIUM = "Medium"
LOW = "Low"
INFORMATIONAL = "Informational"
class StatusEnum(str, Enum):
FAILED = "Failed"
PASSED = "Passed"
NA = "N/A"
def create_finding(
check_id, finding_name, finding_details, resolution, reference, severity, status, region=""
):
"""Create standardized finding format
Args:
check_id: Unique check identifier (for example, BR-01, SM-01, AC-01, AR-01)
finding_name: Name of the finding
finding_details: Detailed description
resolution: Steps to resolve. N/A findings can still include an
explanatory "No action required" or permission-remediation message.
reference: Documentation URL
severity: SeverityEnum value
status: StatusEnum value (Failed, Passed, or N/A)
region: AWS region where the finding was identified
"""
return {
"Check_ID": check_id,
"Finding": finding_name,
"Finding_Details": finding_details,
"Resolution": resolution,
"Reference": reference,
"Severity": severity,
"Status": status,
"Region": region,
}
Add your new function to both SAM templates:
aiml-security-assessment/template.yamlaiml-security-assessment/template-multi-account.yaml ComprehendSecurityAssessmentFunction:
Type: AWS::Serverless::Function
Properties:
FunctionName: !Sub 'aiml-security-${AWS::StackName}-ComprehendAssessment'
CodeUri: functions/security/comprehend_assessments/
Handler: app.lambda_handler
Runtime: python3.12
Timeout: 600
MemorySize: 1024
Environment:
Variables:
AIML_ASSESSMENT_BUCKET_NAME: !Ref AIMLAssessmentBucket
TARGET_REGIONS: !Ref TargetRegions
Policies:
- Statement:
- Sid: ComprehendReportWrite
Effect: Allow
Action:
- s3:PutObject
Resource: !Sub '${AIMLAssessmentBucket.Arn}/comprehend_security_report_*.csv'
- Sid: ComprehendReadPermissions
Effect: Allow
Action:
# Example only: grant the exact operations used by app.py.
- comprehend:ListEndpoints
- comprehend:DescribeEndpoint
Resource: '*'
Add the new service to the Run Security Assessments parallel branch inside the Scan Regions Map state in aiml-security-assessment/statemachine/assessments.asl.json. Also add the function ARN substitution and LambdaInvokePolicy for the new function in both SAM templates.
{
"Parallel Service Assessments": {
"Type": "Parallel",
"Branches": [
{
"StartAt": "Bedrock Security Assessment",
"States": {"Bedrock Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true}}
},
{
"StartAt": "SageMaker Security Assessment",
"States": {"SageMaker Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true}}
},
{
"StartAt": "AgentCore Security Assessment",
"States": {"AgentCore Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true}}
},
{
"StartAt": "AWS Agent Registry Security Assessment",
"States": {"AWS Agent Registry Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true}}
},
{
"StartAt": "Comprehend Security Assessment",
"States": {"Comprehend Security Assessment": {"Type": "Task", "Resource": "arn:aws:states:::lambda:invoke", "End": true}}
}
]
}
}
Add every exact assessment-service action used by the new function to that
function’s Policies block in both SAM templates. Do not use service-wide
wildcards such as comprehend:List*, and do not add assessment-service actions
to
deployment/1-aiml-security-member-roles.yaml,
deployment/2-aiml-security-codebuild.yaml, or
deployment/aiml-security-single-account.yaml. Those templates contain
deployment/orchestration roles, not assessment runtime roles. Add the exact
runtime actions and prefix-scoped report-object access only to the new Lambda’s
policy statements in both SAM templates.
Before merging a new check, validate every new boto3 operation and its exact IAM action against the AWS Knowledge MCP documentation tools. Confirm the client (control-plane versus data-plane), IAM prefix, and any resource or condition-key constraints; add grants to both SAM templates in the same change.
Test your new assessment function locally:
cd aiml-security-assessment
sam build --template template.yaml
sam local invoke ComprehendSecurityAssessmentFunction --event testfile.json
Most day-to-day contributions add or update individual security checks inside the existing assessment packages (Bedrock, SageMaker, AgentCore, or AWS Agent Registry) rather than creating an entire new service package.
Define the check contract before coding: Record the Check_ID, exact resource and regional or account-global scope, AWS client and operation, response field and allowed values, successful empty-response behavior, and what evidence produces Passed, Failed, or N/A. Verify the request/response shape and enum values against the installed botocore service model and AWS documentation. An HTTP 200 or an empty response does not establish compliance; distinguish account-level policies from resource-level settings. Compare the resulting predicate with the catalog description and remediation before writing tests.
Locate the target file: Choose bedrock_assessments/app.py, sagemaker_assessments/app.py, agentcore_assessments/app.py, or agent_registry_assessments/app.py. New checks must follow the existing function structure and naming patterns inside that file.
Implement the check: Write a function that returns a dict with a "csv_data" list of findings. Always pass region=region (or the loop variable) to every create_finding() call. Use the shared schema.py helpers where present. Keep each check’s ID consistent in normal findings and handler fallback/error findings.
status="N/A". Unsupported regional APIs and features are Informational; follow the target package’s severity convention for access-denied and other could-not-assess results (Responsible AI GRC uses Low).is_region_unsupported() and describe_api_error() helpers in bedrock_assessments/app.py (or equivalent patterns) instead of raw string matching.status="Failed" and resolution="No action required".N/A row for each affected check. An access-denied or partial inventory must not claim that no resources exist. Keep Passed resolution text free of remediation instructions.REQUIRE_MARKETPLACE_ENDPOINT_CMK=false), emit N/A + Informational when the hardening gap is observed; reserve Passed only for controls that were checked and satisfied.Inventory scope, pagination, and error isolation: Use get_paginator() or the _agentcore_list_all pattern for list APIs, and test a non-compliant item on a later page. If a safety cap or deadline truncates the list, retain the findings collected and add an N/A incomplete-inventory notice; do not claim “all resources” passed. Wrap per-resource detail calls in individual try/except blocks so one throttle or delete race does not erase other resources’ findings. Run account-global inventories once or de-duplicate them with a truthful Region value.
Synthesized mappings (if applicable): If the new check should also appear under the Agentic AI lens (AG- prefix) or an OWASP category, update the corresponding mapping dictionary. Values in OWASP_CHECK_MAPPINGS are lists because one source check can emit multiple OW- rows. Allocate new AG numbers by hand across the Bedrock, AgentCore, and Agent Registry mapping files and native checks to avoid collisions (current catalog ends at AG-38).
Add tests: Use SDK-shaped responses for compliant/pass, non-compliant/fail (or advisory N/A when that is the defined result), no-resource, access-denied, region-unavailable, and unexpected-error cases, including a successful empty response when the API can return one. Empty-inventory tests alone do not exercise a check’s compliance predicate. Shared inventories need later-page, list-error, and per-resource detail-error cases; account-global inventories need a two-region case. Assert the exact Check_ID on fallback rows and that a failed check leaves the other checks’ rows in the CSV. See tests/test_bedrock_checks.py and tests/test_sagemaker_checks.py for patterns.
Update documentation: Align the check’s actual evidence, resource scope, status behavior, severity, and remediation with docs/SECURITY_CHECKS.md and the applicable docs/SECURITY_CHECKS_*.md catalog. Review README.md, this guide, docs/TROUBLESHOOTING.md, CHANGELOG.md, and relevant scope, sample-report, and diagram documents; update every changed count, table-of-contents entry, prefix table, and deep link. Do not promise a control or report state the implementation does not establish.
ruff check and ruff format --check only on the changed .py files (match CI scope).tests/ (which includes test_consolidate_responsible_ai_grc.py), responsible_ai_grc_tests/, and the report-pipeline session.cfn-lint on any edited templates.create_finding() function for consistent outputPassed: Resources were checked and met the assessed best practiceFailed: Resources were checked and found non-compliantN/A: The check is not applicable, requires manual review, targets an unavailable API or region, or could not determine a result. Use the target package’s severity convention: advisory and unavailable-feature results are Informational, while Responsible AI GRC could-not-assess results are Low.High: Critical security issues requiring immediate attentionMedium: Important security improvements recommendedLow: Minor optimizations suggestedInformational: Advisory information, no-resource results, or unavailable-feature N/A dispositionsN/A incomplete row for that item while preserving findings already
collected. A confirmed delete race may be skipped.N/A “Incomplete” row under the same Check_ID and the handler
still writes its CSV. See _run_check_safely() in
agent_registry_assessments/app.py. An exception escaping at check scope
can abort every subsequent check.Catch path. Returning {"statusCode": 500} without raising
makes the Lambda invocation appear successful and can leave the report
without a visible assessment result.# Test an individual SAM function
cd aiml-security-assessment
sam build --template template.yaml
sam local invoke NewServiceSecurityAssessmentFunction --event test-event.json
# Deploy to test account
sam deploy --stack-name aiml-security-test --capabilities CAPABILITY_IAM
# Execute AWS Step Functions
aws stepfunctions start-execution \
--state-machine-arn arn:aws:states:region:account:stateMachine:TestStateMachine \
--input '{"accountId":"123456789012","enableResponsibleAIGRC":"false","enableOWASP":"false"}'
When adding a new check, extending a lens (AG-), or introducing a new compliance standard (OW-, NR-, EU-, etc.), you must generate the HTML report from the test fixtures and verify the output before creating a pull request. This catches data-routing, template, and rendering issues that unit tests alone may not surface.
# Generate viewable HTML reports from the fixture data
(cd aiml-security-assessment/functions/security/generate_consolidated_report \
&& ../../../../.venv/bin/python -m pytest test_generate_report.py \
-k "generate_viewable_report or generate_multi_account_report" -s --tb=no)
The generated reports are written under aiml-security-assessment/functions/security/generate_consolidated_report/test_reports/. Open the single-account and multi-account HTML files in a browser (desktop and mobile viewports) and verify:
Check_ID (for example BR-41, SM-31, AR-09, AG-33, OW-13, or NR-01) appears in the findings table with the expected Finding name, Severity badge, Status, Region, and Resolution text.If the report looks correct, commit the generated test_reports/*.html files only if your change intentionally updates the canonical fixtures; otherwise they are git-ignored or regenerated on demand.
For detailed troubleshooting guidance, common issues, and debugging tips, see the Troubleshooting Guide.
MaxRegionConcurrencyEnableResponsibleAIGRCAssessment; the Lambda is deployed by default and also runs as a hidden OWASP source dependency when EnableOWASPAssessment is enabledEnableOWASPAssessment; the Lambda is deployed by default but invoked only when enabledReport generation uses a single shared template (report_template.py) for both deployment modes:
aiml-security-assessment/functions/security/generate_consolidated_report/
├── app.py # Lambda handler (single-account)
├── report_template.py # Shared HTML/CSS/JS template
└── ...
consolidate_html_reports.py # CodeBuild script (multi-account)
| Component | Mode | Description |
|---|---|---|
app.py (AWS Lambda) |
mode='single' |
Generates per-account HTML reports during AWS Step Functions execution |
consolidate_html_reports.py |
mode='multi' |
Consolidates all account reports in AWS CodeBuild post-build phase |
Both call generate_html_report() from report_template.py with different parameters.
To update report styling, layout, or features:
report_template.py only - changes apply to both single and multi-account reports../../../../.venv/bin/python -m pytest test_generate_report.py -vget_html_template() - HTML/CSS/JS structuregenerate_table_rows() - Finding row generationgenerate_html_report() - Main entry point with mode parameter (‘single’ or ‘multi’)The assessment_history/ package at the repository root compares each
account’s current run with its previous usable run and writes the changes
report. buildspec.yml runs it in the post-build phase (run_changes_report),
after the existing reports. It reads the findings CSVs in the central bucket
and never fails the build. User-facing behavior is in
Changes Since Last Assessment.
| Module | Role |
|---|---|
models.py |
Record shapes, change states, area routing, CSV columns |
normalize.py |
Day counts and dates blanked out before matching, each tied to the scanner code that writes it |
compare.py |
Pairs two runs’ rows and labels each; no AWS calls |
discover.py |
Groups an account folder’s CSVs into runs, picks the previous usable run (complete files, and a run record that doesn’t say it failed), reads the CSVs |
render_common.py, render_changes.py |
The HTML page, reusing report_template.py’s CSS, escaping, names, and icons |
__main__.py |
python3 -m assessment_history compare: S3 mode for the build, local-folder mode for people |
When you change something it depends on:
Finding_Details: add a
pattern to VOLATILE_PATTERNS in normalize.py, or an IGNORED_SOURCES
entry with a reason. tests/test_assessment_history_normalize.py fails
until you do.PREFIX_TO_MODULE
in discover.py and the module and area lists in models.py. Files with an
unknown prefix are logged and not read.COMPLIANCE_PREFIX_TO_AREA in models.py
must match COMPLIANCE_STANDARDS in report_template.py; a test checks it.report_template.py: the changes page reuses its styling and helpers;
guard tests fail if something it relies on moves.sample-reports/scripts/build_changes_sample.py
and review the diff.Run the package’s tests with its 100% line and branch coverage bar (CI runs the same check):
.venv/bin/python -m pytest tests/test_assessment_history_*.py \
--cov=assessment_history --cov-branch --cov-report=term-missing \
--cov-fail-under=100
The Agentic AI Security lens (AG-01 through AG-38) is synthesized at runtime, not produced by a separate scanner. It re-uses findings from the core Bedrock, AgentCore, and AWS Agent Registry assessments plus a small number of native gateway checks.
bedrock_assessments/app.py → AGENTIC_BEDROCK_CHECK_MAPPINGSagentcore_assessments/app.py → AGENTIC_AGENTCORE_CHECK_MAPPINGSagent_registry_assessments/app.py → AGENTIC_AGENT_REGISTRY_CHECK_MAPPINGSbedrock-agentcore-control client.AG- prefix as its dedicated Agentic AI assessment area. COMPLIANCE_STANDARDS is the separate registry for OWASP and future compliance standards.docs/SECURITY_CHECKS.md and run the full mapping-drift, test-coverage, and gate checklist before merging.The HTML presentation under “By Compliance Standard” is data-driven. Adding a new standard such as NIST AI RMF or the EU AI Act follows the OWASP pattern, but the execution flag, artifact access, and required-artifact checks must also be wired. Concrete steps:
^[A-Z]{2,3}-\d{2}$:
OW- (already implemented)NR-EU-aiml-security-assessment/functions/security/<slug>_assessments/:
owasp_assessments/schema.py verbatim.owasp_assessments/requirements.txt.app.py following the OWASP pattern: read per-service CSVs
from S3, apply your <STANDARD>_CHECK_MAPPINGS dict, run any native
checks, write <slug>_security_report_<execution>_<region>.csv.EnableOWASPAssessment deployment pattern:
deployment/aiml-security-single-account.yaml (parameter definition,
parameter group, CodeBuild env var)deployment/2-aiml-security-codebuild.yaml (parameter definition,
parameter group, CodeBuild env var)buildspec.yml (ENABLE_<STANDARD> export and every applicable
aws stepfunctions start-execution input)aiml-security-assessment/template.yaml (new function resource,
LambdaInvokePolicy, definition-substitution entry, and any SAM-level
parameter exposed for direct deployments)aiml-security-assessment/template-multi-account.yaml (same)aiml-security-assessment/statemachine/assessments.asl.json (flag
reference and enabled/skipped branches; see step 5)
Verify that any declared SAM parameter is consumed by the state machine.
A direct StartExecution must supply the flag in its input when the
state machine expects it; a CodeBuild environment variable alone does not
set that direct-execution input.aiml-security-assessment/template.yaml and
aiml-security-assessment/template-multi-account.yaml, add the grant
to the specific function’s Policies block that actually makes
the call — not to another function’s policy.s3:ListBucket on the relevant prefix when it lists keys and
s3:GetObject on the matching object ARN when it reads them. Check
both SAM templates and any source-CSV readers.deployment/1-aiml-security-member-roles.yaml,
deployment/2-aiml-security-codebuild.yaml, or
deployment/aiml-security-single-account.yaml. Those deployment roles
only create/update the SAM stack, poll executions, and retrieve reports.
New deployment-time AWS operations require a separate least-privilege
review of the affected orchestration role.tests/test_member_role_policy_size.py.aiml-security-assessment/statemachine/assessments.asl.json:
<Standard> Enabled? Choice state that reads
$.OriginalInput.enable<Standard>.<Standard> Security Assessment Task state that invokes the
new function.Run Security Assessments (Parallel) → OWASP Enabled? →
… → <Standard> Enabled? → … → back to the region map end.
Replace End: true with Next on each preceding success, skipped, or
incomplete route that must reach the new choice.OriginalInput, Execution, Region, RegionIndex, and other
downstream fields across success, skipped, and caught-error paths.
Set ResultPath explicitly on result-producing Task/Pass states and
on Catch; a Catch ResultPath does not protect the success path.
Test each path through the next Choice and report-generation state.COMPLIANCE_STANDARDS in report_template.py. Each entry needs
slug, name, prefix, icon, icon_small, reference_url,
section_title, and truthful scope_text. The slug and Check_ID
prefix must be unique. Choose an icon colour distinct from the existing
report navigation colours.
generate_html_report() builds the “By Compliance
Standard” sidebar entry, filter option, dashboard card, findings
section and summary, Assessment Scope chip, and source link when
the standard has rows. Extend this registry instead of copying an
OWASP-specific HTML block.get_html_template() placeholders together. Keep one slug across
the section id, finding-row data-service, filter option value,
summary data-filter-service, and scope-chip data-scope-service.
Review affected CSS/JavaScript selectors. The Assessment Scope block
also uses exact str.replace() anchors after rendering; update those
anchors when changing their source markup, since a missed match fails
silently.Check both routing and artifact completeness. The report generator
(aiml-security-assessment/functions/security/generate_consolidated_report/app.py)
and the multi-account consolidator (consolidate_html_reports.py) both
iterate COMPLIANCE_STANDARDS to initialise service_stats /
service_findings and to route rows by Check_ID prefix. This handles
presentation and row routing, but it does not add the new CSV to
the report Lambda’s per_region_categories / expected_artifacts
validation or grant permission to list and read it. Extend those
selected-artifact checks, map any CSV filename fragment that differs
from the report slug, and verify CodeBuild stages the CSV for the
multi-account consolidator. A selected standard with a missing,
unreadable, or header-only CSV must be reported as incomplete, not
displayed with zero findings.
Update docs: add a SECURITY_CHECKS_<STANDARD>.md in the OWASP
style, bump the check count in README.md and docs/SECURITY_CHECKS.md.
Add tests: mapping emission, native-check behavior, enabled/skipped
Step Functions paths, artifact completeness, routing, and report-template
rendering. Use representative synthetic CSVs with positive, negative,
N/A, and empty-inventory findings; test missing and unreadable artifacts
separately. Include IAM coverage assertions for the new Lambda and
report reader in both SAM templates. See tests/test_owasp_checks.py and
tests/test_report_template_owasp.py as templates.
When you modify the report template or add new features, update the sample reports and screenshots:
After making changes to report_template.py, regenerate sample reports from a fresh assessment run or from the local report test fixtures. The existing test_generate_report.py file is a pytest/unittest test module, not a standalone --mode/--output CLI.
# Generate local viewable reports from fixtures
(cd aiml-security-assessment/functions/security/generate_consolidated_report \
&& ../../../../.venv/bin/python -m pytest test_generate_report.py \
-k "generate_viewable_report or generate_multi_account_report" -s)
The fixture reports are written under aiml-security-assessment/functions/security/generate_consolidated_report/test_reports/. Use them to validate report rendering before refreshing the canonical files in sample-reports/.
The repository includes an automated screenshot capture tool:
# Prepare or verify the optional screenshot environment without changing files
./sample-reports/scripts/capture_screenshots.py --check-dependencies
# Capture and optimize screenshots
./sample-reports/scripts/capture_screenshots.py
The repository-root .venv must already exist and use Python 3.12. The script
re-launches itself with .venv/bin/python, installs
sample-reports/dev-requirements.txt into that environment when needed, and
installs Chromium under .venv/playwright-browsers.
What the script does:
sample-reports/ folderWhat gets generated:
The script captures 4 screenshots:
dashboard-overview-light.png - Executive dashboard in light modedashboard-overview-dark.png - Executive dashboard in dark modefindings-table.png - Detailed findings table with filtersmulti-account-summary.png - Multi-account consolidated viewAll screenshots are automatically optimized (target: 200-300KB each, ~700KB total).
Customization:
Edit sample-reports/scripts/capture_screenshots.py to customize:
# Viewport size
VIEWPORT_WIDTH = 1440
VIEWPORT_HEIGHT = 900
# Image quality
JPEG_QUALITY = 85 # Range: 1-100
PNG_OPTIMIZE = True
# Add new screenshots to SCREENSHOTS list
SCREENSHOTS = [
{
"name": "my-screenshot",
"file": "security_assessment_single_account.html",
"description": "My Custom View",
"actions": [
{"type": "wait", "selector": ".element", "timeout": 2000},
{"type": "click", "selector": ".button"},
{"type": "scroll", "position": 500},
],
"clip": {"x": 0, "y": 0, "width": 1440, "height": 800},
}
]
Available action types:
wait - Wait for selector (for example, {"type": "wait", "selector": ".metrics", "timeout": 2000})click - Click element (for example, {"type": "click", "selector": ".theme-toggle"})scroll - Scroll to position (for example, {"type": "scroll", "position": 500})wait_time - Wait milliseconds (for example, {"type": "wait_time", "ms": 300})Troubleshooting:
| Issue | Solution |
|---|---|
.venv not found |
Create it with python3.12 -m venv .venv, then bootstrap the repository dependencies |
| Playwright or Chromium missing | Run ./sample-reports/scripts/capture_screenshots.py --check-dependencies |
| Sample reports not found | Run from repository root |
| Screenshots too large | Lower JPEG_QUALITY or reduce viewport size |
| Browser launch fails | Run playwright install-deps (Linux only) |
After generating new screenshots, update the README to reference them:
### Sample Assessment Reports
**Preview:**

*Executive summary with severity counts and assessment-area breakdown*

*Interactive findings table with filtering capabilities*
The changes sample and the assessment-history golden saved answers are built from the two sample reports. After regenerating a sample report, run:
.venv/bin/python sample-reports/scripts/build_changes_sample.py
Review the diff under sample-reports/ and
tests/fixtures/assessment_history/golden/. --check verifies without writing.
Then refresh its screenshot with
./sample-reports/scripts/capture_changes_screenshot.py, which captures only
the changes page and doesn’t rewrite any report.
dashboard-overview-light.png, not screenshot1.pngsample-reports/ for easy organizationEvery releasable behavior or deployment change must update the root
CHANGELOG.md under Unreleased; a release does not need to be created for
every merged change. Use one or more of these deployment-impact categories:
deployment/aiml-security-single-account.yaml changed.deployment/1-aiml-security-member-roles.yaml changed. This update must
complete before CodeBuild runs.deployment/2-aiml-security-codebuild.yaml changed.The changelog should list the exact changed deployment templates and give the
required order. When a release is pinned by tag or commit, it should also remind
users to update the GitHubBranch stack parameter. When a version is tagged,
move the accumulated entries into a dated version section and recreate an empty
Unreleased section. The end-user procedure and repository-diff fallback are documented in
Upgrading to a New Release.
GitHub Actions workflows run automatically to validate code quality and security on every pull request.
| Workflow | File | What It Checks |
|---|---|---|
| Python Code Quality | .github/workflows/python-lint.yml |
ruff check (lint) and ruff format --check (formatting) on changed .py files |
| Python Tests | .github/workflows/python-tests.yml |
Runs upstream tests, Responsible AI GRC tests, and report-pipeline tests in separate pytest sessions |
| CloudFormation Lint | .github/workflows/cfn-lint.yml |
Validates deployment and SAM templates with cfn-lint |
| SAM Validate & Build | .github/workflows/sam-validate.yml |
Runs sam validate --lint on both SAM templates and sam build on the single-account template |
| ASH Security Scan | .github/workflows/ash-security-scan.yml |
Scans changed files for secrets, dependency vulnerabilities, and IaC misconfigurations |
Additional workflows run post-merge or on schedule:
| Workflow | File | Trigger |
|---|---|---|
| ASH Full Repository Scan | .github/workflows/ash-full-repository-scan.yml |
Push to main, monthly schedule, manual |
| Labeler | .github/workflows/label.yml |
Auto-labels PRs by changed paths (bedrock, sagemaker, agentcore, deployment, docs) |
cfn-lint suppressions are configured in .cfnlintrc at the repository root for IAM actions not yet in cfn-lint’s database (for example, bedrock-agentcore actions).
Before pushing, run these checks locally to catch issues early:
# Bootstrap or refresh the repository-local virtual environment
.venv/bin/pip install -r tests/requirements.txt \
-r aiml-security-assessment/functions/security/agent_registry_assessments/requirements.txt \
-r aiml-security-assessment/functions/security/agentcore_assessments/requirements.txt \
-r aiml-security-assessment/functions/security/bedrock_assessments/requirements.txt \
-r aiml-security-assessment/functions/security/cleanup_bucket/requirements.txt \
-r aiml-security-assessment/functions/security/responsible_ai_grc_assessments/requirements.txt \
-r aiml-security-assessment/functions/security/generate_consolidated_report/requirements.txt \
-r aiml-security-assessment/functions/security/iam_permission_caching/requirements.txt \
-r aiml-security-assessment/functions/security/owasp_assessments/requirements.txt \
-r aiml-security-assessment/functions/security/resolve_regions/requirements.txt \
-r aiml-security-assessment/functions/security/sagemaker_assessments/requirements.txt
.venv/bin/pip check
# Match CI by checking Python files changed relative to main
changed_py=$(git diff --name-only --diff-filter=ACMR origin/main...HEAD -- '*.py')
.venv/bin/ruff check $changed_py
.venv/bin/ruff format --check $changed_py
# Unit tests. tests/ is one session; the Responsible AI GRC and report-pipeline
# suites need their own sessions because they live outside tests/.
export AIML_ASSESSMENT_BUCKET_NAME=test-assessment-bucket
export AWS_DEFAULT_REGION=us-east-1
export AWS_ACCESS_KEY_ID=testing
export AWS_SECRET_ACCESS_KEY=testing
.venv/bin/python -m pytest tests/ -v --tb=short
.venv/bin/python -m pytest aiml-security-assessment/functions/security/responsible_ai_grc_tests/ -v --tb=short
(cd aiml-security-assessment/functions/security/generate_consolidated_report \
&& ../../../../.venv/bin/python -m pytest test_generate_report.py -v --tb=short)
# CloudFormation lint
.venv/bin/cfn-lint deployment/*.yaml \
aiml-security-assessment/template.yaml \
aiml-security-assessment/template-multi-account.yaml
# SAM validate and build
(cd aiml-security-assessment \
&& sam validate --template template.yaml --lint \
&& sam validate --template template-multi-account.yaml --lint \
&& sam build --template template.yaml \
&& sam build --template template-multi-account.yaml)
This developer guide provides the foundation for extending the AI/ML Security Assessment Framework. As you add new AI/ML services and security checks, please update this documentation to help future contributors understand and build upon your work.