After each assessment run, the framework compares each account’s results with that account’s previous run and writes a “Changes since last assessment” report next to the run’s main report. It shows which findings were resolved, which are still open, which regressed, which are new, and which no longer appear.
For every account whose run completed, in the same folder as that run’s main
report (<AssessmentBucket>/<account_id>/):
| File | What it is |
|---|---|
security_assessment_changes_<YYYYMMDD_HHMMSS>.html |
Interactive page: summary counts, counts by assessment area, and a filterable table of findings by change |
security_assessment_changes_<YYYYMMDD_HHMMSS>.csv |
The same results, one row per finding, with both runs’ details, statuses, and severities |
The timestamp is when the current run’s results were saved, in UTC, so the name identifies the run it describes. Running the comparison again for the same two runs overwrites the file instead of adding another.
In single-account mode the build also writes a small run record,
assessment_history_run_<execution_id>.json, after every run. It says
whether the run’s Step Functions execution succeeded; see
Which run is “the previous run”.
The first run of an account has nothing to compare with, so no changes report is written for it. Every later run gets one, including runs where nothing changed (“No changes since the last assessment”).
A sample page and CSV are in sample-reports/
(security_assessment_changes.html and .csv), built from the
single-account sample report. Resource IDs and the random parts of resource
names in them are made-up values of the same shape, so they don’t match the
sample main report.

The report is controlled by the EnableAssessmentHistory deployment
parameter in both deployment/aiml-security-single-account.yaml and
deployment/2-aiml-security-codebuild.yaml:
| Value | Effect |
|---|---|
true (default) |
A changes report is written after each run |
false |
The comparison is skipped; the build log says Changes report disabled (EnableAssessmentHistory=false) |
The parameter reaches the build as the ENABLE_ASSESSMENT_HISTORY
CodeBuild environment variable. Stacks created before the parameter existed
don’t have the variable, and are treated as true. A change takes effect on
the next build.
Findings CSVs are saved every run whatever the setting, so after turning the report back on, the next run compares with the most recent usable run, including runs made while it was off. Run records are written only while the setting is on, so runs made while it was off are judged by their files alone.
The comparison runs in the CodeBuild post-build phase (buildspec.yml),
after the run’s results are in the central AssessmentBucket:
It can’t fail an assessment run:
WARNING line and the build carries on.WARNING line, and a summary line counts them.It reads only the findings CSVs already in the bucket, and in single-account mode the run records. No changes were made to the scanners, the AWS SAM templates, the Step Functions workflow, or the IAM roles.
Runs are told apart by the Step Functions execution ID in each findings CSV’s name. Execution IDs are random, so runs are ordered by when S3 saved their files; a run’s time is its latest file’s time.
The previous run is the most recent usable run saved before the current run. A run is usable when:
Its run record, if it has one, says it succeeded. In single-account
mode the build writes assessment_history_run_<execution_id>.json after
every run. A Step Functions execution only succeeds when the main report’s
completeness check passes, so a run recorded as failed is never used. Runs
with no record (made before this was added, or while the setting was off)
are judged by their files alone. In multi-account mode the build uploads a
run’s files only after checking them, so no records are written.
Service selection. A service turned off with the deployment’s
Enable*Assessment switches writes no CSV, so its CSVs aren’t required. Which
services a run selected comes from:
ENABLE_BEDROCK, ENABLE_SAGEMAKER,
ENABLE_AGENTCORE, and ENABLE_AGENT_REGISTRY settings (a missing setting
means selected);A selected service with no CSV still makes a run incomplete. A run with all four services off (Responsible AI GRC and OWASP only) is usable.
Runs made before release 2.0.0 have no AWS Agent Registry CSV and no run record, so they are judged by their files: AWS Agent Registry counts as not selected in that run and isn’t compared, and checks added in 2.0.0 are listed as found in only one run.
Only what ran in both runs is compared:
Global rows (account-level findings) are always compared. A row
whose Region lists several regions is compared when every listed region was
scanned in both runs.Check IDs that appear in only one run (for example, after upgrading the framework) are listed in the report as a note.
| Previous run | Current run | Change |
|---|---|---|
| Failed | Passed | Resolved |
| Failed | Failed | Still open |
| Passed | Failed | Regressed |
| Not present, or N/A | Failed | New |
| Failed | Not present | No longer reported |
| Failed | N/A | No longer assessed |
All other combinations have no Failed result in either run (for example,
Passed → Passed or Passed → N/A) and are shown as Not failing.
Failed → N/A is not counted as resolved, because N/A also covers
access-denied and unavailable-region results.
Rows are grouped by assessment area, Region, and Check_ID. The Finding
title isn’t part of the group: many checks use one title when they fail and
another when they pass or can’t be assessed (for example, SM-04 writes
“GuardDuty Not Enabled”, “GuardDuty Enabled”, or “GuardDuty Check Error”).
Some checks write one row per resource (for example, one row per IAM role),
with the resource only in Finding_Details. Within a group, rows are paired
in these steps:
Same title, identical details. Rows with the same Finding and
Finding_Details are paired. Then, rows with the same Finding are paired
whose details are the same once values that change on every run are
blanked out:
| Changing value | Written by |
|---|---|
(N days) |
AC-03 (and AG-17, mapped from it), AR-02 |
N days ago |
FS-31 (and OW-09, mapped from it) |
on YYYY-MM-DD, since YYYY-MM-DD |
BR-14, SM-02 |
Only these values are blanked. Digits in general are not, because
resource names contain digits. A test fails if a scanner starts writing
another day count or date into Finding_Details without it being added
to this list.
A pair is made only when it’s unambiguous: if two rows would pair with the
same row in steps 1 and 2, none of them are paired. The CSV’s Match_Rule
column records which step paired each row (exact, normalized, details,
single-row, check-level, or unmatched), so any result can be traced. A
check-level summary row appears once for each row it’s paired with.
Repeated rows within a run are dropped first, with the same rule the main report uses.
The page uses the main report’s layout, styling, and light/dark setting (the choice carries over between the two reports).
Passed → Failed). Filters for search, Region, Assessment Area, Severity,
and Change. By default the table shows every change state and hides Not
failing rows; choose All Findings in the Change filter to show them.
A Still open row whose two Failed rows were paired only because each run had
one, with different details, reads Failed → Failed · details changed and
has a note under its details.
Clicking an area in the sidebar filters the table to it.The sidebar lists the main report and the changes CSV by file name. The links
work when the files are in the same folder, for example downloaded together
or copied with aws s3 sync; opening the page straight from the S3 console
opens one file at a time. The main report has no link to a single finding:
use its search box, for example with the Check ID.
One row per compared finding, in these columns:
| Column | Contents |
|---|---|
Account_ID |
The account |
Assessment_Area |
bedrock, sagemaker, agentcore, agent-registry, agentic, responsible-ai-grc, or owasp |
Region, Check_ID, Finding |
As in the findings CSVs |
Change |
One of the change states above, including Not failing |
Previous_Status, Current_Status |
Passed, Failed, N/A, or Not present |
Previous_Severity, Current_Severity |
Each run’s severity (blank when not present) |
Previous_Finding_Details, Current_Finding_Details |
Each run’s details (blank when not present) |
Resolution, Reference |
From the current run, or the previous run when the finding is gone |
Match_Rule |
exact, normalized, details, single-row, check-level, or unmatched |
Previous_Execution_ID, Current_Execution_ID |
The two runs compared |
A cell that starts with =, +, -, @, a tab, or a carriage return gets
a leading ' so spreadsheet programs don’t run it as a formula; finding text
can contain resource names chosen by anyone who can create resources.
| Message | Meaning |
|---|---|
Changes report for account <id> |
Start of that account’s comparison |
Compared run <id> (saved <time>) with run <id> (saved <time>), N day(s) apart |
The two runs used |
No changes since the last assessment |
Nothing changed; the report is still written |
Written: s3://... |
Where the report was written |
No previous run for account <id>; changes report skipped. If this stack was redeployed, or earlier results were moved or deleted, they aren't compared. ... |
First run for the account, or no usable earlier run in the bucket |
Skipped run <id> saved <time>: incomplete (...) |
An incomplete run passed over |
Skipped run <id> saved <time>: the assessment run did not succeed (...) |
Its run record says the Step Functions execution didn’t succeed |
Skipped run <id> saved <time>: its run record <file> can't be read |
The run record isn’t valid, so the run isn’t used |
Skipped run <id> saved <time>: unreadable (...) |
One of the run’s CSVs can’t be read; the next older run is tried |
Skipped N more run(s) |
More than five runs were passed over |
Not compared (not selected in both runs): <service> (<which run>) |
A core service selected in only one of the two runs, or in neither; it’s left out of the comparison and the counts |
No core service was selected in both runs, so there are no headline counts; see the changes CSV |
Only Responsible AI GRC or OWASP were compared |
WARNING: ENABLE_<SERVICE>='<value>' is not true or false; the services with CSVs are taken as selected |
A service switch has an unexpected value |
WARNING: ENABLE_<SERVICE> is '<value>', not true or false, so the run record does not list the selected services |
Single-account: the record is still written; later runs judge this run by its files |
WARNING: Could not write the assessment run record s3://... |
Single-account: no record for this run; it will be judged by its files |
WARNING: No execution ID was saved, so no assessment run record was written |
Single-account: the run didn’t start |
Ignored N run(s) saved after the current run: ... |
Runs newer than the current run |
Not read: <file> (unknown findings file type) |
A findings CSV from a module this version doesn’t know |
WARNING: Changes report cannot be completed for account <id>. Reason(s): ... |
No report for that account; the reasons are listed |
WARNING: Changes report skipped: only Ns of build time left |
Too little build time left |
WARNING: The time limit was reached with N of M account(s) not started: ... |
The accounts that got no report this build; the next build starts with a different account |
Starting with account <id>: the account order moves along one account with each build number |
Multi-account: which account went first this build |
WARNING: The changes page for account <id> could not be written; the CSV was. Reason: ... |
Only the CSV was written for that account |
Changes report skipped: the assessment run did not succeed |
Single-account run failed |
Changes report disabled (EnableAssessmentHistory=false) |
The setting is off |
AssessmentBucket. Updating the stack keeps it. Deleting and recreating the
stack creates a new bucket, so history starts over. Switching between
single-account and multi-account uses a different bucket, so the first run
after switching has no previous run.The same code compares any two runs on your computer. Copy each run’s findings
CSVs into its own folder (for example with aws s3 cp --recursive --exclude "*"
--include "*_security_report_<execution_id>*"), then run from the repository
root:
.venv/bin/python -m assessment_history compare \
--account 123456789012 \
--previous-dir ./previous-run \
--current-dir ./current-run \
--output-dir ./changes
Each folder must hold one complete run. The setting is ignored in this mode.
| Path | Contents |
|---|---|
assessment_history/models.py |
Record shapes, change states, CSV columns |
assessment_history/normalize.py |
The changing values blanked out in matching step 1 |
assessment_history/compare.py |
The comparison (no AWS calls) |
assessment_history/discover.py |
Finding and reading an account’s runs in S3 or a local folder |
assessment_history/render_common.py, render_changes.py |
The HTML page, using the main report’s styling |
assessment_history/__main__.py |
The command line the build runs |
tests/test_assessment_history_*.py |
Tests; tests/assessment_history_helpers.py builds test data |
tests/fixtures/assessment_history/ |
A hand-written example account folder, and the golden saved answers (golden/expected_*.json) |
sample-reports/scripts/build_changes_sample.py |
Builds the sample page and CSV and the golden saved answers |
sample-reports/scripts/capture_changes_screenshot.py |
Captures sample-reports/changes-overview.png from the sample page |
Run the tests with the package’s 100% line and branch coverage bar (CI runs the same check):
.venv/bin/python -m pytest tests/test_assessment_history_*.py \
--cov=assessment_history --cov-branch --cov-report=term-missing \
--cov-fail-under=100
The golden tests use both sample reports: each is turned back into the
findings CSVs the scanners write (in a temporary folder; only the saved answers
are committed), and a current run is made by applying a short list of edits
(EDITS in the script) that covers every change state and each matching step.
Values that AWS generated in the sample reports (resource IDs, the random parts
of resource names) are replaced with made-up values of the same shape first,
and a test fails if any of them reaches a committed file. After changing a sample report or a comparison rule,
regenerate and review the diff:
.venv/bin/python sample-reports/scripts/build_changes_sample.py
--check makes no changes and exits 1 if anything is out of date. To
refresh the screenshot afterwards:
./sample-reports/scripts/capture_changes_screenshot.py
It uses the same browser setup as capture_screenshots.py, but captures only
the changes page and doesn’t rewrite any report.