Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
89 changes: 89 additions & 0 deletions .challengeai/challenge-ac.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
# ChallengeAC

Agile artifacts: whether what was asked for can be traced to what was built.

## Covers

Acceptance criteria, requirement traceability, ambiguity in what was requested,
and audit readiness of the delivery record.

## The requirement

SP 800-53 Rev. 5 System and Services Acquisition (SA-3, system development life
cycle) and Configuration Management (CM-3, change control). An assessor asks how
the team knows it built what was asked for, and how a change was decided.

## Traceability without ceremony

The requirement is that a capability can be traced to its implementation and its
test. It is not that a particular artifact exists before work starts.

A traceability record written as the work lands describes the system, so it
survives an assessment as written. A backlog written ahead of the work describes
intent, which drifts from what shipped and has to be reconciled against reality
before an assessor can use it. Either can satisfy the requirement; only one is
accurate by construction.

Where the record claims a capability is complete, that claim is checkable. A
partial state written honestly, with a sentence on the shortfall, reads better in
assessment than a completion that an assessor disproves.

## Change control

An assessor asking how a change was controlled is asking for a durable record
that shows what changed, who approved it, and how it was verified. A pull
request carrying one concern, with review and conversation resolution required,
answers that. A pull request carrying several makes the record of why any one of
them landed ambiguous.

## In this repository

There's no formal traceability record — no requirements table, no linked
issue tracker mapping capability to implementation to test. What exists is
real, structured process discipline for an open-source project of this size:
`.github/pull_request_template.md` requires a description, a change-type
checkbox, a "Related Issues" section (`Fixes #`/`Closes #`/`Resolves #`,
linking the PR to an issue), a changes-made list, and a testing checklist.
`.github/ISSUE_TEMPLATE/` has both bug-report and feature-request forms with
`config.yml` present, giving contributors a structured way to open work
items in the first place — real traceability infrastructure, even without a
formal capability-tracking document behind it.

**One specific gap in that template, worth naming precisely:** the PR
template's testing checklist asks for `npm run lint` and `npm run build`
explicitly, but not `npm test` — even though `npm test` (vitest) is a real,
required CI job (see ChallengeCI). A contributor following the template
literally could check every box without ever having run the unit tests
locally, even though CI would still catch a failure before merge.

`CONTRIBUTING.md` (not audited in depth by this tool file — see its own
content directly) describes the expected contribution flow; whether one PR
per concern, review approval, and conversation resolution are actually
enforced by GitHub branch protection isn't verified — see `profile.yml`'s
`gates` section.

## Evidence

- A traceability record mapping each capability to its implementation and test,
with honest states rather than aspirational ones.
- Change history, one concern per change, with the verification recorded.
- Review approval and conversation resolution enforced rather than optional.

The PR and issue templates are real, structured evidence of intent to trace
changes to issues; no formal capability-to-test mapping exists beyond that.

## Review checklist

- Does the capability that just landed appear in the traceability record,
naming its implementation and its test?
- Does the record claim completion for something only partly built?
- Does the change record say how it was verified?
- Is more than one concern bundled here, making the record of why it landed
ambiguous?
- Was something ambiguous resolved by asking, or by guessing and documenting the
guess?
- **This repository-specific:** does a PR that changes generator or
sanitization logic actually run `npm test` locally, not just the two
boxes the PR template checklist currently asks for? Add the missing
checkbox, or rely on CI catching it — but don't assume the template alone
ensures it.
108 changes: 108 additions & 0 deletions .challengeai/challenge-api.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
# ChallengeAPI

The API: its contract, what it accepts, and what it discloses.

## Covers

API contracts, input validation, interoperability, versioning, rate limiting,
and federal API governance.

## The requirement

- **SP 800-53 Rev. 5** System and Communications Protection (SC) and System and
Information Integrity (SI-10, input validation).
- **The Federal Source Code Policy and API guidance** expect government APIs to
be documented, versioned, and stable for the people who build against them.
- **OMB guidance on open data** expects machine-readable access where the data
is public.

## The contract

An API is a promise to people who cannot be consulted before it changes.

- **Versioned in the path**, moving independently of the product release, so
adding a version does not by itself require a major release of the system.
- **Published as a machine-readable document**, so a client can generate against
it rather than read prose.
- **Validated at the edge of the process.** Every parameter is parsed and
bounded before it reaches a query. Unbounded pagination and unbounded result
sizes are both denial-of-service vectors and cost vectors.
- **Health reported against dependencies**, not just process liveness. A process
that is up and cannot reach its data is not healthy.
- **Rate limited at the edge**, so a burst is refused rather than queued.

## Errors

An error names what was wrong with the request without describing the internals
of the system. A stack trace, a driver message, or a query fragment in a
response body is an information disclosure finding.

## In this repository

There is no API — no backend, no server-rendered route, nothing this
repository exposes over the network beyond the static built site itself
(see ChallengeIaC). Everything runs in the browser: CSV/JSON parsing
(`papaparse`), chart preview rendering, and code generation are all
client-side, with no request/response cycle to a server this project
controls.

**The closest analog worth reviewing with this tool's questions in mind is
the generated code output itself** — what the app hands back to the user as
a download or copy-paste, rather than a network response:

- **Validated at the edge:** partial, and narrower than it looks. Uploaded
CSV gets a real size cap (10MB, strictly enforced) but a weaker type
check than "strict" — `text/plain` is accepted as a valid MIME type
regardless of extension, and any `.csv`-named file is accepted
regardless of its actual MIME type (see ChallengeATO's corrected table).
Sanitization is real but *narrower than "every string value"*:
`CodeGenerator.ts` only calls `sanitizeForCodeGeneration` on `data` and
`selectedOptions` — every other config field, including
`selectedLibrary`, is spread through untouched. `selectedLibrary` can be
set directly from a URL query parameter with no validation
(`Generator.tsx`'s `initializeFromUrl`) and gets interpolated unescaped
into `StaticHTMLGenerator`'s generic-fallback `<title>` tag and script
comments — a real, working XSS path via a crafted share link, not a
theoretical gap. See ChallengeATO for the full writeup and ChallengeEA
for the duplicated-sanitizer angle.
- **Versioning:** not applicable — there's no published contract to version.
The generated code's shape can change between releases with no
compatibility guarantee to anyone who saved output from an earlier version.
- **Published as a machine-readable document:** not applicable — nothing
here is a contract a client generates against.
- **Health / rate limiting:** not applicable — there's no server-side
process or endpoint to report health or limit request rate on. The only
rate-limit-shaped concern is client-side performance on very large CSV
uploads, bounded by the existing 10MB cap.
- **Errors:** `Generator.tsx` surfaces user-facing error messages (invalid
file type, file too large, parse failures) directly in the UI rather than
as an API response — no internal detail (stack traces, file paths) was
found leaking into any user-visible error message during this review.

## Evidence

- The published contract document.
- Route tests exercising the documented behaviour, including failure cases.

Neither applies in the usual sense. `dataTransform.test.ts` (see
ChallengeTDD) is the closest evidence this repository has for validated
behavior, though it covers data transformation, not the sanitization path
specifically.

## Review checklist

- Is every new parameter validated and bounded?
- Does a new route appear in the contract document in the same pull request?
- Does an error response leak an internal detail?
- Does a change alter the shape of an existing response? That is a breaking
change to a published contract, and it belongs behind a new version.
- Is anything personal placed in a URL, where it reaches logs and history?
**This repository already does this** — "Copy Share Link" puts the
entire dataset in the URL query string (see ChallengeSQL, ChallengeATO).
- **This repository-specific:** does every config field that reaches a
generator's template output — not just `data` and `selectedOptions`, but
`selectedLibrary`, `selectedOutputFormat`, and any field added later —
get routed through `sanitizeForCodeGeneration` before interpolation? The
current gap (`selectedLibrary` bypassing it entirely, reachable via a URL
parameter) shows that "sanitize before generating" isn't yet a rule
applied to the whole config object, just to two fields of it.
99 changes: 99 additions & 0 deletions .challengeai/challenge-ato.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
# ChallengeATO

Authorization readiness: whether this system could be assessed today and whether
the evidence would hold.

## Covers

Control traceability, evidence sufficiency, the security package, the privacy
assessment, and bounded penetration testing against a running instance.

## The requirement

- **FISMA** requires federal systems to implement an information security
program and to be authorized before operating.
- **NIST SP 800-37 Rev. 2** defines the Risk Management Framework.
- **FIPS 199** categorizes the system as low, moderate, or high impact. The
categorization drives everything downstream, and it is recorded in
`profile.yml`.
- **NIST SP 800-53 Rev. 5** provides the control baseline the categorization
selects.
- **NIST SP 800-53A Rev. 5** provides the procedures an assessor uses.
- **The E-Government Act** requires a Privacy Impact Assessment where a
system handles information **in identifiable form** — not merely
information "about" individuals in aggregate — or initiates a qualifying
electronic collection from ten or more non-federal persons.

## The package

A security package is a set, numbered so it can be handed over as one. The
System Security Plan describes the system, its boundary, and how each control is
met. Around it sit the assessment plan and report, the plan of action and
milestones, the risk assessment, the contingency and incident response plans,
configuration management, access control policy, continuous monitoring, and the
privacy impact assessment.

The boundary described in the SSP is the boundary the system actually
provisions. Anything outside it is inherited from the provider's own
authorization and is named as inherited rather than claimed.

## In this repository

No security package exists — no `docs/security/`, no SSP, no FIPS 199
categorization, no POA&M, no PIA. `profile.yml` records `control_baseline:
null` and `impact_level: null` because none is pursued, which is
appropriate for a tool with no accounts, no persistence, and no PII (see
ChallengeSQL) — not a gap this system needs to close, unlike a system that
actually handles federal data.

What's worth checking instead is `SECURITY.md`'s own "Security Measures"
list, since it makes specific, checkable claims rather than staying generic:

| Claim (`SECURITY.md`) | What was actually found |
|---|---|
| "Input Sanitization: All user inputs are sanitized to prevent XSS attacks" | **Not fully true — withdrawn.** `sanitizeForCodeGeneration` genuinely escapes `data` and `selectedOptions` before code generation (`CodeGenerator.ts` calls it on exactly those two fields), but `selectedLibrary` and `selectedOutputFormat` are spread into the generator config untouched (`...config` in `CodeGenerator.ts`) and never sanitized at all. `Generator.tsx`'s URL-restore logic (`initializeFromUrl`) sets `selectedLibrary` directly from the `lib` query parameter with no validation against the known library list. An unrecognized `lib` value falls through `StaticHTMLGenerator`'s switch to `generateGenericStaticHTML`, which interpolates `selectedLibrary` unescaped into `<title>Drupal Data Visualization - ${selectedLibrary}</title>` and into script comments. A crafted share link — `/generator?lib=</title><script>...</script>&...` — makes that markup live in the downloaded HTML. Real, working XSS via a URL parameter this claim says is covered. See ChallengeAPI and ChallengeEA for the fix this points to. |
| "File Upload Validation: Strict file type and size validation for uploads" | Partially true. The 10MB size cap is real and strictly enforced. The type check is weaker than "strict," though: `Generator.tsx` accepts a file if its extension is `.csv` **or** its reported MIME type is `text/csv`, `application/csv`, **or `text/plain`** — rejection only happens if both checks fail. `text/plain` is what browsers report for a wide range of non-CSV files, so a non-CSV file reported as `text/plain` passes regardless of extension, and any file merely named `something.csv` passes regardless of its actual MIME type or content. Not "strict," and not "requires both" — it's an either/or check with a permissive MIME option. |
| "Content Security Policy: CSP headers to prevent code injection" | Not backed. No `[[headers]]` block in `netlify.toml`, no `<meta http-equiv="Content-Security-Policy">` in `index.html`, no CSP configuration found anywhere in this repository. |
| "Dependency Management: Regular updates of dependencies with security patches" | Partially backed by process, not by a schedule: `ci.yml`'s `security` job runs `npm audit` on every push/PR (real, and blocking — see ChallengeCI), which would surface a new advisory quickly. There's no Dependabot or Renovate configuration (no `.github/dependabot.yml`) automating the updates themselves, though — the audit finds problems; nothing here fixes them automatically. |
| "Secure Defaults: Secure configuration by default" | Too generic to check against anything specific. |
| "HTTPS Enforcement: All traffic encrypted in transit" | Real in practice — Netlify serves everything over HTTPS by default — but that's Netlify's platform behavior, not a setting this repository configures or could disable if it wanted to. Reasonable to claim, worth knowing it's inherited rather than enforced by this codebase. |

`README.md` carries a `Security: SOC 2` badge linking to
`netlify.com/security` — that's **Netlify's own SOC 2 Type II report about
its hosting platform**, not an audit of this application's code or
generated output. Presented alongside this project's own badges without
that distinction, a reader would reasonably assume it describes this
project. See ChallengeUI for the parallel check on the accessibility badge.

## Evidence

- Pipeline output produced on every run: scan results, test reports, and a
component inventory.
- A traceability record mapping each capability to its implementation and test.
- A reviewable record of the deployed configuration, so the boundary can be
checked rather than taken on description.

CI produces real scan output (`npm audit`, on every run — see ChallengeCI)
and a real test run (16 unit tests in `dataTransform.test.ts` — see
ChallengeTDD), which is more than several other repos in this suite
produce. No component inventory (SBOM) and no traceability record exist.

## Review checklist

- Does the SSP describe the system as it is now, or as it was at the last
release?
- Is every control claim backed by something a reader can open?
- Does the boundary in the SSP match what is actually provisioned?
- Has the data model started handling personal information the PIA does not
mention?
- **A weakness, POA&M item, or vulnerability goes into a security document only
with explicit approval.** Cataloguing theoretical gaps manufactures a record
of insecurity. State what is implemented; independent assessment produces the
authoritative findings. This repository has no security document to add a
finding to — the table above states what was independently checked
against the actual code, not a fabricated POA&M.
- **This repository-specific:** before adding a new claim to `SECURITY.md`
or a new badge to `README.md`, can it be traced to real, checkable code
or configuration in this repository — and if it describes a third
party's own certification (like Netlify's SOC 2), is that attribution
clear rather than implied to be this project's own?
Loading
Loading