Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
200 commits
Select commit Hold shift + click to select a range
064973e
feat: evidence-grade NIM discovery + all-modality cost-quality benchm…
seonghobae Aug 4, 2026
b48f396
merge: stack NIM benchmark on provider-egress security base
seonghobae Aug 4, 2026
0e19b3a
docs(changelog): record NIM benchmark evidence slice
seonghobae Aug 4, 2026
66c8d55
ci: apply reviewed NIM transport hardening
seonghobae Aug 4, 2026
9d546a7
fix(nim): harden egress and equalize policy budgets
seonghobae Aug 4, 2026
785fb15
fix(nim): install evidence and transport hardening
seonghobae Aug 4, 2026
e07a504
test(nim): prove secure transport and equal budgets
seonghobae Aug 4, 2026
d0312ea
docs(nim): document pinned egress and equal total budgets
seonghobae Aug 4, 2026
c7bcad2
chore(ci): remove temporary NIM repair workflow
seonghobae Aug 4, 2026
e1ae745
chore(ci): stage one-time NIM source repair
seonghobae Aug 4, 2026
1015608
chore(ci): remove privileged one-shot NIM repair workflow
seonghobae Aug 4, 2026
07128ac
chore(ci): stage exact-head NIM security repair
seonghobae Aug 4, 2026
a8393d7
chore(ci): harden bounded NIM repair job
seonghobae Aug 4, 2026
4cc0f55
chore(ci): stage bounded NIM source integration
seonghobae Aug 4, 2026
9ad742a
chore(ci): remove overlapping one-shot NIM workflow
seonghobae Aug 4, 2026
4fd96f1
fix(ci): replace write-capable PR repair with isolated publisher
seonghobae Aug 4, 2026
4bb7c85
docs(changelog): record reviewed NIM access evidence
seonghobae Aug 4, 2026
d9bedd7
chore(ci): stage equal-budget regression repair
seonghobae Aug 4, 2026
c4d47d2
ci: complete bounded NIM source integration
seonghobae Aug 4, 2026
db6ebb9
chore(ci): remove redundant NIM test-only repair workflow
seonghobae Aug 4, 2026
dba1d7e
ci: run bounded NIM source integration from PR checks
seonghobae Aug 4, 2026
a3741b5
fix(ci): execute bounded NIM source transformation deterministically
seonghobae Aug 4, 2026
18d3fe7
fix(ci): preserve embedded transformation literals
seonghobae Aug 4, 2026
27eb1b8
fix(security): remove PR-head OIDC source mutation
seonghobae Aug 4, 2026
ba6b34b
fix(security): delete privileged branch-trigger repair workflow
seonghobae Aug 4, 2026
e4578da
fix(security): delete branch-controlled OIDC repair source
seonghobae Aug 4, 2026
f80e503
test(security): pin PR workflow OIDC boundary
seonghobae Aug 4, 2026
bc77bfc
ci(review): export read-only exact-head workspace
seonghobae Aug 4, 2026
a71ebbc
ci(review): include inert transformation evidence
seonghobae Aug 4, 2026
6df72ee
test(ci): require secretless NIM dry-run job
seonghobae Aug 5, 2026
09fd519
fix(ci): isolate NIM dry runs from provider secrets
seonghobae Aug 5, 2026
2b622bd
docs(security): record NIM workflow secret isolation
seonghobae Aug 5, 2026
40b48ba
docs(changelog): record secretless NIM dry runs
seonghobae Aug 5, 2026
348e5e6
docs(benchmark): explain workflow credential isolation
seonghobae Aug 5, 2026
667b311
chore(review): stage verified NIM repair patch
seonghobae Aug 5, 2026
2498648
ci(nim): stage reviewed patch part 1 of 7
seonghobae Aug 5, 2026
4552084
ci(nim): stage reviewed patch part 2 of 7
seonghobae Aug 5, 2026
45c151e
ci(nim): stage reviewed patch part 3 of 7
seonghobae Aug 5, 2026
9dac6b2
ci(nim): stage reviewed patch part 4 of 7
seonghobae Aug 5, 2026
fecd25f
ci(nim): stage reviewed patch part 5 of 7
seonghobae Aug 5, 2026
19be572
ci(nim): stage reviewed patch part 6 of 7
seonghobae Aug 5, 2026
736893c
ci(nim): stage reviewed patch part 7 of 7
seonghobae Aug 5, 2026
07f28a9
ci(nim): apply exact reviewed benchmark integration
seonghobae Aug 5, 2026
7d5d6aa
ci: retrigger reviewed NIM integration with scoped permissions
seonghobae Aug 5, 2026
fd6b074
ci: verify and export reviewed NIM integration read-only
seonghobae Aug 5, 2026
cb7e9b3
ci: bind NIM verification to security ancestor
seonghobae Aug 5, 2026
3878e5b
fix(ci): bind NIM verification to pull-request head
seonghobae Aug 5, 2026
5bbedc2
fix(ci): reconcile reviewed NIM patch per file
seonghobae Aug 5, 2026
d98fe6c
fix(ci): reconstruct reviewed NIM tree from patch base
seonghobae Aug 5, 2026
428736c
fix(ci): bind reviewed NIM reconstruction to exact head
seonghobae Aug 5, 2026
e4e03c4
fix(ci): run reviewed NIM verification read-only on push
seonghobae Aug 5, 2026
6cd2012
fix(ci): use runner temp variable at execution time
seonghobae Aug 5, 2026
7595919
fix(ci): discover the exact reviewed patch base
seonghobae Aug 5, 2026
a93ec20
chore(ci): inspect the staged NIM integration payload
seonghobae Aug 5, 2026
f11193d
fix(ci): replay the reviewed NIM patch stack
seonghobae Aug 5, 2026
83b7906
fix(ci): rehydrate the reviewed patch index
seonghobae Aug 5, 2026
e9bfb85
fix(ci): execute the patch rehydration helper
seonghobae Aug 5, 2026
88274c0
fix(ci): apply the reviewed patch against the rehydrated index
seonghobae Aug 5, 2026
03b30e5
fix(ci): verify the reconstructed reviewed tree
seonghobae Aug 5, 2026
e256a9a
fix(ci): locate the exact reviewed patch context
seonghobae Aug 5, 2026
012f7c7
fix(ci): build the wheel in an isolated backend environment
seonghobae Aug 5, 2026
dde856f
ci(nim): publish the authenticated verified source tree
seonghobae Aug 5, 2026
8049ca8
feat(nim): integrate verified provider-neutral benchmark
seonghobae Aug 5, 2026
d1516a5
docs(nim): record immutable integration provenance
seonghobae Aug 5, 2026
f090ffc
test(ci): require stacked pull-request coverage
seonghobae Aug 5, 2026
d091b69
ci: validate stacked pull requests
seonghobae Aug 5, 2026
0d66e0a
ci: fuzz stacked pull requests
seonghobae Aug 5, 2026
928214c
ci: secure stacked pull requests
seonghobae Aug 5, 2026
c759cd1
test(ci): require isolated wheel builds
seonghobae Aug 5, 2026
8e19d64
fix(ci): isolate benchmark wheel builds
seonghobae Aug 5, 2026
e7f7238
ci(review): export read-only PR90 workspace
seonghobae Aug 5, 2026
c72f1cb
test(benchmark): stage complete-budget preflight red tests part 1
seonghobae Aug 5, 2026
cb20657
test(benchmark): stage complete-budget preflight red tests part 2
seonghobae Aug 5, 2026
636c3e0
fix(benchmark): stage complete request planner part 3
seonghobae Aug 5, 2026
1053f19
docs(benchmark): stage complete-budget evidence and verification part 4
seonghobae Aug 5, 2026
638a126
ci(benchmark): apply complete-budget preflight test-first
seonghobae Aug 5, 2026
1f6da53
ci(review): stage exact PR90 request-plan patch
seonghobae Aug 5, 2026
8a31de7
ci(review): refresh exact PR90 request-plan patch
seonghobae Aug 5, 2026
87496c8
ci(review): verify and publish exact PR90 request-plan fix
seonghobae Aug 5, 2026
b3c0471
ci(review): inspect exact PR90 patch hash
seonghobae Aug 5, 2026
2efcf08
ci(review): export decoded PR90 patch evidence
seonghobae Aug 5, 2026
a06f9c5
ci(review): verify exact staged PR90 patch
seonghobae Aug 5, 2026
77cc8cc
ci(review): publish verified PR90 request-plan fix
seonghobae Aug 5, 2026
796e6e9
test(benchmark): align scheduled complete-plan contract
seonghobae Aug 5, 2026
5956214
test(benchmark): require request-plan provenance schema
seonghobae Aug 5, 2026
9b72816
fix(nim): require a complete catalog benchmark plan
seonghobae Aug 5, 2026
a2af363
merge: integrate provider-egress security base into NIM benchmark
github-actions[bot] Aug 5, 2026
c5b3876
docs(adr): record exact NIM security integration evidence
seonghobae Aug 5, 2026
d43adaa
test(benchmark): expose undersized locked evaluation manifest
seonghobae Aug 5, 2026
16b5016
feat(benchmark): expand locked manifest to evidence floor
seonghobae Aug 5, 2026
ac13a56
test(benchmark): bind evidence floor to existing hard ceiling
seonghobae Aug 5, 2026
62e0e59
docs(benchmark): record thirty-task evidence floor
seonghobae Aug 5, 2026
43c629f
docs(benchmark): align request plan and evidence floor
seonghobae Aug 5, 2026
ff51584
docs(benchmark): doctor thirty-task evidence floor
seonghobae Aug 5, 2026
bec1129
test(benchmark): expose undersized default request caps
seonghobae Aug 5, 2026
cdfe08e
ci(benchmark): test and repair default request caps
seonghobae Aug 5, 2026
37c78b4
ci(benchmark): rearm default request-cap repair
seonghobae Aug 5, 2026
545cd64
ci(benchmark): add exact-head reopen trigger
seonghobae Aug 5, 2026
e56d1a3
ci(benchmark): use resolvable checkout action
seonghobae Aug 5, 2026
5216171
ci(benchmark): repair complete-plan default finalizer
seonghobae Aug 5, 2026
8cae8d4
test(benchmark): bind evidence floor to canonical request planner
seonghobae Aug 5, 2026
e31f901
ci(benchmark): cover all complete-plan regression anchors
seonghobae Aug 5, 2026
da8ec5f
ci(benchmark): repair invalid one-shot YAML
seonghobae Aug 5, 2026
1b97e48
ci(benchmark): split workflow-authorized publication
seonghobae Aug 5, 2026
0ebb1d5
fix(benchmark): align complete-plan source evidence
github-actions[bot] Aug 5, 2026
9e4b50b
fix(benchmark): align manual request cap with complete plan
seonghobae Aug 5, 2026
821b58e
ci(benchmark): remove completed one-shot workflow
seonghobae Aug 5, 2026
7b23bac
test(nim): cover buyer-facing complete request plan
seonghobae Aug 5, 2026
5c1a3f6
ci(nim): include request-plan regression in quality gate
seonghobae Aug 5, 2026
f5a494f
test: require NIM parser Atheris instrumentation
seonghobae Aug 5, 2026
48f007d
fix: instrument NIM parser imports for Atheris
seonghobae Aug 5, 2026
2736113
docs: distinguish historical NIM integration evidence
seonghobae Aug 5, 2026
b60ec5d
test: cover NIM review regressions
seonghobae Aug 5, 2026
bdc9673
chore: defer broader NIM regression batch
seonghobae Aug 5, 2026
646c417
docs: make NIM receipt head-stable
seonghobae Aug 5, 2026
318feb9
docs: record NIM fuzz instrumentation repair
seonghobae Aug 5, 2026
63438d0
test(nim): require completion-dependent README evidence status
seonghobae Aug 5, 2026
6fc53f1
docs(nim): clarify completion-dependent evidence status
seonghobae Aug 5, 2026
4a79964
test(nim): normalize README contract whitespace
seonghobae Aug 5, 2026
1217fc9
docs(changelog): record NIM evidence-status clarification
seonghobae Aug 5, 2026
11353d4
test(security): inspect every pull-request workflow structurally
seonghobae Aug 5, 2026
67c5c84
test(nim): enforce ordered non-empty secret-isolation steps
seonghobae Aug 5, 2026
989d7fe
test(security): parse pull-request branch filters structurally
seonghobae Aug 5, 2026
4329081
test(nim): lock review regressions before repair
seonghobae Aug 5, 2026
715ee76
fix(fuzz): require normalized NIM catalog depth failures
seonghobae Aug 5, 2026
8c4ee86
test: skip fixture-bound standalone NIM cases
seonghobae Aug 5, 2026
4ca0448
fix: close NIM evidence review gaps
seonghobae Aug 5, 2026
aed3e7c
docs: record NIM evidence hardening
seonghobae Aug 5, 2026
a5d529b
test: require NIM review regressions in coverage gate
seonghobae Aug 5, 2026
8603bcf
ci: cover current NIM review regressions
seonghobae Aug 5, 2026
e9f2072
docs: record exact NIM regression gate
seonghobae Aug 5, 2026
ffd03ba
test(nim): require complete CSV model-assignment evidence
seonghobae Aug 5, 2026
a9bf8bc
feat(nim): preserve model assignment evidence in CSV artifacts
seonghobae Aug 5, 2026
d031f53
feat(nim): fail closed until CSV evidence is complete
seonghobae Aug 5, 2026
554a220
ci(nim): gate CSV evidence adapter at 100% coverage
seonghobae Aug 5, 2026
2db0737
docs(nim): record CSV assignment evidence parity
seonghobae Aug 5, 2026
dfc300a
docs(nim): doctor CSV assignment evidence boundary
seonghobae Aug 5, 2026
5c2c695
refactor(nim): preserve CSV mode without cleanup branches
seonghobae Aug 5, 2026
7dd5245
test(nim): cover CSV evidence fail-closed edge cases
seonghobae Aug 5, 2026
816707f
ci(nim): execute CSV evidence edge regressions
seonghobae Aug 5, 2026
00483f3
test(nim): make routing disclaimer contract case-insensitive
seonghobae Aug 5, 2026
88b0e96
ci(nim): preserve benchmark docstring contract and gate adapter
seonghobae Aug 5, 2026
c9fa27c
test(ci): require pull request workflows to checkout exact heads
seonghobae Aug 5, 2026
5e46772
ci: checkout exact pull request heads in test jobs
seonghobae Aug 5, 2026
f7320aa
ci: checkout exact pull request heads in fuzz jobs
seonghobae Aug 5, 2026
16f1409
ci: checkout exact pull request heads in security jobs
seonghobae Aug 5, 2026
692e49a
docs(ci): record exact-head pull request verification boundary
seonghobae Aug 5, 2026
1ef2c97
docs(ci): doctor exact-head versus merge-test evidence
seonghobae Aug 5, 2026
f042953
merge: synchronize NIM benchmark stack with PR #96
seonghobae Aug 5, 2026
2b84e1e
test(nim): define transactional artifact publication contract
seonghobae Aug 6, 2026
ef28784
fix(nim): publish complete evidence sets transactionally
seonghobae Aug 6, 2026
ed52bc8
test(nim): route CSV wrapper fakes through private staging
seonghobae Aug 6, 2026
5771957
ci(nim): include transactional publication regressions
seonghobae Aug 6, 2026
e1f19ff
test(nim): write fake artifacts into precreated staging
seonghobae Aug 6, 2026
c446251
test(nim): reuse wrapper-created staging directory
seonghobae Aug 6, 2026
30c046f
test(nim): cover transactional publication failure edges
seonghobae Aug 6, 2026
6ed1fbf
ci(nim): include publication edge coverage
seonghobae Aug 6, 2026
5e46ba5
test(nim): cover duplicate split output argument
seonghobae Aug 6, 2026
1cea0c0
docs(nim): record transactional evidence-set publication
seonghobae Aug 6, 2026
3373561
test(benchmark): require strict complete-answer scorers
seonghobae Aug 7, 2026
9e92b1e
feat(benchmark): add explicit strict scoring composition root
seonghobae Aug 7, 2026
a26195c
test(benchmark): cover explicit strict scoring lifecycle
seonghobae Aug 7, 2026
d3a0dee
feat(benchmark): activate strict scoring in supported CLI
seonghobae Aug 7, 2026
e5699f9
test(benchmark): enforce strict scorer coverage and packaging
seonghobae Aug 7, 2026
d255a3d
fix(benchmark): validate derived strict manifest through public loader
seonghobae Aug 7, 2026
e87fe54
docs(benchmark): record strict answer-scoring validity boundary
seonghobae Aug 7, 2026
3ba5977
docs(changelog): record strict NIM scoring evidence
seonghobae Aug 7, 2026
e052426
refactor(benchmark): keep strict scorer branches fully testable
seonghobae Aug 7, 2026
fd75801
test(benchmark): verify strict scoring through supported publication CLI
seonghobae Aug 7, 2026
08b89e6
test(benchmark): run strict scoring integration in quality gate
seonghobae Aug 7, 2026
5ed6f82
test(benchmark): preserve lazy optional-import boundary
seonghobae Aug 7, 2026
c1388ec
test(ci): preserve stacked exact-head workflow contract
seonghobae Aug 7, 2026
01b17d2
docs(changelog): preserve current security-base evidence
seonghobae Aug 7, 2026
a9a37e2
docs(changelog): make stacked-base reconciliation conflict-free
seonghobae Aug 7, 2026
9780f1b
test(benchmark): require task-specific strict text semantics
seonghobae Aug 7, 2026
33fec53
feat(benchmark): preserve task-specific strict text semantics
seonghobae Aug 7, 2026
720556d
fix(benchmark): declare strict aliases and disambiguating context
seonghobae Aug 7, 2026
28bb2a9
docs(benchmark): document task-specific strict scoring semantics
seonghobae Aug 7, 2026
a235459
docs(changelog): record task-specific strict scorer semantics
seonghobae Aug 7, 2026
3c71220
test(benchmark): bound strict scorer input resources
seonghobae Aug 7, 2026
f584651
fix(benchmark): bound strict scorer input resources
seonghobae Aug 7, 2026
1d58f89
test(benchmark): run strict scorer resource bounds in quality gate
seonghobae Aug 7, 2026
b6eadde
docs(benchmark): record strict scorer resource boundary
seonghobae Aug 7, 2026
2384230
docs(changelog): record strict scorer resource cap
seonghobae Aug 7, 2026
7084325
test(benchmark): require named deterministic clock seams
seonghobae Aug 7, 2026
63c7d22
test(benchmark): keep review regressions behavior-focused
seonghobae Aug 7, 2026
b2a0091
test(benchmark): reject strict answer leakage before egress
seonghobae Aug 7, 2026
07f731b
fix(benchmark): reject strict answer leakage before egress
seonghobae Aug 7, 2026
53ba32a
test(benchmark): cover strict leakage failure boundaries
seonghobae Aug 7, 2026
2a471f2
test(benchmark): include strict leakage in permanent quality gate
seonghobae Aug 7, 2026
af03057
docs(benchmark): record strict prompt leakage boundary
seonghobae Aug 7, 2026
fcf39c4
docs(changelog): record strict prompt leakage rejection
seonghobae Aug 7, 2026
0bbd869
test(benchmark): expose strict publication failure evidence
seonghobae Aug 7, 2026
99ea9ab
test(benchmark): require deterministic assignment step IDs
seonghobae Aug 7, 2026
bf204e0
fix(benchmark): canonicalize empty assignment step IDs
seonghobae Aug 7, 2026
7c8629e
test(benchmark): cover assignment step-ID collision recovery
seonghobae Aug 7, 2026
2afc383
test(benchmark): include assignment step-ID contracts in quality gate
seonghobae Aug 7, 2026
2b3fc0a
fix(benchmark): preserve runtime trace step identities
seonghobae Aug 7, 2026
1503a3c
test(benchmark): cover runtime assignment step IDs
seonghobae Aug 7, 2026
26f8d8d
test(scoring): cover unrepresentable answer-key exponents
seonghobae Aug 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion .github/workflows/fuzz.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]
schedule:
# Weekly deeper run (longer per-target budget via FUZZ_SECONDS).
- cron: "41 4 * * 2"
Expand All @@ -26,6 +25,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand All @@ -50,6 +50,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand Down Expand Up @@ -83,6 +84,9 @@ jobs:
- name: Fuzz orchestration engine
run: python fuzz/fuzz_orchestration.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/orchestration

- name: Fuzz NIM model-catalog parser
run: python fuzz/fuzz_nim_catalog.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/nim_catalog

- name: Upload crash artifacts
if: failure()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
Expand Down
149 changes: 149 additions & 0 deletions .github/workflows/nim-benchmark.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
name: NIM benchmark

# Evidence-grade NVIDIA NIM model discovery and cost-quality benchmark.
# Manual dispatch defaults to a deterministic dry run. Monthly scheduled runs
# are live but use a conservative hard request cap. Dry execution receives no
# provider credential; only the live step can read NVIDIA_NIM_API_KEY.

on:
workflow_dispatch:
inputs:
dry_run:
description: "Dry run without contacting NVIDIA"
type: boolean
default: true
max_total_requests:
description: "Hard cap on provider requests for this run"
type: number
default: 2000
pricing_scenario:
description: "Optional reviewed pricing-scenario JSON path"
type: string
default: ""
schedule:
- cron: "23 3 5 * *"

permissions:
contents: read

concurrency:
group: nim-benchmark
cancel-in-progress: false

jobs:
dry_run_benchmark:
name: Deterministic NIM benchmark dry run
if: github.event_name == 'workflow_dispatch' && inputs.dry_run == true
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Run dry benchmark
env:
MAX_REQUESTS: ${{ inputs.max_total_requests }}
PRICING_SCENARIO: ${{ inputs.pricing_scenario }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
set -euo pipefail
extra_args=()
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
--dry-run \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload dry-run artifacts
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-dry-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 30
if-no-files-found: error

live_benchmark:
name: Live NIM catalog benchmark
if: github.event_name == 'schedule' || (github.event_name == 'workflow_dispatch' && inputs.dry_run != true)
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Resolve live parameters
id: live_params
env:
EVENT_NAME: ${{ github.event_name }}
INPUT_MAX_REQUESTS: ${{ inputs.max_total_requests }}
INPUT_PRICING: ${{ inputs.pricing_scenario }}
run: |
set -euo pipefail
if [ "$EVENT_NAME" = "schedule" ]; then
echo "max_requests=2000" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=" >> "$GITHUB_OUTPUT"
else
echo "max_requests=${INPUT_MAX_REQUESTS}" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=${INPUT_PRICING}" >> "$GITHUB_OUTPUT"
fi

- name: Run live benchmark
env:
NVIDIA_NIM_API_KEY: ${{ secrets.NVIDIA_NIM_API_KEY }}
MAX_REQUESTS: ${{ steps.live_params.outputs.max_requests }}
PRICING_SCENARIO: ${{ steps.live_params.outputs.pricing_scenario }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
set -euo pipefail
extra_args=()
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload live benchmark artifacts
if: always()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-live-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 90
if-no-files-found: error
3 changes: 2 additions & 1 deletion .github/workflows/security.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]
schedule:
- cron: "17 3 * * 1"
workflow_dispatch:
Expand All @@ -34,6 +33,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Initialize CodeQL
Expand All @@ -54,6 +54,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand Down
62 changes: 61 additions & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]

permissions:
contents: read
Expand All @@ -21,6 +20,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand All @@ -36,3 +36,63 @@ jobs:

- name: Run full test suite
run: python -m pytest -q

nim_benchmark_quality:
name: NIM benchmark coverage, docstrings, and package smoke
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install hash-locked quality tools
run: python -m pip install --require-hashes -r requirements-opencode-review-ci.txt

- name: Prove complete benchmark coverage and public docstrings
run: |
set -euo pipefail
python -m coverage erase
python -m coverage run --branch \
--source=contextual_orchestrator.nim_benchmark,contextual_orchestrator.nim_csv_evidence,contextual_orchestrator.nim_strict_scoring \
-m pytest \
tests/test_nim_benchmark.py \
tests/test_nim_benchmark_budget_view.py \
tests/test_nim_benchmark_release_acceptance.py \
tests/test_nim_benchmark_review_regressions.py \
tests/test_nim_benchmark_workflow_contract.py \
tests/test_nim_artifact_publication.py \
tests/test_nim_artifact_publication_edges.py \
tests/test_nim_assignment_step_ids.py \
tests/test_nim_csv_evidence.py \
tests/test_nim_csv_evidence_edges.py \
tests/test_nim_strict_scorer_validity.py \
tests/test_nim_strict_scoring_bounds.py \
tests/test_nim_strict_scoring_integration.py \
tests/test_nim_strict_scoring_leakage.py \
-q
python -m coverage report \
--include=contextual_orchestrator/nim_benchmark.py,contextual_orchestrator/nim_csv_evidence.py,contextual_orchestrator/nim_strict_scoring.py \
--show-missing \
--fail-under=100
python -m interrogate -f 100 contextual_orchestrator/nim_benchmark.py
python -m interrogate -f 100 contextual_orchestrator/nim_csv_evidence.py
python -m interrogate -f 100 contextual_orchestrator/nim_strict_scoring.py

- name: Build, install, and import the wheel
run: |
set -euo pipefail
rm -rf dist "$RUNNER_TEMP/nim-wheel-site"
python -m pip wheel --no-deps . --wheel-dir dist
python -m pip install --no-deps \
--target "$RUNNER_TEMP/nim-wheel-site" \
dist/contextual_orchestrator-*.whl
cd "$RUNNER_TEMP"
PYTHONPATH="$RUNNER_TEMP/nim-wheel-site" \
python -c "import contextual_orchestrator; import contextual_orchestrator.nim_benchmark; import contextual_orchestrator.nim_csv_evidence; import contextual_orchestrator.nim_strict_scoring"
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,9 @@ tempcred.txt

# hypothesis fuzzing DB
.hypothesis/

# coverage measurement data
.coverage

# local benchmark artifacts (uploaded by CI, not committed)
benchmark_artifacts/
12 changes: 12 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,19 +6,31 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and

## [Unreleased]

### Added

- Add an optional provider-neutral NVIDIA NIM benchmark harness that dynamically discovers the live `/v1/models` catalog, probes every discovered model under bounded concurrency and a complete hard request plan, and compares every eligible direct model, route-once, bounded-conduct, and reviewed pricing-scenario policy under equal total token and call envelopes.
- Add deterministic no-egress dry runs, valid media fixtures, secret-redacted transactional JSON/CSV/Markdown artifacts, complete role/agent/model assignment evidence, paired bootstrap uncertainty, evidence sufficiency, and quality-latency and reviewed-hypothetical-cost Pareto frontiers.
- Add validation-time public-address-pinned provider transport, original-host authority/SNI/certificate verification, redirect and ambient-proxy rejection, credential-forwarding prevention, 8 MiB response bounds, and live-secret isolation to the NIM benchmark path.
- Add versioned strict complete-answer scoring for locked tasks: finite full-response decimal comparison, NFC-normalized declared text alternatives with task-specific case semantics, explicit alias and prompt-disambiguation evidence, bounded value and prompt inputs, decimal-equivalent and declared-alias prompt-leakage rejection, complete token boundaries that avoid larger-word false positives, zero-score handling for oversized or unrepresentable model output, fail-before-egress manifest validation, derived-manifest provenance, and lazy optional-adapter activation.
- Add permanent benchmark contracts for 100% production statement/branch coverage, 100% public docstrings, fuzzing, package build/install/import, exact contributor-head workflows, evidence-status semantics, and release acceptance without automatic routing authorization.

### Security

- Restrict the private plain-HTTP provider seam to `localhost` or literal loopback IP addresses, reject URL userinfo before connection, dial directly without ambient proxy lookup, reject all redirect responses, and close failed resources deterministically.
- Pin each HTTPS provider connection to the exact public addresses approved during validation, preserve the original hostname for TLS verification, bypass environment proxy resolution, and reject redirects to close DNS-rebinding and credential-forwarding SSRF paths.
- Bound every provider response to 8 MiB of cumulative consumed bytes, including SSE iteration, reject oversized declared lengths before body consumption, fail closed on malformed or conflicting `Content-Length` and ambiguous `Content-Length` plus `Transfer-Encoding`, redact header-inspection failures, and never silently truncate an untrusted response.
- Integrate DNS-pinned provider dispatch directly into `ModelClient` so package import performs no optional-adapter monkey-patching or order-dependent class mutation.
- Reject provider hosts that resolve to any non-globally-routable address, including RFC 6598 shared address space, while retaining explicit multicast, private, loopback, link-local, and reserved-address protections.
- Document narrowly scoped Semgrep suppressions for parameter-bound database queries, the explicit development-only TLS verification opt-out, and provider URLs that pass the egress guard.

### Changed

- Pin Atheris by Python interpreter so the Python 3.11 fuzz job and the newer central coverage-evidence image both install a published, hash-locked wheel.
- Run repository Tests, Fuzz, and Security workflows for stacked pull requests targeting any branch, bind every checkout to the literal contributor-head SHA, and keep checkout credentials non-persistent so local evidence cannot silently become absent or synthetic-merge-only evidence.

### Documentation

- Add APA 7 doctoring for Python environment-marker semantics, Atheris artifact availability and hashes, and the supported-platform uncertainty boundary.
- Add provider-response resource-bound doctoring covering the 8 MiB fail-closed limit, HTTP framing preflight, bounded SSE reads, batch-output partitioning, incident handling, and operational rollback.
- Add pull-request exact-head workflow doctoring covering stacked-base support, contributor-head identity, untrusted-code execution, merge-tree separation, cancellation handling, and rollback.
- Record the CI trust boundary between generic coverage and native fuzz execution, including the evidence-preserving retry rule for branch-referenced reusable workflows.
30 changes: 29 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -199,9 +199,36 @@ is read from a **KV config store**, never `os.getenv`.
`pg_tiktoken` counting, and the production batch backend without adding a
repository split here.

Grounding papers (LLM cost, routing, load balancing) live in
Grounding papers (LLM cost, routing, load balancing, evaluation) live in
[docs/papers](docs/papers/README.md) with citations.

### NIM cost-quality benchmark (optional harness)

Evidence-grade benchmark of the routing policies against a **dynamically
discovered** NVIDIA NIM catalog. It probes chat, completions, Responses,
embeddings, image/video/audio understanding, transcription, and speech; compares
direct, route-once, and bounded-conduct cells under one equal total-token and
call budget; records paired uncertainty and Pareto frontiers; and keeps reviewed
actual endpoint-access evidence separate from optional hypothetical paid rates.
The bundled manifest contains thirty locked tasks. It may reach
`evidence_review_required` when at least 90% of policy-task cells complete and
the paired-task floor is met; otherwise it reports `insufficient_evidence`. No
benchmark artifact automatically changes production routing.

The adapter is lazy and optional: ordinary `import contextual_orchestrator` does
not import or mutate it. Deterministic `--dry-run` receives no network access or
NVIDIA secret. Live execution resolves `NVIDIA_NIM_API_KEY` from the credential
registry, pins HTTPS connections to validation-time public addresses, rejects
redirects and proxy routing, and fails closed on missing/expired evidence. See
[docs/nim_benchmark.md](docs/nim_benchmark.md) and the
[engineering decision record](docs/doctoring/nim-benchmark-evidence-grade.md).

```bash
python -m contextual_orchestrator nim-benchmark --dry-run \
--pricing-scenario examples/nim_pricing_scenario.json \
--output-dir benchmark_artifacts
```

## Design Artifacts

- [Library research](docs/library_research.md)
Expand Down Expand Up @@ -255,6 +282,7 @@ python tests/test_paper_contracts.py
python tests/test_admin_contract.py
python tests/test_conventions.py
python tests/test_api_contract.py
python tests/test_nim_benchmark.py
python tests/test_security_hardening.py
python tests/test_repository_security_metadata.py
python tests/test_product_planning_contract.py
Expand Down
1 change: 1 addition & 0 deletions conductor/tracks.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,3 +4,4 @@
|---|---|---|
| 001-paper-grounded-orchestrator | active | Implement the source-backed orchestration contract with TDD, DDD, and CDD |
| 002-enterprise-design-foundation | active | Add paper-grounded screen design, user stories, REST API, code/DB conventions, and i18n |
| 003-nim-cost-quality-benchmark | active | Evidence-grade NIM catalog discovery, all-modality capability probes, and the route/conduct/single-worker cost-quality benchmark (docs/nim_benchmark.md) |
Loading
Loading