I contribute fixes upstream, not only to my own repositories.
Shipping your own repository is one skill. Landing a change in someone else's — under their conventions, their review and their standards — is a different one. Each entry below is a real pull request, shown with the problem it addressed, what I actually changed, the tests the request documents, and the status it has right now.
6
merged upstream
1
open, not merged
6
upstream repositories
24
files changed
// +765 / −95 lines across 24 files · state, dates and diff sizes verified against the GitHub API on
Where the work landed
6 repositories across 6 organizations.
Upstream repositories · 6
Tap a repository to show only its pull requests.
Filter contributions
7 contributions
Upstream contributions
AcademySoftwareFoundation/dna #195
Fix SPI v2 Uvicorn app target
Merged
Opened
Merged
Open for
2 days
The problem
The documented startup command for the SPI v2 backend named the wrong ASGI target, so a reader following the README exactly could not start the service.
What I changed
Corrected the Uvicorn target in the service README to the FastAPI application object. A one-line change, and the smallest contribution here — but a wrong getting-started command is the first thing a new contributor hits.
Tests and checks
Documentation-only change; the corrected command was checked to name the application object.
The MCP server's loop_estimate_cost re-implemented the loop-cost realistic-mix table by hand instead of importing it, and the copy never read pattern.cost.early_exit_required. For early-exit patterns that overestimated the realistic cost by roughly three times against the loop-cost library the documentation points users at.
What I changed
Added the loop-cost package as a sibling file: dependency — the same pattern mcp-server already uses for loop-audit, loop-gate and loop-context — and called its estimateCost() directly instead of maintaining a second mix table, so the two tools cannot drift apart again.
Tests and checks
A regression test added to test/server.test.mjs with an early-exit fixture pattern, checked to fail against the pre-fix code and pass against the fix.
npm test passing for the MCP server (26/26) and for loop-cost (12/12).
Harden loop-action command execution against unquoted shell expansion
Merged
Opened
Merged
Open for
2 days
The problem
The action spliced inputs.command into its run: script as raw text, before bash ever parsed it. For a multi-line command: | block that silently broke sandbox isolation — only the first line reached loop-sandbox run --, and every later line executed as a bare top-level statement on the runner, outside both the sandbox and the circuit breaker.
What I changed
Wrote the command to a temporary script passed through env:, so the workflow YAML never re-templates it, and invoked it as a single quoted bash "$SCRIPT" call on both the sandboxed and the direct path. The README gained safe-usage documentation and an explicit warning against embedding untrusted input in command.
Tests and checks
No tests or checks are documented in this pull request.
Fix AGUIAdapter.dump_messages reordering ToolReturnPart after UserPromptPart
Merged
Opened
Merged
Open for
14 days
The problem
A ModelRequest carrying both a ToolReturnPart and a UserPromptPart came out of AG-UI dumping in the wrong order — the user message first, the tool message after it. Providers that require a tool result to follow its tool call immediately, Anthropic and Bedrock among them, then rejected the replayed history.
What I changed
Preserved the original ModelRequest.parts ordering through dumping, with a _flush_user_content() helper that flushes buffered user content at ordering boundaries rather than at the end, and added regression coverage for the ToolReturnPart-then-UserPromptPart case.
Tests and checks
Regression test test_dump_messages_preserves_part_order added to tests/test_ag_ui.py and run under pytest.
ruff check run over the changed adapter and test file.
Repeat count could only be set globally. Measuring how non-deterministic one test case was therefore meant repeating the entire evaluation, paying for every other test to run again as well.
What I changed
Added tests[].options.repeat, which overrides the global repeat for that test only. Each repeat index gets its own cache namespace — including index 0 — so repeated runs neither reuse nor overwrite the ordinary non-repeat cache entry. Scenario config now inherits all options fields, matching the existing vars/metadata/assert inheritance, with scenario test options still taking precedence. Invalid values fall back to the effective global repeat and are diagnosed by the public Zod schema and the generated JSON schema.
Tests and checks
467 tests passed across six evaluator, type and config-schema suites under vitest.
Generated JSON schema regenerated byte-for-byte clean; typecheck, lint, format and the repository architecture check all passed.
Production documentation site build passed.
Exercised end to end through the local CLI, covering scenario precedence and a fractional repeat value falling back to a single run.
The doctor command gave no way to tell a missing optional dependency apart from a broken core installation: absent optional tooling failed the check rather than reporting itself, and Rich markup meant the extras being discussed did not even render literally in the output.
What I changed
Added deep-mode checks for the optional extras and for the semgrep, pip-audit and opa command-line tools, plus an API-key readiness warning. Missing optional tooling now warns instead of failing a core-only install, and the output labels are escaped so the extras render as written.
Tests and checks
Doctor tests added for both the missing and the present optional-toolchain cases; pytest tests/test_doctor.py reported 6 passed.
Tell a collapsed axis apart from a uniform offset in describeBiasPattern
Open
Opened
Status
open — not merged
The problem
The diagnostic reported a collapsed vertical channel — predicted y nearly constant while target y swept most of the screen — as a uniform offset, and advised the user to recalibrate. Recalibration cannot repair a collapsed axis. The function checked only direction consistency and magnitude relative to accuracy, and both failure shapes point every bias arrow the same way, so the uniformity test could not separate them.
What I changed
Regressed bias against target position independently per axis, which does separate them: a uniform offset gives a slope near zero, a collapsed axis a slope near minus one. The slope check runs before the uniformity check, a collapsed axis is now named as such per axis, and the message points at the eye-zoom and resolvable-step diagnostics instead of recommending recalibration. Existing thresholds were left untouched, and the one new threshold is justified against the two data points actually available.
Tests and checks
Four tests added to the diagnostics bundle suite: the recorded session classified as a collapsed axis rather than an offset, the other axis reported independently, a mirrored fixture proving the detection is not hardcoded to one axis, and the existing uniform-offset fixture re-asserted as still an offset.
Core suite 75/75 and native suite 173/173 passing; the core workspace typecheck clean.
Status, dates and diff sizes are read from the GitHub API and cached in the content layer, so the page never depends on GitHub being reachable and never shows a state nobody checked.
An open pull request is shown as open everywhere — badge, lifecycle and evidence — and content validation rejects a merged-typed proof on anything that is not merged.
The problem and change for each entry are written from that pull request’s own description. Where a request documents no tests, the card says so rather than leaving a gap that reads like evidence.
Nothing here counts reviews, reactions, stars or downstream impact. Those are not in the records, so they are not on the page.