The problem
Finding a suitable repository requires combining developer interests with repository health and contributor opportunities. Fetching GitHub data on every widget request also couples profile rendering to API latency and rate limits.
Context and constraints
An SVG embed in a profile README exposes precomputed activity and candidate repositories without running a GitHub query for every profile view.
- GitHub data acquisition must remain outside the widget request path.
- Owned, starred and hidden repositories should not reappear as candidate recommendations.
- A temporary GitHub failure needs bounded retries; authentication and application errors should fail promptly.
What I built
A TypeScript monorepo with a pure matching package, batch repository indexing and user-stat jobs, precomputed recommendation storage, a Next.js dashboard and an SVG widget endpoint. Matching combines language/topic overlap and repository health, with feedback and contributor-readiness adjustments. GraphQL requests retry transient failures with bounded exponential backoff.
Architecture
Batch jobs acquire GitHub data and compute matches; the web app and widget read stored results.
- GitHub GraphQL
- Repositories and activity
- Scheduled indexer
- Acquire data, compute matches
- Matching rules
- Eligibility and weighted scores
- Precomputed data
- Profiles and recommendations
- Web and SVG views
- Dashboard and profile embed
My contribution
Author and maintainer. Matching and eligibility rules, repository indexing, feedback adjustments, onboarding/dashboard flows, SVG widget delivery and transient-request regression tests.
Technical decisions
What was chosen, why, and what it cost.
Precompute data outside widget requests
- Decision
- Acquire GitHub data in scheduled jobs and serve widgets from stored activity and recommendations.
- Why
- Profile rendering avoids live GitHub queries and their rate-limit exposure.
- Trade-off
- Data can be stale between successful indexing runs.
Keep matching rules testable without I/O
- Decision
- Implement scoring and eligibility as pure functions in a separate package.
- Why
- Ranking behaviour can be tested independently of API and database availability.
- Trade-off
- Tested heuristics do not establish that users find the recommendations useful.
Retry only recoverable GitHub failures
- Decision
- Allow three attempts for network failures and transient HTTP responses; keep authentication and GraphQL application errors fail-fast.
- Why
- Recover brief outages without hiding permanent errors behind repeated requests.
- Trade-off
- Long outages can still exhaust the retry budget.
Verification
How the implementation was checked, and how much of that can be shown publicly.
Pure matcher and SVG testsevidenced
Inspected eligibility, scoring, feedback and widget tests. No recommendation-quality outcome is inferred from unit tests.
Transient-request regression testsevidenced
Own-repository PR #1 covers recovery, retry exhaustion, HTTP 401 and GraphQL errors with mocked responses.
Scheduled run evidenceevidenced
A recent Nightly Index & Match run completed successfully at the reviewed commit.
Broader matching-quality reviewnot publicly evidenced
The progress report leaves the broader user evaluation gate incomplete.
Results
- Widget serving reads precomputed data rather than calling GitHub directly.
- Transient API retries are bounded and covered by focused tests.
Limitations and disclosure
What this project does not do, and what cannot be shown publicly.
- Matching is heuristic, not an evaluated semantic recommender. Broader user-quality review remains incomplete.
- Upstash caching, semantic matching, translated summaries and digest email are not implemented.
- A successful scheduled run confirms job execution, not the quality of every recommendation or reliability during an extended GitHub outage.
Own-repository PRs and workflow runs are project validation evidence. They are not included in merged upstream contribution counts.