Notebook — Course Memory and Recall
Course and lecture workspace with rich-text notes, drawing blocks, document ingestion, hybrid retrieval, cited answers, and study workflows. Includes PostgreSQL-backed integration tests and deterministic course-memory fixtures.
The problem
Course evidence is spread across lecture notes, slides and uploaded documents. Search and recall need to retain the lecture, page, source type and user boundary instead of returning an untraceable answer.
Context and constraints
A lecture workspace keeps student notes and source materials together, while answers retain links to the actual evidence used.
- Uploads and notes must stay scoped to the authenticated user and course.
- Editing should not wait for extraction, indexing or model calls.
- Search-only use must remain available without a configured model.
- Model-generated citation identifiers need validation against stored source identity.
What I built
A Next.js workspace with Tiptap notes and drawing blocks, local draft autosave, server synchronization and version history. PostgreSQL stores courses, documents, chunks and durable background jobs. Retrieval fuses full-text, fuzzy and optional current-model vector results. Intent-routed recall builds cited answers, while source-grounded study sessions retain attempt history.
Architecture
Uploads and notes become course-scoped evidence; retrieval and context construction feed cited recall without blocking editing.
- Lecture workspace
- Notes, drawing, materials
- Indexing jobs
- Extract, normalize, chunk
- Course memory
- PostgreSQL and source identity
- Hybrid retrieval
- Keyword, fuzzy, optional vector
- Cited recall
- Context and citation validation
My contribution
Author and maintainer. Course workspace, editing and autosave, ingestion and indexing jobs, retrieval and citation handling, study workflows and test fixtures.
Technical decisions
What was chosen, why, and what it cost.
Keep source identity attached to retrieval chunks
- Decision
- Carry lecture, file, page/slide, source class and note-section identity through retrieval and context building.
- Why
- An answer can link back to the evidence the student can inspect.
- Trade-off
- Valid citation numbers establish source identity, not the truth of every generated statement.
Move ingestion out of the editing path
- Decision
- Persist extraction and indexing work in PostgreSQL background jobs and keep draft edits locally before syncing.
- Why
- Document processing and model latency should not block note taking.
- Trade-off
- Local draft recovery is not a complete offline application shell.
Ignore vectors from a different embedding model
- Decision
- Store embedding identity and restrict semantic retrieval to the current configured model.
- Why
- Vectors from incompatible model spaces should not silently enter one ranking.
- Trade-off
- Switching models requires re-embedding; keyword and fuzzy retrieval remain available meanwhile.
Verification
How the implementation was checked, and how much of that can be shown publicly.
Database-backed integration and memory testsevidenced
Inspected real-PostgreSQL API tests and a deterministic seven-lecture course with typo, abbreviation, source-class and missing-topic cases. Suites were not rerun for this portfolio update.
Model boundary and citation testsevidenced
Stubbed model tests exercise prompt construction, context budgets, invalid citation removal and course-only versus explanation modes.
Real embedding evaluationnot publicly evidenced
The semantic test block is opt-in and requires a configured local embedding endpoint; no general retrieval-quality benchmark is claimed.
Results
- Implemented course-scoped retrieval and cited recall with inspectable source identity.
- Search and extractive recall can operate without a configured language model.
Limitations and disclosure
What this project does not do, and what cannot be shown publicly.
- Scanned-document OCR, handwriting recognition and a full offline app shell are not implemented.
- Real embedding checks are opt-in; default model tests use a stub. Citation identity checks do not prove every answer is factually correct.
- No measured educational outcomes or production deployment are established. Login rate limiting and an S3 storage adapter remain unbuilt.
The application lives in the notebook subdirectory of 30-day-code. It is a separate project from the archived Smart Notebook and is not counted as an upstream contribution.