Recall turns PDFs, images, and pasted text into a searchable, source-grounded personal knowledge base. Local Qwen3 classifies and summarises every item, then answers questions from retrieved evidence.
This repository is an AI Product Manager portfolio MVP. It deliberately runs without paid services: the API includes local SQLite storage and a deterministic hybrid retrieval fallback. OpenAI-compatible generation can be enabled with environment variables.
- Upload a text/scanned PDF, upload an image, or paste text.
- Recall extracts or OCRs the content and preserves PDF page numbers.
- Local Qwen3 automatically classifies and summarises the document.
- Ask across the personal library.
- Receive a grounded synthesis and clickable source cards.
- If evidence is weak, Recall says the library is insufficient instead of inventing an answer.
- Open an original PDF from its citation or delete an item together with its vectors and local file.
apps/web— Next.js product interfaceapps/api— FastAPI ingestion and RAG APIdocs— product decisions, evaluation plan, and database schema
cd apps/api
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-ocr.txt
cp .env.example .env
uvicorn app.main:app --reload --port 8000Run Ollama in another terminal. The project uses the installed qwen3:latest model by default:
ollama serve
ollama pull qwen3:0.6b
ollama pull qwen3-embedding:0.6b
ollama listcd apps/web
cp .env.example .env.local
npm install
npm run devOpen http://localhost:3000. The frontend loads built-in demo content when the API is unavailable, so the product can still be reviewed as a UI prototype.
GitHub Actions automatically builds and publishes a frontend-only product demo after every push to main. It uses three prebuilt knowledge items and grounded answers for the suggested questions, so reviewers can explore the complete interface without installing Ollama. Free-form ingestion and live model inference remain local-only.
Online demo: https://nbrhyq.github.io/Recall_RAG/
For the product story, decisions, evaluation evidence, and retrospective, see Portfolio Case Study.
GET /healthGET /api/documentsGET /api/documents/{id}/file— open the locally stored original PDF/imageDELETE /api/documents/{id}— remove the item, chunks, vectors, and original filePOST /api/pdfs— multipart PDF upload, maximum 20MBPOST /api/images— PNG/JPG/WebP upload, maximum 10MBPOST /api/texts— manually entered title and contentPOST /api/ask
Interactive documentation is available at http://localhost:8000/docs.
PDF pages first use native text extraction; pages without text and uploaded images use local PaddleOCR. Classification, summaries, and grounded answers run through qwen3:latest on the user's local Ollama service, so knowledge content stays local.
Retrieval uses a three-stage local pipeline:
- Qwen3 Embedding 0.6B writes semantic vectors to persistent Qdrant local storage.
- Keyword and vector results are merged into a hybrid candidate set.
- Qwen3 reranks candidates before generating a citation-grounded answer.
See docs/PRODUCT.md for scope and docs/EVALUATION.md for the quality framework.
The checked-in evaluation set contains 45 manually labelled questions: 35 answerable and 10 that should be refused. The final hybrid pipeline reached 85.7% Hit@1, 100% Hit@3/5, and 100% abstention accuracy on this set. See the full results and Bad Case Analysis, including limitations and latency trade-offs.
Reproduce the comparison after importing the evaluation PDF:
cd apps/api
source .venv/bin/activate
python evaluation/run_retrieval_eval.py --versions v0_keyword v1_embedding v2_hybrid_rerank --force