Local research workspace: search open-access papers (OpenAlex and arXiv), upload PDFs, ask questions with on-page highlights, compare two papers on a claim, and export citations.
Papers, embeddings, and chat history stay on disk. Citations jump to the PDF. Chat needs an OpenAI API key; pairwise compare uses TypeSafe Jev. Ingest and retrieval run locally (Docling, Chroma, BGE).
Copy backend/.env.example to backend/.env, set OPENAI_API_KEY, EMAIL, and TYPESAFE_API_KEY (compare), then:
docker compose up --buildOpen http://127.0.0.1:5000. First build is large; first ingest downloads Hugging Face weights inside the container. docker compose down keeps the thesys-data volume (papers and chat). docker compose down -v deletes it.
Chat over your research papers. Answers cite sources in the document. Click a citation to jump the preview. Full highlights the whole retrieved passage; Focused (default) paints a tighter highlight on the same passage. Toggle in the PDF header.
Reader mode. Open a paper full-width, select text, and ask about that span.
Document summaries. Follows a general format of problem statement, methodology, key findings, limitations, and metrics. Citations still jump to the highlighted passage.
Figures and diagrams. Ingest describes plots, pipelines, and other figures so chat and summaries can cite them. Click a citation to highlight the figure on the page.
Compare two papers. Select two library PDFs, turn Compare on, and ask about one claim. TypeSafe Jev classifies each retrieved pair as corroboratory, contradictory, or neutral; the side panel lists those pairs with confidence. Without TYPESAFE_API_KEY, compare is unavailable.
Discover papers. Search open-access literature from OpenAlex and arXiv in one pool, then preview a record, open the link, or add it to the library.
Library and export. Session files in one place. Bibliography export as BibTeX, RIS, EndNote, or CSV.
- Frontend: React (Vite) on port
5000 - Backend: FastAPI on port
8000 - Local RAG: Docling ingest → Chroma + BGE embeddings → BGE reranker
- OCR: off by default (
DO_OCR=false); setDO_OCR=truefor scanned PDFs - Chat: OpenAI via LangChain
- Compare: TypeSafe Jev (
TYPESAFE_API_KEY) - Paper search: OpenAlex + arXiv (open-access; optional OpenAlex API key)
- Citations: CiteAs (uses
EMAIL)
PDFs, vectors, and chat history stay on disk under backend/data/. First ingest downloads embedding/rerank weights (Hugging Face). GPU helps; CPU works and is slower.
Needs Python 3.11+, Node.js 20+, and an OpenAI API key (or use Docker above).
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
pip install -r requirements.txtcp backend/.env.example backend/.env
cp frontend/.env.example frontend/.envSet at least OPENAI_API_KEY and EMAIL in backend/.env. TYPESAFE_API_KEY is required for Compare (TypeSafe Jev). HF_TOKEN is optional (Hugging Face rate limits). OPENALEX_API_KEY is optional. DO_OCR defaults to false; set it to true to OCR scanned / image-only PDFs on ingest (slower).
cd backend
python main.pycd frontend
npm install
npm run devOpen http://127.0.0.1:5000. The UI talks to VITE_API_BASE_URL (default http://127.0.0.1:8000).
- Optional Cohere embeddings and reranker for users who want higher retrieval quality. Local BGE stays the default.
- Per-chat settings: saved title, selected model, and reasoning effort.
AGPL-3.0. Maintained by Sreehari. Issues and pull requests are welcome; start from an issue if the change is large.








