Projects01 of 15
ResearchLens
Evidence-grounded literature retrieval
What it does
ResearchLens is a research assistant over 106 indexed papers, split into 15,664 passages. It answers from evidence, cites the passages it used, and refuses when it has none.
- Ask across the whole corpus, or confine a question to papers you choose.
- Add your own PDFs. They are parsed into an index that belongs to your session alone and is never merged into the shared one.
- Ask what work exists like a paper you added.
- Ask about what is current, and the search reaches past the corpus to arXiv, PubMed and OpenAlex.
Reading a PDF properly
A PDF is not a string, so the parser recovers the structure a reader sees:
- headings are found by font geometry;
- two-column reading order is recovered from word coverage;
- words broken across lines are de-hyphenated against the document’s own vocabulary.
Finding the right passages
Retrieval is hybrid. BM25 and dense retrieval run side by side, their rankings are fused by reciprocal rank, and a cross-encoder reranks the result.
Grounding is enforced in code
The model is not asked to be faithful; the pipeline makes it so.
- Retrieval runs first, and the model is never called without evidence.
- Citation markers are resolved against the passages actually retrieved. An invented one is deleted before it reaches the reader.
- An answer left with nothing supported becomes a refusal.
Three kinds of evidence, kept apart
Results from arXiv, PubMed and OpenAlex are marked as abstracts rather than passages. An abstract supports what a paper claims, not what it measured.
ResearchLens also reads this site and its textbook, fetched fresh. A question about my work, or about a mechanism I have written up, is answered from those and cited to the page or the proposition. Each source is labelled, because a self-description, a piece of teaching and a measured result are three different kinds of claim.
Built with
Python, FastAPI, ONNX and Gradio, with a benchmark harness scored against hand-labelled ground truth and 293 tests. It is deployed on Hugging Face and runs locally with no API key at all.