Retrieval layer for arxiv-complete
A small prototype that turns the full arXiv metadata population (3.15M papers, one parquet file, ~1.6 GB) into a searchable corpus — so you can answer "what is the existing research on topic X?" with real recall and immediacy, instead of the navigational links (blog posts, GitHub lists) a web search returns.
It layers semantic (embedding-based) retrieval on top of keyword search, because keyword-only search misses synonyms.
What it is
| File | Role |
|---|---|
data.py |
Lazy loader for the arxiv-complete metadata parquet (polars scan_parquet). Streams rows so only the matching subset is materialized. |
embed.py |
Pluggable embedding providers: an OpenAI-compatible remote provider (production) and a dependency-free local hashing fallback (offline, no API key). |
index.py |
Builds and persists a FAISS flat L2 index + metadata sidecar (paper_id → vector, title, year, category, URL, abstract). Incremental adds. |
query.py |
The interface. Embeds the query, runs ANN search, ranks, applies filters. Falls back to keyword search if no index is available. |
search.py |
CLI tying it together. |
build_index.py |
Embeds the corpus and persists the index. |
Data source
- Dataset:
secemp9/arxiv-complete - Config used:
metadata— 3,148,796 rows, one per arXiv paper. - Columns:
paper_id, title, authors, abstract, categories, primary_category, submitter, license, doi, journal_ref, comments, report_no, msc_class, acm_class, proxy, n_versions, first_version_date, latest_version_date, oai_datestamp, oai_sets, arxiv_abs_url. - Snapshot through 2026-08-26. License per paper is mixed (CC-BY, CC-BY-NC-ND, PD, …) and carried in the
licensecolumn. - Related configs exist too (
versions,files, and content blobs likelatex/source/paper_text/ps/pdf) — this prototype only consumesmetadata.
Setup
cd aisys/retrieval
uv venv ../.venv && source ../.venv/bin/activate
uv pip install -r requirements.txt
Requires Python 3.10+. faiss-cpu is needed for the ANN path; the keyword
fallback works with just polars + numpy.
Running
python search.py --query "speculative decoding"
Options
--query, -q Natural-language query (required)
--top-k Max results (default 10)
--year-min Only papers from this year onward
--year-max Only papers up to this year
--category Filter by primary_category, e.g. cs.LG
--survey-only Only survey/overview/taxonomy papers
--keyword-only Skip embeddings; pure keyword fallback
--embedder "hashing" | "remote" (default: auto)
--index-name Persisted index name (default "arxiv")
--parquet Override the metadata parquet source
Worked example
# Semantic search (builds the index on first run)
python search.py --query "speculative decoding" --top-k 5
# Keyword-only, no embeddings
python search.py --query "speculative decoding" --keyword-only
# Surveys only, in a category
python search.py --query "speculative decoding" --survey-only --category cs.LG
Embeddings: remote (production) vs local fallback
This machine (Strix Halo, ~104 GiB unified RAM) has OOM-killed local models before, so we deliberately do not import torch or run a large local embedding model here.
Production path — set these env vars and the CLI uses a remote OpenAI-compatible embeddings endpoint:
export EMBEDDING_API_BASE=https://api.openai.com/v1
export EMBEDDING_API_KEY=sk-...
export EMBEDDING_MODEL=text-embedding-3-small
Any endpoint with the OpenAI embeddings shape works (OpenAI, Together, Groq, a local vLLM/Ollama server, …).
Offline fallback — with no key configured, the code automatically uses
HashingEmbedder, a deterministic, dependency-free embedding that projects a
bag of character n-grams into a fixed-size vector via the hashing trick. It needs
no weights, no GPU, no internet and no API key, so the whole ANN pipeline is
fully exercisable for demos and CI. (It is not semantically meaningful — synonyms
are not mapped together — but it shares enough lexical overlap to cluster near
the right topics.)
Swapping in a real local model (if you have the RAM)
Add a class in embed.py:
class LocalEmbedder(Embedder):
DIM = 768
def __init__(self):
from sentence_transformers import SentenceTransformer
self._model = SentenceTransformer("all-MiniLM-L6-v2")
def embed(self, texts):
return self._model.encode(texts, normalize_embeddings=True)
and pass it via embedder=LocalEmbedder() to the index/query builders. The
rest of the pipeline is agnostic to where the vectors come from.
How the index works
build_index.pyscans the metadata (optionally filtered by category/year), embedstitle + abstractfor each paper, and adds the vectors to a FAISSIndexFlatL2.- It persists
<index-name>.index(FAISS),<index-name>.meta.pkl(parallel arrays: id, year, category, url, title, abstract), and<index-name>.meta.json(provenance). search.pyloads that index, embeds the query, runs ANN search, and attaches metadata to each hit. Re-runningbuild_index.pyadds only new papers (incremental, keyed bypaper_id).
Notes / decisions
- The corpus is large (3.15M rows), so the data layer streams via lazy
scan_parquetand pushes filters (year, category, survey) into the scan rather than loading everything into RAM. --survey-onlymatches abstracts containingsurvey,comprehensive,overview, ortaxonomy. Combined with semantic search this isolates the handful of survey/overview papers about a topic (e.g. ~28 for "speculative decoding").- The
licensecolumn is loaded but not yet surfaced in results — a natural follow-up for any downstream licensing filter.