Models
- Llama 2 7B
- Flair
- Hugging Face
- Self-hosted vision LLMs
- LightRAG
Computational methods
We treat historical texts as relational data. Every collection runs through a reproducible pipeline that combines language models, expert human validation, and network analysis, with self-hosted compute on our own infrastructure.
High-resolution capture with color and scale references. Vision models assess each page's physical condition and produce restoration reports.
Language models correct OCR errors and normalize transcriptions while preserving historical spelling where it carries meaning.
Automated extraction of people, places, institutions, and events, with temporal alignment and semantic embeddings across the corpus.
Every annotation goes through expert review in our own tooling and in Argilla. No entity reaches the graph without human and vocabulary control.
Concurrence graphs, community detection, and comparison against synthetic realizations to infer social dynamics from partial observations.