install it as scholion — the command line also answers to crossread: read your sources against each other
A local engine that reads your genome, lab forms, prescriptions and wearable data against each other and shows where every statement came from. A scholium is a note in the margin of a manuscript saying where something was taken from: the program does not replace the source, it hangs provenance off it.
Attach it to a chat with Claude or ChatGPT and say «set this up» — the model reads it and leads, one small step at a time, explaining each before it happens. Your data stays on your machine; the safety rules outrank every other instruction the model is given.
And if you run Ouroboros Desktop, there is nothing to install by hand: the skill has been accepted into OuroborosHub and installs from there in one click (Skills → OuroborosHub → Install), pip package and all.
Want the full picture instead? the complete bundle with the reference texts · every way to install, for people at home in a terminal
What is different
A report where everything is green is easy to sell. What is hard is saying what your data does not entitle you to claim — and that is what decides whether the rest is worth believing.
Coverage is measured per gene: a gene read at 70 % yields the same zero as a gene read at 100 %. Large deletions are not seen by short reads at all, and that is said plainly rather than passed over.
Before showing a “pathogenic” variant from a database, the engine checks: is such a condition plausible given your picture, are enough copies of the gene affected, does the variant work in the direction assumed, is the effect distinguishable from the method's own error.
No reference line on your form — the marker is shown without a flag, not with a range borrowed from elsewhere. No complete panel — no biological age. A gap stays a gap.
Every genetic statement carries an evidence level. A polygenic percentile depends on the model and the reference population; it is a position in a population, not a probability of disease, and that is written next to the number.
When half the rows are highlighted, people stop looking at highlighting. A threshold that fires on almost everyone is treated as broken and rebuilt: it measures a property of the data rather than of the objects.
The changelog answers not “which files were touched” but what changed in the conclusions and what was withdrawn. An old formulation lives on in your head until it is explicitly taken back.
How it is built
Numbers, flags and connections are computed by local code: reproducibly and checkably. A language model adds only the wording — a coherent reading, priorities, questions for the doctor. There is not a single call to a model anywhere in the core, and the application proves it by scanning its own source right on the Assistant tab — and a separate Guide tab explains, screen by screen, what every colour and label means, so a screen is never left unexplained just because the source is not at hand.
VCF, PDF lab forms, prescriptions, a wearable export — they sit in your own folder.
Flags, trends, CPIC phenotypes, ClinVar findings, PGS percentiles, coverage.
Optional: a reading with sources and questions for the doctor. Any model.
The web interface, the command line (scholion / crossread), a skill — in a chat or in the shared skills folder — an MCP tool server, an Ouroboros plugin, and a plugin package for ChatGPT desktop, Codex and Cursor.
Who it is for
One product, but everyone has their own reason to open it a second time. Below, honestly, what each of them gets today.
You have a VCF or a BAM from whole-genome sequencing and you do not want to hand it to somebody else's cloud. You get: ClinVar findings by tier, ACMG secondary findings, pharmacogenomics down to star alleles and HLA, polygenic scores — and a list of what your data does not let you say. A consumer test export is read as well — 23andMe, AncestryDNA, MyHeritage, Living DNA, FamilyTreeDNA, including inside the archive the provider hands you — and the class of the input is measured while it is read: a chip is not passed off as a whole genome.
Nextcloud, Home Assistant, Nightscout — and health as one more layer that should not depend on somebody else's server. The core runs on the Python standard library plus a single dependency for reading PDF lab forms (pdfplumber); there is not one network call for any computation, and the whole thing is readable in an evening.
Thyroid, iron, lipids, carbohydrate metabolism. A panel every quarter, prescriptions, follow-up draws. Here the benefit arrives on the day of the appointment: the new panel checked against the last one, what to monitor on the current regimen, which questions to ask.
The knowledge base lives apart from the code; every entry carries its source and its evidence rule, and an edit to the base is versioned as an event of its own. A disputable threshold is visible because it is written as a number in a file rather than hidden in code.
New in 0.5.0 and 0.5.1 — the main thing
The thyroid used to appear on the page three times — among the deviations, in the line about a drug, and in the list of what to test — and nowhere was it said that these are one subject. A body system is now the unit things assemble around. The tab is called Radar: the figure, two rings per system, one block each, and the full card one click away. Since 0.5.1 every segment also carries a panel of positions somebody wrote, standing in front of the base.
System: Lipids · patient's register
this build cannot answer for this list — the genome is not
readable on this profile — see the genome status
1. Laboratory now 100/100 — markers measured: 3 of 4
never taken: Cholesterol, total
2. Movement 100 → 100 (0) against 2025-11
3. Genetics base GenCC 2026-09-06 — genes: 38;
positions of the curated panel: 7; read: 0, unread: 0
positions whose phrase is signed by the panel's author,
and by no clinician: 7 (signed 2026-09-13)
4. Prescriptions · 5. Clinician's target · 6. What to test · 7. Questions
Next step, in three baskets
test · read in the genome · ask the clinician
no genome is attached — a full genome would close for this
system: 40 genes of the list not read here
The synthetic demo profile of a fictional person, shipped with the package: scholion system lipids. That profile has no genome, so nothing of the genetic half is read — and the card says so rather than reporting silence as health. The same card comes back from a click on a radar segment, from an organ on the figure, from the local API and from the assistant's tool.
How much of the laboratory panel has been measured, and how much of the genetic half has been read. They are different quantities: somebody who has taken the whole panel and had no gene read should not see what somebody in the opposite position sees. A system whose genetic list is not composed shows that ring dotted — an absence, not a zero.
1469 genes across the fifteen systems from the Gene Curation Coalition export: who asserted the gene–disease link, how strongly, under which mode of inheritance, and on what date. Two groups disagreeing about one gene are shown disagreeing rather than averaged. An assertion classified Limited, Disputed or Refuted is never presented as a finding, and one copy of an allele in a recessive gene is carriership and a question — never a risk line.
214 positions across fourteen of the systems, authored rather than generated. The unit of a row is a position, not a gene — an rsID with its HGVS, the allele the author named, the phrase for one copy and for two. A row whose phrase for the state actually found is missing is kept and printed as pending: somebody put the position here, and what follows from that genotype is still to be written.
scholion recompute finds that step, says what it still needs, and runs it with the progress visible.What it looks like in use
Every number in these screenshots comes from the synthetic demo profile of a fictional person, which ships with the package. You can look at the product before loading anything of your own: pip install scholion && scholion init --demo && scholion serve.
Markers over years, flags, movement, the link to the genome — and a focus of attention: the one thing you are working on right now.
For a drug in your regimen — what the guidelines say as applied to your data: pharmacogenomics, what to monitor, interactions with what you already take. The result is a list of questions, not of instructions.
ClinVar findings laid out by clinical weight; polygenic risks as navigation for screening; the longevity layer as context rather than a verdict.
The form for the next draw assembles itself — from the current regimen, the clinical thresholds you have crossed, what you agreed with your doctor, and what is missing for the computations.
“I think late coffee makes me sleep worse” is a testable statement. The n-of-1 module turns it into an experiment with the protocol fixed in advance.
Why come back
A local tool has no subscription and no notifications. There are exactly three reasons to open it again, and all three arrive on their own.
| What arrives by itself | What Scholion does with it |
|---|---|
| A new ClinVar release | A scheduled monthly reanalysis: the fresh database, a diff against your variants, a repeat ACMG scan, a review of PGS candidates. A report of the form “this many were reclassified last quarter, and these are yours”. |
| A new lab panel | The new panel read, checked against the previous one, and the checklist for the next draw updated. Forms keep arriving; the genome is loaded once. |
| A new prescription | Checked against the current regimen and the pharmacogenomics, with a list of questions ready for the date of the appointment. |
Privacy & licences
Genome, labs, prescriptions and metrics are stored only on your machine. The code reaches outward only to public reference databases (Ensembl, RxNorm, CPIC) and translation services (for cross-language drug names), and only for general information — your data never goes into the request. SCHOLION_OFFLINE=1 forbids even that.
The depersonalised build is produced by a sanitiser: if any sign of personal data reaches it, the build stops. The package holds the code, the knowledge bases, the tests and the demo profile.
More than a thousand tests on the standard library, a backward-compatibility check of the public contract (56 commands, 20 snapshots), and a separate set of rules about the safety of the conclusions. Anyone who takes the package runs them with one command — inside the package, not only in our repository.
Code — Apache-2.0; the curated knowledge base — CC BY 4.0. The provenance of every file that is not ours is recorded in ATTRIBUTION.md; data whose licence forbids commercial use is not bundled at all.
Six doors, one core — a zero-install folder, pip install scholion, a skill for a model with no terminal, an MCP tool server, an Ouroboros plugin, or a plugin package for ChatGPT desktop, Codex and Cursor. Exact, tested commands are in the Installation section below.
Reports print English by default and switch with --lang ru or SCHOLION_LANG=ru. Recognition of Russian lab forms is not a setting but a feature: a marker's name comes from the form in your hand, so it is shown in the language it was printed in.
The assistant is any model you like. A skill for Claude, an Ouroboros plugin and a plugin package for ChatGPT, Codex and Cursor ship with it; for anything else, scholion skill prints the instruction and scholion assistant --context assembles a snapshot of the state with the list of commands. Through a skill, the model gets no access to your machine — it asks you to run a command and reads the output. Through the tool server, it calls tools that run on your machine and read only your own data.
Data used for testing
A tool tested only on the person who built it passes its own tests and fails on the world. Three open sources make the testing possible — and not one of them is republished here.
The standing reference run covers 36 sets from eleven providers — chips, whole genomes, panels and split call sets. PGP participants consented to publishing their data openly precisely so that it would be worked on — and without it this product would only ever have seen the author's own machine. Thank you. Nothing of theirs lives here: no genotypes, no findings, no identifiers, no medical records. What travels out of that work is the behaviour of the engine and aggregate numbers.
A synthetic FHIR R4 bundle the import is tested against: a portal export with clinical records, produced by a real system rather than by the same person who wrote the parser. The patient is generated — which is why the file may live in the repository at all: nobody's consent is involved. Apache-2.0.
The input-format detector — detection.py and text_io.py — vendored under Apache-2.0 with attribution, every change marked in place and a way to update it. It decides what a file is by its content rather than by its name: BAM, CRAM, VCF, gVCF told apart from VCF, paired FASTQ, and the exports of 23andMe, AncestryDNA, MyHeritage, FamilyTreeDNA and Living DNA — gzip, zip and tar wrappers included.
Full provenance for every third-party file is in ATTRIBUTION.md, and a single command compares the vendored code with upstream and reports drift in either direction.
Installation
The same analysis runs underneath all six; what changes is what you need to already have, and who ends up typing the commands.
pip install scholion
scholion init --demo
scholion overview
scholion serve
Python 3.10+ and nothing else needed first. This is the pip package — it gets you the local web interface and the full command line in one step; pdfplumber, needed for reading PDF lab forms, comes with it. The command line answers to two names, scholion and crossread, both installed by the same package. An optional pip install "scholion[genome]" adds faster VCF access via pysam; without it the built-in reader still works, just slower.
Python 3.10+ is the only requirement; every line of analysis runs on the standard library. (Reading PDF lab forms is the exception — that needs pdfplumber, which this delivery does not bundle.) Start with ./bin/crossread --help. This is also where the genome-preparation tooling lives — FASTQ → VCF, PharmCAT/PyPGx — for building a VCF from raw reads; scholion tools reports which external programs (bcftools, samtools) that needs and how to get them.
The easiest way in — no terminal skills needed: download the skill file and attach it to your chat with Claude or ChatGPT, then say «set this up». The model reads it and takes it from there — one small step at a time, explaining each before it happens. There is also the full bundle with the reference texts and the safety rules, and — if the package is already installed — scholion skill --full prints the same instruction. The model gets no access to your machine, your profile never leaves it, and the safety rules take precedence over every other instruction the model is given.
scholion/ouroboros_tools.py registers 35 sch_* tools — second opinion on a drug, lab analysis, locus lookup, polygenic scores, longevity, goals and more. Ouroboros discovers tool modules by scanning its own tools package, so one line goes there once: pip install scholion, then put from scholion.ouroboros_tools import get_tools into <ouroboros>/ouroboros/tools/scholion_tools.py and set SCHOLION_REPO_DIR. The module itself is not copied, so an upgrade needs nothing more. python3 -m scholion.ouroboros_tools prints the tool list on its own, outside Ouroboros.
New in 0.5.3. The folder agent-plugin/ is Scholion in the Agent Plugins format, which five vendors agreed on in August 2026. ChatGPT desktop, Codex, Cursor, VS Code, GitHub Copilot and Kiro read it. One import brings the tools and the instruction that governs them, and the two cannot be installed apart. The package installs nothing: install the engine once with pipx install scholion, and if the engine is missing, the launcher refuses and prints that command. ChatGPT marks the plugin Desktop only, and that is correct: the engine reads files on your own disk.
scholion mcp serves the same 35 tools over standard input and output. It opens no port and needs no key. It speaks the July 2026 revision of the protocol, which has no handshake, and the handshake revisions back to 2024. Thirteen tools also return their answer as a structure: the fields their command prints with --json, plus the report with its qualifications. An assistant no longer has to parse text back into numbers.
All of them run the same analysis. The one difference worth knowing before you pick: pip install scholion and the unpacked folder both give you the full application, but genome-preparation tooling — turning raw sequencer reads into a VCF — ships only in the source tree, alongside the external bioinformatics tools it orchestrates. If you already have a VCF, the pip package is everything this page shows.