local-first · genome · labs · prescriptions · wearables

Scholion

install it as scholion — the command line also answers to crossread: read your sources against each other

A local engine that reads your genome, lab forms, prescriptions and wearable data against each other and shows where every statement came from. A scholium is a note in the margin of a manuscript saying where something was taken from: the program does not replace the source, it hangs provenance off it.

🔒 Your files stay on your machine🧬 A full genome, and chips 🧪 Labs and trends💊 Prescriptions and interactions ◎ Fifteen body systems⌚ Wearables🤖 An assistant, if you want one 📦 Open source, Apache-2.0🔗 Source on GitHub
🧩 The fastest way in — no terminal needed

Download the skill file

Attach it to a chat with Claude or ChatGPT and say «set this up» — the model reads it and leads, one small step at a time, explaining each before it happens. Your data stays on your machine; the safety rules outrank every other instruction the model is given.

And if you run Ouroboros Desktop, there is nothing to install by hand: the skill has been accepted into OuroborosHub and installs from there in one click (Skills → OuroborosHub → Install), pip package and all.

Want the full picture instead? the complete bundle with the reference texts  ·  every way to install, for people at home in a terminal

⚠️ A research and educational tool. Not a medical device: it does not diagnose, does not start or stop therapy, does not adjust doses and does not replace a physician. Everything it produces is material for your own study and for a conversation with your doctor.

What is different

Everyone sells confidence. Here you get honesty about coverage

A report where everything is green is easy to sell. What is hard is saying what your data does not entitle you to claim — and that is what decides whether the rest is worth believing.

📏

“No findings” means “none in the part that was read”

Coverage is measured per gene: a gene read at 70 % yields the same zero as a gene read at 100 %. Large deletions are not seen by short reads at all, and that is said plainly rather than passed over.

🧭

A frightening label is not yet your diagnosis

Before showing a “pathogenic” variant from a database, the engine checks: is such a condition plausible given your picture, are enough copies of the gene affected, does the variant work in the direction assumed, is the effect distinguishable from the method's own error.

Nothing is filled in for you

No reference line on your form — the marker is shown without a flag, not with a range borrowed from elsewhere. No complete panel — no biological age. A gap stays a gap.

📉

Next to the number, how much to trust it

Every genetic statement carries an evidence level. A polygenic percentile depends on the model and the reference population; it is a position in a population, not a probability of disease, and that is written next to the number.

🚦

Few red marks — and each one earns its place

When half the rows are highlighted, people stop looking at highlighting. A threshold that fires on almost everyone is treated as broken and rebuilt: it measures a property of the data rather than of the objects.

↩️

A retracted conclusion is called retracted

The changelog answers not “which files were touched” but what changed in the conclusions and what was withdrawn. An old formulation lives on in your head until it is explicitly taken back.

How it is built

Code computes, the assistant phrases — and the assistant is optional

Numbers, flags and connections are computed by local code: reproducibly and checkably. A language model adds only the wording — a coherent reading, priorities, questions for the doctor. There is not a single call to a model anywhere in the core, and the application proves it by scanning its own source right on the Assistant tab — and a separate Guide tab explains, screen by screen, what every colour and label means, so a screen is never left unexplained just because the source is not at hand.

1 · Your files

VCF, PDF lab forms, prescriptions, a wearable export — they sit in your own folder.

2 · The engine

Flags, trends, CPIC phenotypes, ClinVar findings, PGS percentiles, coverage.

3 · The assistant

Optional: a reading with sources and questions for the doctor. Any model.

4 · Ways in

The web interface, the command line (scholion / crossread), a skill — in a chat or in the shared skills folder — an MCP tool server, an Ouroboros plugin, and a plugin package for ChatGPT desktop, Codex and Cursor.

Scholion · demo profile
The Assistant page, in the menu: the application checks itself — it scans its own code and prints how many calls to language models it contains (zero). Below it, the exact addresses it can reach and only on request.
The Assistant page, in the menu: the application checks itself — it scans its own code and prints how many calls to language models it contains (zero). The only addresses it can reach, and only on request: api.cpicpgx.org, api.mymemory.translated.net, mor.nlm.nih.gov, rest.ensembl.org, rxnav.nlm.nih.gov, translate.googleapis.com.
Scholion · demo profile
The Guide, in the menu: what every colour, badge and label in the interface means, explained once and in place.
The Guide, in the menu: what every colour, badge and label means — explained once, in place, so a screen is never left unexplained just because the source is not at hand.

Who it is for

Four different people — four different first screens

One product, but everyone has their own reason to open it a second time. Below, honestly, what each of them gets today.

Someone who owns their genome

You have a VCF or a BAM from whole-genome sequencing and you do not want to hand it to somebody else's cloud. You get: ClinVar findings by tier, ACMG secondary findings, pharmacogenomics down to star alleles and HLA, polygenic scores — and a list of what your data does not let you say. A consumer test export is read as well — 23andMe, AncestryDNA, MyHeritage, Living DNA, FamilyTreeDNA, including inside the archive the provider hands you — and the class of the input is measured while it is read: a chip is not passed off as a whole genome.

An engineer who keeps their own things at home

Nextcloud, Home Assistant, Nightscout — and health as one more layer that should not depend on somebody else's server. The core runs on the Python standard library plus a single dependency for reading PDF lab forms (pdfplumber); there is not one network call for any computation, and the whole thing is readable in an evening.

Someone living with a chronic line

Thyroid, iron, lipids, carbohydrate metabolism. A panel every quarter, prescriptions, follow-up draws. Here the benefit arrives on the day of the appointment: the new panel checked against the last one, what to monitor on the current regimen, which questions to ask.

A clinician-researcher or bioinformatician

The knowledge base lives apart from the code; every entry carries its source and its evidence rule, and an edit to the base is versioned as an event of its own. A disputable threshold is visible because it is written as a number in a file rather than hidden in code.

New in 0.5.0 and 0.5.1 — the main thing

The body as fifteen systems, not as a stack of test results

The thyroid used to appear on the page three times — among the deviations, in the line about a drug, and in the list of what to test — and nowhere was it said that these are one subject. A body system is now the unit things assemble around. The tab is called Radar: the figure, two rings per system, one block each, and the full card one click away. Since 0.5.1 every segment also carries a panel of positions somebody wrote, standing in front of the base.

scholion system lipids · demo profile
System: Lipids · patient's register

this build cannot answer for this list — the genome is not
readable on this profile — see the genome status

1. Laboratory now   100/100 — markers measured: 3 of 4
   never taken: Cholesterol, total
2. Movement         100 → 100 (0) against 2025-11
3. Genetics         base GenCC 2026-09-06 — genes: 38;
   positions of the curated panel: 7; read: 0, unread: 0
   positions whose phrase is signed by the panel's author,
   and by no clinician: 7 (signed 2026-09-13)
4. Prescriptions · 5. Clinician's target · 6. What to test · 7. Questions

Next step, in three baskets
   test · read in the genome · ask the clinician
   no genome is attached — a full genome would close for this
   system: 40 genes of the list not read here

The synthetic demo profile of a fictional person, shipped with the package: scholion system lipids. That profile has no genome, so nothing of the genetic half is read — and the card says so rather than reporting silence as health. The same card comes back from a click on a radar segment, from an organ on the figure, from the local API and from the assistant's tool.

Two rings, never merged

How much of the laboratory panel has been measured, and how much of the genetic half has been read. They are different quantities: somebody who has taken the whole panel and had no gene read should not see what somebody in the opposite position sees. A system whose genetic list is not composed shows that ring dotted — an absence, not a zero.

🧬

The base of the genetic half has a version

1469 genes across the fifteen systems from the Gene Curation Coalition export: who asserted the gene–disease link, how strongly, under which mode of inheritance, and on what date. Two groups disagreeing about one gene are shown disagreeing rather than averaged. An assertion classified Limited, Disputed or Refuted is never presented as a finding, and one copy of an allele in a recessive gene is carriership and a question — never a risk line.

In front of the base, a panel somebody wrote

214 positions across fourteen of the systems, authored rather than generated. The unit of a row is a position, not a gene — an rsID with its HGVS, the allele the author named, the phrase for one copy and for two. A row whose phrase for the state actually found is missing is kept and printed as pending: somebody put the position here, and what follows from that genotype is still to be written.

What it looks like in use

Five scenarios

Every number in these screenshots comes from the synthetic demo profile of a fictional person, which ships with the package. You can look at the product before loading anything of your own: pip install scholion && scholion init --demo && scholion serve.

1

One picture instead of a stack of PDFs

Markers over years, flags, movement, the link to the genome — and a focus of attention: the one thing you are working on right now.

trendsranges from your own formsfocus of attention
Scholion · demo profile
Overview: the numbers worth seeing first, the body figure beside the radar, and what is out of range now (demo).
Overview: the numbers worth seeing first, the body figure beside the radar, and what is out of range now (demo).
Scholion · demo profile
Labs: what is out of range first, with flags, sparklines and movement; the values in range are folded (demo).
Labs: what is out of range first, with flags, sparklines and movement; the values in range are folded (demo).
Scholion · demo profile
Lifestyle: the few numbers first, then anthropometry, activity and recovery with their trends (demo).
Lifestyle: the few numbers first, then anthropometry, activity and recovery with their trends (demo).
2

Going to the doctor with the question already prepared

For a drug in your regimen — what the guidelines say as applied to your data: pharmacogenomics, what to monitor, interactions with what you already take. The result is a list of questions, not of instructions.

CPIC with the citationinteractionswhat to monitor
  • The system does not prescribe and does not cancel: it quotes the source and formulates the question; the decision stays with the physician.
  • You can see what is missing for an answer: an unknown genotype is called unknown, not “normal”.
Scholion · demo profile
Medicines: one drug checked against genome × labs × interactions × ClinVar (demo).
Medicines: one drug checked against genome × labs × interactions × ClinVar (demo).
Scholion · demo profile
Radar: the block of one system — the verdict, the measurements as a table, the genotype as findings, the questions for the doctor (demo).
Radar: the block of one system — the verdict, the measurements as a table, the genotype as findings, the questions for the doctor; everything else is under «More» (demo). More in the section above.
3

The genome, read with its caveats

ClinVar findings laid out by clinical weight; polygenic risks as navigation for screening; the longevity layer as context rather than a verdict.

ClinVar × your VCFACMG SFPGS percentilesper-gene coverage
Scholion · demo profile
Genome: findings by tier, polygenic risks and the longevity layer — with an evidence level on every statement (demo).
Genome: findings by tier, polygenic risks and the longevity layer — with an evidence level on every statement (demo).
4

What to test next time

The form for the next draw assembles itself — from the current regimen, the clinical thresholds you have crossed, what you agreed with your doctor, and what is missing for the computations.

in stepsage of the last valuetube and preparation
  • In steps, not as one long sheet: the expensive test only if the cheap one showed something.
  • With the age of the last value, so you do not pay twice for what you had done a month ago.
  • Computed indices are not ordered — their inputs go on the form instead.
Scholion · demo profile
What to test: the suggestions with their reasons and priority, at the end of the Radar (demo).
What to test: the suggestions with their reasons and priority, at the end of the Radar (demo).
Scholion · demo profile
Medicines: the current regimen as the single point of truth for every check; what was stopped is folded below (demo).
Medicines: the current regimen as the single point of truth for every check; what was stopped is folded below (demo).
5

Testing your own hypotheses

“I think late coffee makes me sleep worse” is a testable statement. The n-of-1 module turns it into an experiment with the protocol fixed in advance.

pre-registrationrandomised blockspermutation testprotocol breaks accounted for
  • First, whether the design can answer at all: sensitivity is bounded by the number of blocks, not the number of days. A classic four-block ABAB can never reach significance, and that is printed before the trial starts.
  • Comparison by periods, not by days: neighbouring nights resemble each other, and counting day by day turns noise into a discovery.
  • Protocol breaks are cut out along with the following day, and a day with no entry counts as unknown rather than as adhered to.

Why come back

It is not the genome that changes, it is what is known about it

A local tool has no subscription and no notifications. There are exactly three reasons to open it again, and all three arrive on their own.

What arrives by itselfWhat Scholion does with it
A new ClinVar releaseA scheduled monthly reanalysis: the fresh database, a diff against your variants, a repeat ACMG scan, a review of PGS candidates. A report of the form “this many were reclassified last quarter, and these are yours”.
A new lab panelThe new panel read, checked against the previous one, and the checklist for the next draw updated. Forms keep arriving; the genome is loaded once.
A new prescriptionChecked against the current regimen and the pharmacogenomics, with a list of questions ready for the date of the appointment.

Privacy & licences

Nothing leaves without an action of yours

🔒

Local by construction

Genome, labs, prescriptions and metrics are stored only on your machine. The code reaches outward only to public reference databases (Ensembl, RxNorm, CPIC) and translation services (for cross-language drug names), and only for general information — your data never goes into the request. SCHOLION_OFFLINE=1 forbids even that.

📦

A package with its own audit

The depersonalised build is produced by a sanitiser: if any sign of personal data reaches it, the build stops. The package holds the code, the knowledge bases, the tests and the demo profile.

🧪

Checkability

More than a thousand tests on the standard library, a backward-compatibility check of the public contract (56 commands, 20 snapshots), and a separate set of rules about the safety of the conclusions. Anyone who takes the package runs them with one command — inside the package, not only in our repository.

⚖️

Licences

Code — Apache-2.0; the curated knowledge base — CC BY 4.0. The provenance of every file that is not ours is recorded in ATTRIBUTION.md; data whose licence forbids commercial use is not bundled at all.

⌨️

Installation

Six doors, one core — a zero-install folder, pip install scholion, a skill for a model with no terminal, an MCP tool server, an Ouroboros plugin, or a plugin package for ChatGPT desktop, Codex and Cursor. Exact, tested commands are in the Installation section below.

🌍

Two languages

Reports print English by default and switch with --lang ru or SCHOLION_LANG=ru. Recognition of Russian lab forms is not a setting but a feature: a marker's name comes from the form in your hand, so it is shown in the language it was printed in.

The assistant is any model you like. A skill for Claude, an Ouroboros plugin and a plugin package for ChatGPT, Codex and Cursor ship with it; for anything else, scholion skill prints the instruction and scholion assistant --context assembles a snapshot of the state with the list of commands. Through a skill, the model gets no access to your machine — it asks you to run a command and reads the output. Through the tool server, it calls tools that run on your machine and read only your own data.

Data used for testing

Tested on data that is not ours

A tool tested only on the person who built it passes its own tests and fails on the world. Three open sources make the testing possible — and not one of them is republished here.

🧬

Personal Genome Project (Harvard)

The standing reference run covers 36 sets from eleven providers — chips, whole genomes, panels and split call sets. PGP participants consented to publishing their data openly precisely so that it would be worked on — and without it this product would only ever have seen the author's own machine. Thank you. Nothing of theirs lives here: no genotypes, no findings, no identifiers, no medical records. What travels out of that work is the behaviour of the engine and aggregate numbers.

🧾

Synthea

A synthetic FHIR R4 bundle the import is tested against: a portal export with clinical records, produced by a real system rather than by the same person who wrote the parser. The patient is generated — which is why the file may live in the repository at all: nobody's consent is involved. Apache-2.0.

🔍

Genomi

The input-format detector — detection.py and text_io.py — vendored under Apache-2.0 with attribution, every change marked in place and a way to update it. It decides what a file is by its content rather than by its name: BAM, CRAM, VCF, gVCF told apart from VCF, paired FASTQ, and the exports of 23andMe, AncestryDNA, MyHeritage, FamilyTreeDNA and Living DNA — gzip, zip and tar wrappers included.

Full provenance for every third-party file is in ATTRIBUTION.md, and a single command compares the vendored code with upstream and reports drift in either direction.

Installation

One core, six doors — pick the one you already have

The same analysis runs underneath all six; what changes is what you need to already have, and who ends up typing the commands.

terminal
pip install scholion
scholion init --demo
scholion overview
scholion serve

Python 3.10+ and nothing else needed first. This is the pip package — it gets you the local web interface and the full command line in one step; pdfplumber, needed for reading PDF lab forms, comes with it. The command line answers to two names, scholion and crossread, both installed by the same package. An optional pip install "scholion[genome]" adds faster VCF access via pysam; without it the built-in reader still works, just slower.

📁

No installation — a folder you unpack

Python 3.10+ is the only requirement; every line of analysis runs on the standard library. (Reading PDF lab forms is the exception — that needs pdfplumber, which this delivery does not bundle.) Start with ./bin/crossread --help. This is also where the genome-preparation tooling lives — FASTQ → VCF, PharmCAT/PyPGx — for building a VCF from raw reads; scholion tools reports which external programs (bcftools, samtools) that needs and how to get them.

🧩

A skill, for a language model

The easiest way in — no terminal skills needed: download the skill file and attach it to your chat with Claude or ChatGPT, then say «set this up». The model reads it and takes it from there — one small step at a time, explaining each before it happens. There is also the full bundle with the reference texts and the safety rules, and — if the package is already installed — scholion skill --full prints the same instruction. The model gets no access to your machine, your profile never leaves it, and the safety rules take precedence over every other instruction the model is given.

🔌

A plugin for Ouroboros

scholion/ouroboros_tools.py registers 35 sch_* tools — second opinion on a drug, lab analysis, locus lookup, polygenic scores, longevity, goals and more. Ouroboros discovers tool modules by scanning its own tools package, so one line goes there once: pip install scholion, then put from scholion.ouroboros_tools import get_tools into <ouroboros>/ouroboros/tools/scholion_tools.py and set SCHOLION_REPO_DIR. The module itself is not copied, so an upgrade needs nothing more. python3 -m scholion.ouroboros_tools prints the tool list on its own, outside Ouroboros.

🧳

A plugin package — ChatGPT, Codex, Cursor

New in 0.5.3. The folder agent-plugin/ is Scholion in the Agent Plugins format, which five vendors agreed on in August 2026. ChatGPT desktop, Codex, Cursor, VS Code, GitHub Copilot and Kiro read it. One import brings the tools and the instruction that governs them, and the two cannot be installed apart. The package installs nothing: install the engine once with pipx install scholion, and if the engine is missing, the launcher refuses and prints that command. ChatGPT marks the plugin Desktop only, and that is correct: the engine reads files on your own disk.

🔗

An MCP tool server

scholion mcp serves the same 35 tools over standard input and output. It opens no port and needs no key. It speaks the July 2026 revision of the protocol, which has no handshake, and the handshake revisions back to 2024. Thirteen tools also return their answer as a structure: the fields their command prints with --json, plus the report with its qualifications. An assistant no longer has to parse text back into numbers.

All of them run the same analysis. The one difference worth knowing before you pick: pip install scholion and the unpacked folder both give you the full application, but genome-preparation tooling — turning raw sequencer reads into a VCF — ships only in the source tree, alongside the external bioinformatics tools it orchestrates. If you already have a VCF, the pip package is everything this page shows.