Frequently asked questions
- Why did you build this?
- How does it work?
- Can this protect Wikipedia from LLMs?
- How do you link the article to knowledge from other sources?
- Who are the friends, and how is their work licensed?
- What are the challenges?
Why did you build this?
I build this for many different reasons, which makes it hard to explain. But among others:
- “Open knowledge” has become such a diffuse thing that it is hard even for its advocates to visualize it. I wanted something that shows it whole.
- Wikipedia is isolated, and that’s a problem. The best way I know to fix that is by demonstrating what would be cool about a wiki with strong ties to the rest of open.
- I wanted to understand the possibilities of a tightly-knit open. The best way I know to learn has always been to learn-by-doing, and that’s more fun than ever right now.
How does it work?
When you load an article, this service scans it and extracts Wikidata information, citations, and other sources of metadata. It then uses those data points to find what our friends have to say.
This is not trivial, so it can be slow and would need to be re-engineered to work on the real Wikipedia. But the information is all as real and accurate as Wikipedia and our friends have made it.
Can this protect Wikipedia from LLMs?
Maybe! One theory of the near-future is that human curation will be rarer, but more valuable. If that’s true, one way to strengthen open knowledge might be to highlight — and tighten — the connections between curators.
We’re blessed to live in a time of great abundance of such people. iNaturalist has built an awesome community — we should elevate them as peers in our work. The Met and the Rijksmuseum have some of the most skilled curators on the planet, and their material is freely given to all of us. A world in which we treat them as peers who work with Wikipedia, rather than just sources to import, will be a better one for open knowledge.
How do you link the article to knowledge from other sources?
Most Wikipedia articles contain one or more citations and wikilinks, each pointing at something outside the encyclopedia. Two kinds of statement turn a pointer into a card:
Identifier
The article states an ISBN, DOI, OCLC, LCCN, PMID or arXiv id, and a collection answers to exactly it. The strongest claim a card can make.
Statement
Wikidata states the connection outright — this painting is Met object 11417, this species is iNaturalist taxon 48662, this place is here. The card credits the property.
Each group of results also says who asked: when one friend answers several of the article’s links, its results split into one labeled group per link — and in the opening section, works by the subject are kept separate from works merely cited there.
Who are the friends, and how is their work licensed?
Books and papers
Internet Archive
Books you can borrow, discovered through a footnote’s ISBN.
openness? Public-domain scans free to read; in-copyright books lent, not copied.
Open Library
A book’s editions, and which are free to read, discovered through its ISBN.
openness? Open bibliographic data, downloadable in bulk.
OpenAlex
A free, legal copy of a cited paper, discovered through its DOI or PMID.
openness? Catalog CC0. Only papers with an open copy are shown — each card names its license; closed ones are counted, not carded.
arXiv
Preprints in physics, maths and computing, discovered through the arXiv id in a citation.
openness? Metadata CC0; each paper names its own license. in their words
Museums and image collections
The Met
The museum’s own record of an object — title, artist, date, and often an image — discovered through a Wikidata statement naming it.
openness? Public-domain works released CC0, images included. in their words
Art Institute of Chicago
The museum’s own record of a painting — title, artist, date, and often an image — discovered through a Wikidata statement naming it.
openness? Public-domain images CC0, served over open IIIF. in their words
Rijksmuseum
The museum’s own record of a work — title, date, and a photograph at full resolution — discovered through a Wikidata statement naming it.
openness? Works out of copyright carry the public-domain mark; images served over open IIIF, catalog data CC0. in their words
Cleveland Museum of Art
The museum’s own record of a work — title, date, and a photograph — discovered through a Wikidata statement naming it.
openness? Works out of copyright are released CC0 — data and images both; the per-object share_license_status flag is the museum’s own word. in their words
J. Paul Getty Museum
The museum’s own record of a work — title, date, and a photograph over IIIF — discovered through a Wikidata statement naming it.
openness? Open Content Program images are CC0; each object page states its own license, and the page’s catalog text is CC BY 4.0. in their words
IIIF collections
A manuscript or artwork’s own manifest — title, often an image, and the holding institution’s own credit — discovered through a Wikidata statement naming it.
openness? Terms set per object by its holding institution, stated in each manifest.
the Smithsonian
3D scans and museum records, discovered through the scientific name Wikidata states for a species, or a pair of statements naming the museum and its own accession number.
openness? Open Access items are CC0: no rights reserved at all. in their words
Union catalogs
DPLA
Items from US libraries, archives and museums, discovered through the subject heading a cataloger filed them under.
openness? Metadata CC0; each item’s rights stated by its holder. in their words
Europeana
Items from European museums, libraries and archives, discovered through a Wikidata statement naming Europeana’s own entity for the subject.
openness? Metadata CC0; only openly licensed items are shown, and each card names its license. in their words
DigitalNZ
Items from New Zealand libraries, archives and museums, discovered through that same heading, in the way NZ catalogers spell it.
openness? Each item states in plain words what a reader may do with it — but the API’s metadata is non-commercial by default. in their words
The living world and the map
iNaturalist
Photographs of species, discovered through a Wikidata statement naming the species’ iNaturalist taxon.
openness? Each photo carries its observer’s chosen license; only openly licensed ones are shown here. in their words
GBIF
Maps of where a species has been recorded, discovered through a Wikidata statement naming its GBIF dataset.
openness? Records CC BY-NC, CC BY or CC0, stated per dataset. in their words
OpenStreetMap
A map of a place, discovered through the coordinates Wikidata states for it.
openness? Map data ODbL: share-alike, credit the contributors. in their words
The public record
Free Law Project
The court’s opinion in full, discovered through the case citation already in the article.
openness? Court opinions are public domain: nobody owns the law. in their words
Wikidata & Wikipedia
The hosts: one writes the article that convenes everyone; the other makes the introductions — it knows every friend’s name for every thing.
openness? Article text CC BY-SA 4.0; Wikidata CC0.
What are the challenges?
This is a demo and not intended for production. Among other challenges:
There is nowhere for most of this to go
Each article page carries a closed panel — Who helped, and who Wikipedia doesn’t show — sorting the friends who filled it into three states: shown and credited, a link only, or invisible. Most data points are either ‘link only’ or ‘invisible’.
This is because Wikipedia requires most external links to be fairly plain, and because external media must be hosted on Commons. There are good reasons for both of these rules, but they make it hard to surface information in rich ways, and make it hard to be a good partner to our friends.
Page layout
Arbitrary content means great layout is somewhere between difficult and impossible. Work with designers on this challenge would be necessary (though even rudimentary implementations, like this one, would likely be very enjoyable for certain types of data nerds!)
Content curation
Sources can return thousands of responses. (Think the Smithsonian on the Apollo Program, for example.) A gallery with a thousand items is not very helpful to the reader, so some sort of curation (or at least ability to tune algorithmic prioritization) would be necessary before widespread deployment.
Source curation
Similarly, there are many collections of open content these days. Picking and prioritizing them would be an important challenge if we wanted to expand this.
Metadata gaps
Metadata quality leaves a fair amount to be desired. For example, Internet Archive’s recent scan of thousands of theses will be nice sources of information for articles — once it has metadata. Ideally the fix is to deploy Wikipedian energy to other repositories to improve the metadata, not have it curated only inside Wikipedia.
Rights are complicated at best, murky or unknown at worst
Some items arrive with an honest non-answer: the institution has recorded that the rights status is unknown, or not yet evaluated. These render here with a small ? mark and the institution’s own words behind a click — treated, for now, as peers of the openly licensed material, because a recorded open question is a fact about the collection and silence would hide it. At scale this is a real challenge: a reader wants to know what they may do, and “nobody knows” satisfies no one. The durable fix is rights-clearing work of the kind CopyClear and Dominio Público en América Latina do on Wikidata; a demo can only keep the question visible.
Open data increasingly arrives through priced pipes
The catalog behind every paper card here — OpenAlex — is free to download and openly licensed, but as of February 2026 the convenient way to read it, its API, requires a key and bills by usage. A demo like this fits comfortably inside the free daily allowance, and charging for a service while keeping the data open is a defensible way to keep the lights on. But expect more of this: running an API costs money that open licenses do not pay, so even institutions with genuinely open data will increasingly meter or put terms on the pipe — DigitalNZ’s metadata API, non-commercial by default, is the same problem in a different shape. Anything Wikipedia-scale built on lookups like these would need formal agreements, or its own copies of the open datasets, rather than goodwill rate limits.
Bot volume and caching
Because of the volume of Wikipedia, to be deployable at any sort of scale, this would likely need extensive caching and likely formal agreements with the other data providers.