Help From Our Friends · an experiment in visualizing open knowledge, by Luis Villa

Wikipedia is not alone.

Alongside every Wikipedia article there is a wider open world: libraries that lend, museums that publish their own collections, scientists who post their papers openly, mappers and naturalists who chart the planet for free. This experiment invites them in. Pick an article, and see who else is out there.

A variety of articles, chosen to show off both the cool stuff they bring in, and the open institutions and technologies they build on:

Frequently asked questions

Why did you build this?

I build this for many different reasons, which makes it hard to explain. But among others:

  • “Open knowledge” has become such a diffuse thing that it is hard even for its advocates to visualize it. I wanted something that shows it whole.
  • Wikipedia is isolated, and that’s a problem. The best way I know to fix that is by demonstrating what would be cool about a wiki with strong ties to the rest of open.
  • I wanted to understand the possibilities of a tightly-knit open. The best way I know to learn has always been to learn-by-doing, and that’s more fun than ever right now.

How does it work?

When you load an article, this service scans it and extracts Wikidata information, citations, and other sources of metadata. It then uses those data points to find what our friends have to say.

This is not trivial, so it can be slow and would need to be re-engineered to work on the real Wikipedia. But the information is all as real and accurate as Wikipedia and our friends have made it.

Can this protect Wikipedia from LLMs?

Maybe! One theory of the near-future is that human curation will be rarer, but more valuable. If that’s true, one way to strengthen open knowledge might be to highlight — and tighten — the connections between curators.

We’re blessed to live in a time of great abundance of such people. iNaturalist has built an awesome community — we should elevate them as peers in our work. The Met and the Rijksmuseum have some of the most skilled curators on the planet, and their material is freely given to all of us. A world in which we treat them as peers who work with Wikipedia, rather than just sources to import, will be a better one for open knowledge.

Most Wikipedia articles contain one or more citations and wikilinks, each pointing at something outside the encyclopedia. Two kinds of statement turn a pointer into a card:

Identifier

The article states an ISBN, DOI, OCLC, LCCN, PMID or arXiv id, and a collection answers to exactly it. The strongest claim a card can make.

Statement

Wikidata states the connection outright — this painting is Met object 11417, this species is iNaturalist taxon 48662, this place is here. The card credits the property.

Each group of results also says who asked: when one friend answers several of the article’s links, its results split into one labeled group per link — and in the opening section, works by the subject are kept separate from works merely cited there.

Who are the friends, and how is their work licensed?

Books and papers

Internet Archive

Books you can borrow, discovered through a footnote’s ISBN.

openness? Public-domain scans free to read; in-copyright books lent, not copied.

Open Library

A book’s editions, and which are free to read, discovered through its ISBN.

openness? Open bibliographic data, downloadable in bulk.

OpenAlex

A free, legal copy of a cited paper, discovered through its DOI or PMID.

openness? Catalog CC0. Only papers with an open copy are shown — each card names its license; closed ones are counted, not carded.

arXiv

Preprints in physics, maths and computing, discovered through the arXiv id in a citation.

openness? Metadata CC0; each paper names its own license. in their words

Museums and image collections

The Met

The museum’s own record of an object — title, artist, date, and often an image — discovered through a Wikidata statement naming it.

openness? Public-domain works released CC0, images included. in their words

Art Institute of Chicago

The museum’s own record of a painting — title, artist, date, and often an image — discovered through a Wikidata statement naming it.

openness? Public-domain images CC0, served over open IIIF. in their words

Rijksmuseum

The museum’s own record of a work — title, date, and a photograph at full resolution — discovered through a Wikidata statement naming it.

openness? Works out of copyright carry the public-domain mark; images served over open IIIF, catalog data CC0. in their words

Cleveland Museum of Art

The museum’s own record of a work — title, date, and a photograph — discovered through a Wikidata statement naming it.

openness? Works out of copyright are released CC0 — data and images both; the per-object share_license_status flag is the museum’s own word. in their words

J. Paul Getty Museum

The museum’s own record of a work — title, date, and a photograph over IIIF — discovered through a Wikidata statement naming it.

openness? Open Content Program images are CC0; each object page states its own license, and the page’s catalog text is CC BY 4.0. in their words

IIIF collections

A manuscript or artwork’s own manifest — title, often an image, and the holding institution’s own credit — discovered through a Wikidata statement naming it.

openness? Terms set per object by its holding institution, stated in each manifest.

the Smithsonian

3D scans and museum records, discovered through the scientific name Wikidata states for a species, or a pair of statements naming the museum and its own accession number.

openness? Open Access items are CC0: no rights reserved at all. in their words

Union catalogs

DPLA

Items from US libraries, archives and museums, discovered through the subject heading a cataloger filed them under.

openness? Metadata CC0; each item’s rights stated by its holder. in their words

Europeana

Items from European museums, libraries and archives, discovered through a Wikidata statement naming Europeana’s own entity for the subject.

openness? Metadata CC0; only openly licensed items are shown, and each card names its license. in their words

DigitalNZ

Items from New Zealand libraries, archives and museums, discovered through that same heading, in the way NZ catalogers spell it.

openness? Each item states in plain words what a reader may do with it — but the API’s metadata is non-commercial by default. in their words

The living world and the map

iNaturalist

Photographs of species, discovered through a Wikidata statement naming the species’ iNaturalist taxon.

openness? Each photo carries its observer’s chosen license; only openly licensed ones are shown here. in their words

GBIF

Maps of where a species has been recorded, discovered through a Wikidata statement naming its GBIF dataset.

openness? Records CC BY-NC, CC BY or CC0, stated per dataset. in their words

OpenStreetMap

A map of a place, discovered through the coordinates Wikidata states for it.

openness? Map data ODbL: share-alike, credit the contributors. in their words

The public record

Free Law Project

The court’s opinion in full, discovered through the case citation already in the article.

openness? Court opinions are public domain: nobody owns the law. in their words

Wikidata & Wikipedia

The hosts: one writes the article that convenes everyone; the other makes the introductions — it knows every friend’s name for every thing.

openness? Article text CC BY-SA 4.0; Wikidata CC0.

What are the challenges?

This is a demo and not intended for production. Among other challenges:

There is nowhere for most of this to go

Each article page carries a closed panel — Who helped, and who Wikipedia doesn’t show — sorting the friends who filled it into three states: shown and credited, a link only, or invisible. Most data points are either ‘link only’ or ‘invisible’.

This is because Wikipedia requires most external links to be fairly plain, and because external media must be hosted on Commons. There are good reasons for both of these rules, but they make it hard to surface information in rich ways, and make it hard to be a good partner to our friends.

Page layout

Arbitrary content means great layout is somewhere between difficult and impossible. Work with designers on this challenge would be necessary (though even rudimentary implementations, like this one, would likely be very enjoyable for certain types of data nerds!)

Content curation

Sources can return thousands of responses. (Think the Smithsonian on the Apollo Program, for example.) A gallery with a thousand items is not very helpful to the reader, so some sort of curation (or at least ability to tune algorithmic prioritization) would be necessary before widespread deployment.

Source curation

Similarly, there are many collections of open content these days. Picking and prioritizing them would be an important challenge if we wanted to expand this.

Metadata gaps

Metadata quality leaves a fair amount to be desired. For example, Internet Archive’s recent scan of thousands of theses will be nice sources of information for articles — once it has metadata. Ideally the fix is to deploy Wikipedian energy to other repositories to improve the metadata, not have it curated only inside Wikipedia.

Rights are complicated at best, murky or unknown at worst

Some items arrive with an honest non-answer: the institution has recorded that the rights status is unknown, or not yet evaluated. These render here with a small ? mark and the institution’s own words behind a click — treated, for now, as peers of the openly licensed material, because a recorded open question is a fact about the collection and silence would hide it. At scale this is a real challenge: a reader wants to know what they may do, and “nobody knows” satisfies no one. The durable fix is rights-clearing work of the kind CopyClear and Dominio Público en América Latina do on Wikidata; a demo can only keep the question visible.

Open data increasingly arrives through priced pipes

The catalog behind every paper card here — OpenAlex — is free to download and openly licensed, but as of February 2026 the convenient way to read it, its API, requires a key and bills by usage. A demo like this fits comfortably inside the free daily allowance, and charging for a service while keeping the data open is a defensible way to keep the lights on. But expect more of this: running an API costs money that open licenses do not pay, so even institutions with genuinely open data will increasingly meter or put terms on the pipe — DigitalNZ’s metadata API, non-commercial by default, is the same problem in a different shape. Anything Wikipedia-scale built on lookups like these would need formal agreements, or its own copies of the open datasets, rather than goodwill rate limits.

Bot volume and caching

Because of the volume of Wikipedia, to be deployable at any sort of scale, this would likely need extensive caching and likely formal agreements with the other data providers.