
The wiki changed while I was reading it
I built a cognee connector that turns any MediaWiki, from Wikipedia to an internal company wiki, into a knowledge graph you can query. Ingesting was easy. Staying in sync was the real work.
- mediawiki
- wikipedia
- knowledge-graph
- cognee
- python
- open-source
- rag
Every company I've seen has a wiki. Runbooks, onboarding docs, the page about why the billing service has that one weird retry loop. Years of stuff nobody wants to lose.
And almost nobody reads it. Search is keyword-only, so you have to already know what the page is called to find it. The answer to your question is usually split across three pages written by three people who've since left. Half of them are out of date.
The knowledge is in there. The wiki just can't connect it for you.
So for the Mergetober hackathon I wrote a MediaWiki connector for cognee. Point it at a wiki, it pulls the pages into a knowledge graph, and you ask questions in plain English. MediaWiki is what Wikipedia runs on, and it's also what a lot of internal wikis run on, so one connector covers both.
cognee — an open-source memory layer for AI apps. You give it documents, it extracts entities and relationships into a knowledge graph, and you query that graph with natural language.
MediaWiki — the wiki software behind Wikipedia, Fandom, and thousands of private company wikis. Every install exposes the same Action API at
/w/api.php.
Pulling pages in was the easy bit. The hard bit, and most of this post, was something I didn't expect going in: a wiki is never finished. Someone is always editing, renaming, deleting. A connector that only works on day one is a snapshot, not a sync.
Five lines, one question
Here's the whole thing from the user's side. Three Wikipedia pages, one question:
import cognee
from cognee_community_connector_mediawiki import mediawiki_source
await cognee.remember(
mediawiki_source(
"https://en.wikipedia.org/w/api.php",
titles=["Ada Lovelace", "Charles Babbage", "Analytical engine"],
user_agent="my-app/1.0 (me@example.com)", # Wikimedia wants contact info
),
dataset_name="wiki",
primary_key="id",
write_disposition="merge",
)
results = await cognee.recall(
"What did Ada Lovelace write about Babbage's analytical engine?", datasets=["wiki"]
)
And the answer it gave back, from the ingested pages:
Ada Lovelace translated Menabrea's description of Babbage's Analytical Engine and added seven extensive notes. She described an algorithm for calculating Bernoulli numbers and argued that the engine could manipulate symbols, including letters and musical notes, not just numbers.
That answer needs facts from more than one page: who Menabrea was, what the Analytical Engine was, what Lovelace's notes contained. Keyword search would get you to the right page. It wouldn't stitch them together.
You aren't limited to a list of titles either. namespaces=[0] takes every article on a wiki, categories=["Physics"] takes everything in a category, and for a private wiki you make a bot password at Special:BotPasswords and pass username and password.
What actually goes into the graph
One document per wiki page. Each one has the page's text, its categories, and its recent edit history (who changed it, when, and the edit summary), because "who last touched the deploy runbook" is a fair question to ask.
I didn't write a wikitext parser, on purpose. Wikitext is full of templates and parser functions that only make sense once the wiki expands them, and anyone who's tried to parse it by hand knows how that goes. The connector asks the wiki to render the page itself, using the TextExtracts extension if it's installed or action=parse if it isn't, and works from what comes back.
action=parse came with its own cleanup job. Real Wikipedia HTML includes inline <style> blocks, hidden short descriptions, [edit] links, [1] reference markers and giant navboxes at the bottom. All of that gets dropped. Infobox facts, headings and lists stay, because that's where the useful facts are.
Three pages turned into this:

500 nodes, 1,731 edges. The pipeline view shows how it's built up: 2 text documents, split into 10 chunks, which produced 475 entities, 2 entity types, and 10 summaries.

Following a wiki that never stops changing
Re-downloading the whole wiki every time is the obvious approach. It's also how you burn through your LLM budget, since every page that goes back into cognee gets its entities extracted again.
MediaWiki has a feed for this called recentchanges. It lists every edit, page creation and log action (deletes, moves, merges) since a given time. So the plan is: remember where you stopped, ask for everything after that, and re-check only those pages.
I believed the plan would work as written for about an hour. Here's what actually happened.
The feed isn't in order
The MediaWiki docs say it quietly: recentchanges entries can show up slightly out of timestamp order. So if your cursor says "I've seen everything up to 10:00:00" and an entry stamped 09:59:58 lands after you read, you'll never see it. That edit is just gone.
The fix is to not trust the cursor exactly. Every run replays from the last server time minus an overlap window, then re-checks the pages it finds. That means some pages get looked at twice, which would be a problem if re-checking had side effects. So re-checks compare revision ids first. If the revision hasn't changed, nothing happens.
There's a test that inserts an entry behind the cursor. With no overlap it gets missed. With the overlap it's caught. I like having that one written down, because it's precisely the kind of bug you'd never reproduce by hand.
Deletes live in a different place
Edits show up as normal entries. Deletes, moves and merges show up as log entries, each with a slightly different shape. I checked these against real test.wikipedia.org data rather than trusting the docs:
| Log entry | What it carries |
|---|---|
delete/delete |
the deleted page's pageid (still there after deletion) |
move |
the new title in logparams.target_title |
merge |
the destination in dest_title |
The tempting thing is to write one handler per type. I didn't. For every entry, the connector collects every page the entry mentions and re-checks all of them against the live wiki. Edited? Re-render. Gone? Tombstone it. Came back? Ingest it again. One code path handles edits, deletes, restores, moves and merges, and there's no special case to get wrong.
Rename by id, not by title
This one is easy to miss. If you key documents by title, then renaming "Deploy guide" to "Deploy guide (2025)" looks like one page deleted and a new one created. cognee forgets the old document and re-extracts the new one from scratch, so you pay LLM costs for a page whose content didn't change.
MediaWiki keeps the page id through a rename. So the page id is the key, and a rename is just an update to the title field on the same document.
The sneaky ones
Those three would have been enough for a working connector. These two are why it's a correct one.
Categories change without anyone editing the page
Say you sync categories=["Runbooks"]. Someone edits a template so it adds [[Category:Runbooks]], and that template is used on 200 pages. All 200 are now in the category.
None of those 200 pages was edited. The feed has one entry, for the template. Nothing in it names the pages you now care about.
The only reliable fix I found is boring: every run lists the category's members and diffs that against what the connector already holds. New members get ingested and departed ones get forgotten. I tested both directions on a live wiki, a template adding a category and the same template dropping it, because I didn't trust it until I'd seen it happen.
The feed quietly forgets
recentchanges doesn't keep history forever. $wgRCMaxAge defaults to 90 days, and on Wikimedia sites it's roughly 30. So if your sync doesn't run for five weeks, the feed no longer reaches back to your cursor and you have a gap you can't see.
Worse, the API doesn't tell you the retention window. siteinfo doesn't expose it.
So the connector checks for itself. It looks at the oldest entry the feed still has. If that's newer than the cursor, there's a hole, and the run falls back to a full reconcile: list everything in scope and diff it. That's slower, but it's correct, and it only happens when it has to.
Forgetting is a feature
If a page is deleted on the wiki, it should disappear from the graph too. Otherwise you get confident answers built on a runbook someone deleted because it was wrong.
The connector emits a _deleted tombstone for pages that are gone, and cognee's merge removes that document and its graph content. Here's all three real runs against Wikipedia, with the third one dropping "Analytical engine" from the list of titles:

| Run | What happened |
|---|---|
| 1st | full sync, 3 pages, about 2 minutes, correct answer |
| 2nd (nothing changed) | incremental, 0 re-ingested, 6 requests, about 39 s |
| 3rd (one title removed) | 1 page forgotten, 0 document nodes left for it in the graph |
The second run is the one I care most about. Nothing changed upstream, so nothing got re-extracted. Six API calls and no LLM work.
There's a safety rule that goes with this. If a private wiki is synced without credentials, MediaWiki answers with readapidenied. A careless connector would read that as "zero pages", diff it against the existing state, and delete everything. This one raises an error instead, and there's a test named test_empty_enumeration_never_mass_deletes so that rule stays in place.
Being a decent API citizen
Wikipedia is busy. In one 10-minute window I counted about 1,170 logged changes. My first version re-checked every page in the feed, whether I'd synced it or not, which came to around 45 requests for one incremental run.
Most of those pages had nothing to do with my three titles. Filtering the feed down to pages that could actually matter (in scope, or already held) took a run from ~45 requests to ~8. On a quiet local wiki, a no-change run costs 3.
Beyond that, it does the basics Wikimedia asks for: a real User-Agent with contact info, maxlag so it backs off when the servers are under load, and retries with a limit for transient failures.
Testing it against a real wiki
Unit tests with a fake API only go so far, because the fake behaves however I imagined the API behaves. So I ran a throwaway MediaWiki 1.43 in Docker and used MediaWiki's own maintenance scripts to change the wiki between syncs:

Edit, rename, delete, undelete (same page id comes back), a draft moved into the article namespace, a page moved out of it, a template adding and dropping a category, a forced reconcile, and a private wiki behind a bot password. All of them passed.
Plus 46 tests in the package: a fake Action API for the edge cases, a real dlt merge, and an end-to-end test that deletes a page upstream and checks that its graph content is gone from cognee.

What I'd tell myself before starting
Reading data from an API is a day of work. Staying in sync with it is the actual project, and most of that work is in the places the docs only mention once, or not at all: entries out of order, a feed that gets pruned, category changes no edit ever records.
The approach that held up was the same each time. Don't trust the cursor exactly. Make every re-check cheap and idempotent so doing extra work is harmless. When you can't tell whether you missed something, reconcile the whole scope.
If you keep one thing
Ingesting a wiki is a weekend; following one is the real job, and the trick is making "check again" so cheap and harmless that you never have to be sure you didn't miss anything.
Try it
- PR: topoteretes/cognee-community#294
- Issue: topoteretes/cognee#4724
- Code:
packages/connector/mediawiki/in cognee-community, with a runnableexamples/example.py - More of my open-source work: /open-source
Sources
- My own runs: real English Wikipedia with OpenAI, a local MediaWiki 1.43 in Docker, and the package's test suite. Every number in this post comes from those.
- MediaWiki Action API docs, especially
list=recentchanges(the out-of-order warning) and$wgRCMaxAge. - Log entry shapes checked against live data on test.wikipedia.org, not just the docs.
- cognee and cognee-community.