Denny Vrandečić: Connecting the World’s Knowledge with Abstract Wikipedia – Episode 32

photo of Denny Vrandečić, creator of the Abstract Wikipedia project
Denny Vrandečić

As the founder of Wikidata, Denny Vrandečić has thought a lot about how to better connect the world’s knowledge.

His current project is Abstract Wikipedia, an initiative that aims to let anyone anywhere on the planet contribute to, and benefit from, the world’s collective knowledge, in their native language.

It’s an ambitious goal, but – inspired by the success of other contributor-driven Wikimedia Foundation projects – Denny is confident that community can make it happen

We talked about:

  • his work as Head of Special Projects at the Wikimedia Foundation and his current projects: Wikifunctions and Abstract Wikipedia
  • the origin story of his first project at Wikimedia – Wikidata
  • a precursor project that informed Wikidata – Semantic MediaWiki
  • the resounding success of the Wikidata project, the most edited wiki in the world, with half a million contributors
  • how the need for more expressivity than Wikidata offers led to the idea for Abstract Wikipedia
  • an overview of the Abstract Wikipedia project
  • the abstract language-independent notation that underlies Abstract Wikipedia
  • how Abstract Wikipedia will permit almost instant updating of Wikipedia pages with the facts it provides
  • the capability of Abstract Wikipedia to permit both editing and use of knowledge in an author’s native language
  • their exploration of using LLMs to use natural language to create structured representations of knowledge
  • how the design of Abstract Wikipedia encourages and facilitates contributions to the project
  • the Wikifunctions project, a necessary precondition to Abstract Wikipedia
  • the role of Wikidata as the Rosetta Stone of the web
  • some background on the Wikifunctions project
  • the community outreach work that Wikimedia Foundation does and the role of the community in the development of Abstract Wikipedia and Wikifunctions
  • the technical foundations for his
  • how to contribute to Wikimedia Foundation projects
  • his goal to remove language barriers to allow all people to work together in a shared knowledge space
  • a reminder that Tim Berners-Lee’s original web browser included an editing function

Denny’s bio

Denny Vrandečić is Head of Special Projects at the Wikimedia Foundation, leading the development of Wikifunctions and Abstract Wikipedia. He is the founder of Wikidata, co-creator of Semantic MediaWiki, and former elected member of the Wikimedia Foundation Board of Trustees. He worked for Google on the Google Knowledge Graph. He has a PhD in Semantic Web and Knowledge Representation from the Karlsruhe Institute of Technology.

Connect with Denny online

Resources mentioned in this interview

Video

Here’s the video version of our conversation:

Podcast intro transcript

This is the Knowledge Graph Insights podcast, episode number 32. The original plan for the World Wide Web was that it would be a two-way street, with opportunities to both discover and share knowledge. That promise was lost early on – and then restored a few years later when Wikipedia added an “edit” button to the internet. Denny Vrandečić is working to make that edit function even more powerful with Abstract Wikipedia, an innovative platform that lets web citizens both create and consume the world’s knowledge, in their own language.

Interview transcript

Larry:
Hi, everyone. Welcome to episode number 32 of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show Denny Vrandecic. Denny is best known as the founder of Wikidata, which we’ll talk about more in just a minute. He’s currently the Head of Special Projects at the Wikimedia Foundation. He’s also a visiting professor at King’s College London. So welcome, Denny. Tell the folks a little bit more about what you’re up to these days.

Denny:
Thank you so much for having me, Larry. It’s really a pleasure and honor. I enjoy listening to your podcast a lot, and I’m very happy to be here too. So these days I’m with the Wikimedia Foundation and, as I said, called Head of Special Projects. There are working on two new projects, one called Wikifunctions and Abstract Wikipedia, which are really very much tied together, and we’ll get to those both in a moment, I think.

Larry:
Yeah, I’m really excited about both projects. I can’t wait to get to them, but let’s talk a little bit about Wikidata first because you started that 2012, is that correct?

Denny:
That’s right, yes.

Larry:
What was the impetus for that? What motivated you to start that project?

Denny:
Well, this goes actually back to 2005. Markus Krötzsch and I were PhD students in Karlsruhe and Wikimania was coming to Frankfurt, which is really close to Karlsruhe. It was the very first Wikimania at all. We were both Wikipedians and we wanted to go there and we thought, “What could we do?” And so we connected our research topic, which was the Semantic Web, with Wikipedia and made a proposal there. We didn’t really think it would go anywhere. We were just like, “This would be really cool if this happened.”

Denny:
But over the next few years, there was so much interest in that people actually started implementing our ideas. We picked up on that. Semantic MediaWiki came out of it. And eventually when I was finishing my PhD, I was asked by Mark Reeves, who was working for Paul Allen’s Vulcan back then, he was asking if I would like to make this happen for real. And so we approached the Wikimedia Foundation, we approached Wikimedia movement, and there was great excitement about it. We got the funding aligned and then we started working on Wikidata. This was really a dream come true basically for us who’ve been working on this idea of bringing structured data and Wikipedia together for more than seven years at that point.

Larry:
That’s so interesting because my interest, I mean obviously it goes back aways, but my history of this kind of picks up with Wikidata, so that prehistory of it, connecting Wikipedia to the Semantic Web, which is obviously you’re going to end up with something like Wikidata. And you were backed by the Vulcan Foundation or by Paul Allen’s foundation?

Denny:
Yes.

Larry:
I did not know that. I lived in Seattle for a long time, so I walked by their building a lot. Well, that’s really fascinating. So from 2005 to 2012, was it like simmering or were you doing things like precursors to the launch of Wikidata?

Denny:
We were doing precursive work. We were developing an extension called Semantic MediaWiki, which it is still quite widely used. There are two conferences per year about Semantic MediaWiki users. NASA is using it, for example, on the ISS. Microsoft and many others were also using or are still using it, which is actually integrating structured data into a MediaWiki installation and allowing everyone to build small knowledge graphs to query it and so on.

Denny:
For Wikidata, we took a lot of those lessons. We knew that we needed a little bit different data models. We started actually a different software project where we didn’t build it on Semantic MediaWiki but rather something even more structured. Semantic MediaWiki is really good if you want to interleave the text together with the structured data, with the annotations. Whereas, Wikidata really builds a pure knowledge graph, items connecting it and giving it values and so on. But originally we were thinking, “Oh, we’ll just switch on Semantic MediaWiki on the Wikipedias. I’m very glad we didn’t do that.

Denny:
Actually, just recently, the 10-year anniversary of Wikidata was coming up recently. It’s also already three years ago. Markus and Lydia Pintscher, who’s the product lead for Wikidata at Wikimedia Deutschland, and I wrote a paper about the history of Wikidata. We were actually going into detail about these topics and how Wikidata came around.

Larry:
Oh, I’ll have to link to that paper in the show notes. I’d love to read it. Well, then that’s interesting. And then so Wikidata, that was sort of the original… not original, but it was one of the first realizations of the promise of the Semantic Web, and it continues to be in the sense of the unique identifiers and entity resolution and things like that. I assume you consider it a success. It seems like it’s such an important part of the knowledge part of the internet.

Denny:
If you ask me, yes, definitely, Wikidata is a resounding success, obviously. It’s certainly a much bigger success than we expected. More than half a million people have contributed to Wikidata. If you had asked me in 2010/2011 what the number of people will be who will contribute to such a project, I would be off by more than 10X. I would never have assumed that half a million people would actually contribute to such a project. So I’m really happy. Wikidata is now the most edited wiki in the world by far, even beating English Wikipedia. It is also just very large, very comprehensive. I’m more than excited about how it has developed, and I’m very happy to see how really I was continuing to work after I left Wikidata.

Larry:
Nice. For all of its success though, you see more that could be done in this area, right? Is that where your current projects come from?

Denny:
Yes, absolutely. So Wikidata is a classical knowledge graph. Actually, we went beyond the classical data model already in Wikidata, right? So we are not just like triples, it’s not just subject predicate object, we also introduced the ability to have qualifiers on each of those statements. We introduced the ability to have references for every statement and so on. So there was a number of things that we added.

Denny:
There was already a number of things that we added on top of what knowledge graphs usually have and made it already more complicated for people to reuse it. But nevertheless, the expressivity of Wikidata remains, well, still limited. If you look at the Wikipedia article, at an encyclopedic article, there’s a lot of knowledge in there which would be very difficult to express in a system like Wikidata. My favorite examples are usually things like Jupiter is the largest planet in the solar system. This is something that you wouldn’t have explicitly in a knowledge graph. You would have the sizes of the planets in the solar system. You would know which are the planets of a solar system, but that something is the largest one was maybe implicit because the sizes would suggest that, but it wouldn’t be actually explicitly stated. But in Wikipedia, if you look for the different language additions of the Jupiter article, you will find this information in each of them explicitly. So there’s something in that knowledge which we really want to express, which we want to say, but can’t be expressed explicitly in a knowledge graph.

Denny:
Another example that I like to use is to say that Marie Curie is the only person to have ever received two Nobel Prizes in two different sciences. Again, information which can be implicitly in the knowledge graph but wouldn’t be expressed explicitly. So natural language goes well beyond what you can express in knowledge graph. This is where Abstract Wikipedia comes in, where we are aiming to actually allow the community to create the expressiveness, to capture statements like these and to make them available in the structured format.

Larry:
As you say that, “to let the community capture,” and part of this is the community writ large, like the whole planet, every language, every culture contributing to this, which everybody can do now in Wikipedia, but it’s like you said, it’s captured explicitly in each individual instance of an article or a Wikidata entry, I guess. Can you just give us a quick overview of Abstract Wikipedia and how the, I don’t know, the architecture, just the outline of it?

Denny:
So first you have to realize that Wikipedia exists in more than 300 different languages. The articles in those languages are mostly independent of each other. So if you contribute something to the article about Seattle, for example, in the English Wikipedia, there’s actually no mechanism that propagates that knowledge to the article on the German Wikipedia or the French Wikipedia or the Hausa Wikipedia, for example, since those articles are mostly done, I said, independently of each other. We’ve done a few analysis to see how does knowledge propagate between the languages, and most of it is either very slow, it can take up to half a year, for example, for a change to come to, or it happens never actually. It simply doesn’t happen.

Denny:
In Abstract Wikipedia, the goal is very similar to what Wikidata does. We already have some pieces of knowledge captured in Wikidata, for example, who is the mayor of a city, what is the population of a city and so on. This information can already be accessed by the language Wikipedias. So you will find that there are quite a few language additions that when a mayor changes of a city actually get updated because they get the information about who the mayor is from Wikidata. They just send a query to Wikidata and display the actual result.

Denny:
In Abstract Wikipedia, we want to go further for that. We basically want to express the whole article in an abstract notation, structured way so that when you edit the article, you edit the structure immediately. It’s the same thing like Wikidata, if you edit Wikidata using the Japanese interface… So you can edit Wikidata in any language in about 400 languages, I think, I would need to look it up, you can edit Wikidata in about 400 languages, but you’re always working on the same knowledge graph. So any edit that you make, say, in Japanese, will be immediately visible, say, on the Portuguese version because it’s the same knowledge graph underneath. The same thing will be true for Abstract Wikipedia. So we are going to build the article slowly, piece by piece in an abstract language-independent notation, so this is what we mean of abstract, and have this represented in a structured format. Then if your Wikipedia doesn’t have an article about a topic, we can refer to the Abstract Wikipedia version and display this as natural language text in your language and show the article.

Larry:
Oh, interesting. So you can almost instantly fill those… be there’s big gaps across languages in Wikipedia now. And so you could say something like basic facts about a city or a person. Is that the first prototypical instantiation of Abstract Wikipedia?

Denny:
Absolutely. Exactly that, that we can fill in a lot of the gaps that we currently have, particularly in small and underrepresented languages. This also allows the communities of those languages to really focus on the topics that they want to contribute to, and they can then defer to Abstract Wikipedia to have, for example, articles about every city, country, element, planet, and so on, and they don’t have to write them all themselves. They can always overwrite the article if they want to in their own language, but if they don’t want to and they want to focus on specific topics, they can just say, “Hey, we’ll take this from Abstract Wikipedia. We don’t have to worry about this anymore.” This is now being maintained by the global community that they can also contribute to, by the way.

Larry:
Yeah, interesting. How many contributors are there? Are you to the point of accepting contributions to Abstract Wikipedia, is that-

Denny:
No, we’re not there yet. That part hasn’t launched yet. We’re still working towards that.

Larry:
Yeah, I don’t know what the difference is between Wikipedia editors and Wikidata contributors. I guess you’re probably expecting some similar difference in the number of contributors. But it seems like the number of contributors could go up really quickly. I guess, what does that look like to the editors? Is it similar to how you would edit a Wikipedia article or is it a little more… I’m assuming it’s more like Wikidata, isn’t it?

Denny:
It is closer to Wikidata. It’s more like a drag-and-drop form-based editing. But we’re also thinking about using actually LLMs and let you just write a natural language and then using LLM to, for example, create a structured representation out of that.

Denny:
Since this hasn’t deployed yet, we’re currently designing and working on the user interface that will allow you to actually contribute and maintain Abstract Wikipedia. But in the end, the goal is to make it very accessible. So the number of contributors to Wikidata, as you point rightfully out, is in general higher than to most Wikipedias. There are more contributors to Wikidata than there are on most, besides maybe the biggest Wikipedia. So English Wikipedia obviously has more contributors, but German, French, Italian are all similar to Wikidata, and most of the language editions that we have actually fewer contributors than Wikidata does.

Larry:
As you say that, I’m wondering, is there an opportunity for this to explode? Meaning that, if you had a natural language interface facilitated by an LLM or something, because LLMs are so good at writing code however you’re encoding the stuff into Abstract Wikipedia, I mean are you… I get immediately excited about that. Should I calm down or am I reasonably excited about this?

Denny:
I would recommend that you are reasonably excited about that. I think it’s very fair to be, because it’ll basically allow us to have people contribute across language barriers on the same article, which is a huge thing. We have a lot of research that shows that if people from different points of view contribute together on a Wikipedia article, the article becomes more comprehensive, more neutral, and better on basically every metric. If we take away the language barrier and allow people to work on the same text across languages, we expect the same benefit here. So we can actually hope for more neutral and more comprehensive articles in Wikipedia using this approach.

Larry:
Interesting. I’ve read a lot of business research and other research about the benefits of diverse teams and just the more points of view, the better. And this sounds like a way to really accelerate that in your data and knowledge ecosystem there. That’s really exciting. Where are we at with Abstract Wikipedia? Is this something that people can go play with now? How can people learn more about it, I guess?

Denny:
Well, to learn more about it, there is plenty of websites that we’ll put in the show notes on the Wikipedia pages. What we currently have is Wikifunctions, which is a necessary precondition to Abstract Wikipedia. There the community can create functions that, for example… There can be all kinds of functions, but what is relevant for Abstract Wikipedia is that they can have functions which write text. So given the structured data that we are building in Abstract Wikipedia, how do we actually turn that into a natural language text that the community can read, again that normal users can read?

Denny:
Because we don’t expect anyone to read structured data and get information out of it. What people want to read is text. This is where Wikifunctions come in, that they take the structured representations that the community is creating and maintaining in Abstract Wikipedia and turn it into text so at the end it looks just like normal Wikipedia article in their language, but it’ll be more comprehensive, it’s more up to date, it’ll be more correct and so on, because the larger community is actually working on the topic, maybe not even speaking that language and contributing to the knowledge that you have now available in your language.

Larry:
I heard the term Rosetta Stone come up someplace in my research on this. Is that what you’re talking about there?

Denny:
Kind of. The Rosetta Stone was used for reading and for deciphering languages, but maybe, yeah. I like to think of Wikidata particularly as the Rosetta Stone of the web because it connects to literally thousands of different databases, and they identify this for the different topics, which is exactly if you… Since we were talking about the Semantic Web earlier, in the Semantic Web we have all these identifiers for the different topics and so on, and Wikidata really pulls those together and allows you to understand what this database will be mapped to what in that database over there, so that you can actually connect the data from all over the web together.

Larry:
Nice. Hey, I want to go back to Wikifunctions a little bit because you mentioned that it’s a prerequisite to Abstract Wikipedia, which makes sense. But Wikifunctions already existse. There’s, what, a couple thousand of them out there? Can you talk a little bit about how that project arose and what it’s looking like right now?

Denny:
Yes. So currently we have still a starting budding community of a few dozen contributors who are working on Wikifunctions. They have created already a catalog of more than 2,000 functions and made them available. What we are working on is introducing more data types for those functions, more access to Wikidata, particularly to the lexicographic data in Wikidata. So what we have in Wikidata is not only just the ontological space, the data about topics and things and items, but also we have a lot of data about words, what these words mean, connecting back to the ontology in Wikidata, and their inflections and so on. So if this lexicographical data being available for Wikifunctions, you can actually connect all of this together and start creating correct sentences that explain basically the topics of the different Wikipedia articles. So really, Abstract Wikipedia is a place where all these projects come together in order to provide more knowledge in more languages and allow everyone to contribute to this common knowledge that we have in Wikimedia.

Larry:
That is so interesting. This is the original promise of the Semantic Web in some ways. It seems like that anybody in the world can access and contribute to it in that sense. That’s really exciting. How do you build this? Wikipedia is obviously many years old and it’s well-established, and Wikidata as well, so you have the historical legacy there. But do you have a proactive outreach to try to get more contributors to Wikifunctions and people to test Abstract Wikipedia? Can you tell me a little bit about the development of these new initiatives?

Denny:
Yes. Yes, we have a whole team that’s dedicated to community outreach at the Wikimedia Foundation that is about communicating with the community. First, I can tell you that a lot of members of the community are very excited about what we’re doing, and they like to contribute to our work. They help us building a system. When we have difficult technical questions, we sometimes go to the community. We get a really great sounding board and feedback. We are just starting these days about a very difficult, some might call it an implementation detail, but a very difficult question of where the content of Abstract Wikipedia is actually supposed to live. Because as I said, it’s connecting Wikidata, Wikifunctions, Wikipedia, but where will it be located, where will it be stored and so on? Which is something we actually don’t know what the best answer is.

Denny:
And so we go to the community and discuss this with them and show the ideas that we have, hope for more ideas than that they might have. A number of times we had actually contributors from the community coming to us with better ideas than anything we came up with and we went very happily with those solutions. So there are a lot of people who are very excited about what we’re building. On the other side, we’re also reaching out to the communities in order to see how they can use Abstract Wikipedia in their Wikipedia, how they can use Wikifunctions in their Wikipedia, and how it can make their life better, how it can make their projects better, and how they can use it there. So there are a number of interlocking approaches here that we’re following.

Larry:
Yeah, that’s probably a whole other podcast about the infrastructure that underlies all of this. Because I’ve been doing content modeling for many years, and so I’m always aware that somewhere under that content management system or whatever there’s a repository where everything actually lies, something close to the physical storage of the data. So Abstract Wikipedia will provide for a way for anybody in any language to edit and contribute to this, and then it goes … someplace, I guess. And that’s still TBD, is that sort of the idea?

Denny:
It’s all TBD. That’s exactly what we want to decide in the upcoming few weeks together with the community, because there’s also a lot of governance questions around that. Because depending on where exactly it’s stored, there’s also the community that actually has to maintain this content and keep it up to date and create the rules about creating this content. Do we want this to be close to Wikifunctions, which is a very technical community? Do we want this to be closer to Wikidata? Do we want a whole new project dedicated to this? I really don’t know what the answer is right now.

Larry:
Interesting. So I guess just technically, would all of that data that people contribute to Abstract Wikipedia, is it all just triples? I guess, tell me a little bit about that because I guess that’s going to lead to… It sounds like you’re counting on the community to help you solve this problem of where do we put this.

Denny:
I absolutely am. I mean, it can all be turned into triples, just like with Wikidata. Wikidata isn’t stored natively as triples. It’s stored natively as JSON files, so basically as tree structures and maps. The same thing is true for Wikifunctions, everything there is stored as JSON files. We will do the same thing for Abstract Wikipedia. Obviously you can turn all of that into triples for Wikidata. We do that and we store it into a very large triple store where we then allow SPARQL queries over it and so on. We could do the same thing for Wikifunctions and Abstract Wikipedia. But natively it’s all stored as individual JSON files in a MariaDB that is then accessed by MediaWiki.

Larry:
Okay, yeah. And again, it’s early days. I can’t wait to see how it evolves. We’re almost out of time, but how can people get involved if they want to contribute to or get involved with any of these projects? What’s the best way for folks to…

Denny:
So one way to join us is just to come to wikifunctions.org, the newest Wikimedia project, and there they can contribute to functions, they can see already the existing catalog of functions, they can contribute to…. Another way is to contribute for Wikidata, because more metadata, especially more lexicographic data on Wikidata, is super useful in order to make this project happen. Then we have a website on Meta-Wiki, which actually describes the project as a whole, to give an overview of the project and the different places. All our software is open source. We really welcome people to come and join us also in developing the software. So there are numerous places where people can contribute to, and I’m very confident we’ll put all of these in the show notes.

Larry:
Absolutely, that’ll all get in there. But hey, before we wrap up, Denny, I want to make sure, is there anything last, anything you want to revisit from the conversation or that you just want to make sure we share before we wrap up?

Denny:
I really like to share that this is a project by and for the communities to allow them to reach that goal of the Wikimedia movement. Our division of the Wikimedia movement is really to imagine a world where everyone can share in the sum of all knowledge. And with Abstract Wikipedia, we are trying to unlock this everyone part so that you can share no matter what your language background is and work together on a common shared knowledge base that everyone else can also benefit from and that everyone else can maintain together with you, so that we can tear down those language barriers and allow people to work together on the shared knowledge space.

Denny:
Really, the important thing here for me is to understand that Wikipedia and all the projects of the Wikimedia movement are not just projects for consumption. They are also projects that allow everyone to join us, to collaborate with us, to work together. Wikipedia’s knowledge is not whatever the big industry companies hand us down. It’s all of us together working on it and each of us being able to contribute to it and to contribute our point of view, our knowledge and everything. So this is really the thing that for me really is important. I hope that it’s something that we won’t lose with technology, which only allows to consume.

Larry:
That’s beautiful. I don’t know if we talked about this before we arranged this, but my whole intent in this podcast is to democratize all of this stuff, like how you do this, what’s going on, how to contribute. And this seems like a perfect place for people who want to make the web a better place, so thanks for putting this together.

Denny:
Tim Berners-Lee, since he created the web, was always about this two-way street, right? In the original web browser that he wrote, there was an added feature in it that you could create. And then with Mosaic and Netscape, we lost this ability actually, and it took a few years until what Cunningham came up with, the Wiki website, we did have, again, an Edit button and so on. But really, the history of the web seems to be driven by idealists who really want to point to this everyone can contribute space and by forces that just want to turn it into a place of consumption on the other side. We didn’t talk about this before, but I’m really happy that your goal is also to watch the democratization of the way that we can share knowledge and share other things over the web.

Larry:
Great. We got to check in a year or so to see how everything’s going on this. But hey, one very last thing, Denny, if folks want to connect with you or follow you online, what’s the best place to find you?

Denny:
Oh, the best way, that’s difficult. So denny@wikimedia.org is my email address, but I’m really bad at answering emails, so I’m not sure this is the easiest way to reach out. Otherwise, you can find me on Mastodon, you can find me on LinkedIn and some other of those places. A very good place is actually just to find me on the Wikimedia projects. I’m user Denny, and if you reach out on my top pages, that one I’ll see and usually respond to it.

Larry:
I’ll emphasize the Wikimedia one because we want to get people involved, so-

Denny:
Yes.

Larry:
… I’ll focus on that. Well, thank you so much, Denny. This was an awesome conversation.

Denny:
Thank you, Larry, for having me. This was really fun.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top