Dean Allemang: Semantic Web for the Working Ontologist – Episode 6

photo of Dean Allemang, author of the book Semantic Web for the Working Ontologist
Dean Allemang

Dean Allemang literally wrote the book on the semantic web. “Semantic Web for the Working Ontologist” is now in its third edition.

In the book, Dean and his co-authors, James Hendler and Fabien Gandon, show how to apply web standards to build a meaningful web of global, connected knowledge.

More recently, Dean has conducted research with his colleagues at data.world that shows how using knowledge graphs can triple the accuracy of LLM-based question-answering systems.

We talked about:

  • his role as a principal solutions architect at data.world
  • the meaning of the “semantic web” and its intent of sharing meaning across the web
  • the long history of knowledge representation and how the connectedness of the semantic web adds to it
  • the crucial difference between documents about things and the strings that describe them
  • the contrast between the persistent nature of enterprise data and the ephemerality of the applications that use the data
  • the power of the simple structure of RDF, its mathematical affordances, and the ease of distribution it permits
  • the impact of newer AI tech on knowledge graph building and querying
  • the research that he and Juan Sequeda have conducted that shows how using knowledge graphs can triple the accuracy of LLM-based question-answering systems
  • his thoughts on the yet-to-be-resolved one-way or two-way ontology question
  • the crucial role of trust in AI and how replacing LLMs with knowledge graphs as the point of contact in AI systems could build more trust

Dean’s bio

Dean Allemang has been active in the field of Artificial Intelligence (AI) since the 1980s. With a notable emphasis on Semantic Web, he is the author of the book “Semantic Web for the Working Ontologist.” His passion for understanding and implementing knowledge graphs led to a significant publication about using LLMs to answer queries over structured data, which introduced a new benchmark for evaluation.

In his current role as a Principal Solutions Architect at data.world, he contributes extensively to the development of the AI Context Engine product, which is inspired by his recent research (with Juan Sequeda and Bryon Jacob), and underscores his commitment to practical application of theoretical principles.

For a span of about a decode, Dean operated as an independent consultant, utilizing knowledge graph solutions to address challenges in industries such as Media, Finance, and Life Sciences. This diverse experience has cultivated a broad perspective on applying AI and Semantic Web principles.

Influenced by Sir Tim Berners-Lee’s concept of linked data and data sharing, Dean Allemang’s work reflects a consistent focus on these principles. His contributions have advanced the field of AI and his current interest lies in how knowledge graphs can make generative AI more effective.

Connect with Dean online

Resources mentioned in this interview

Video

Here’s the video version of our conversation:

Podcast intro transcript

This is the Knowledge Graph Insights podcast, episode number 6. Long before the introduction of the semantic web – the innovation that added meaning and metadata to documents on the web – AI pioneers like Dean Allemang had been thinking about how knowledge could be formalized to help people do their work. The web itself, along with the W3C standards that power its semantic capabilities, gave Dean and his peers the ability to scale and connect existing practices and technologies to build a more meaningful web.

Interview transcript

Larry:
Things. Hi everyone. Welcome to episode number six of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show, Dean Allemang. Dean is a principal solutions architect at data.world, a company in this space. He’s also the author of The Semantic Web for the Working Ontologist, kind of the original operating manual and textbook for this field. So welcome, Dean. Tell the folks a little bit more about what you’re up to these days.

Dean:
Hi, Larry. It’s great to be here. So I’ve been at data.world for about three years now. I was an independent consultant for a bit before that. What I have been up to lately, not surprisingly, is figuring out how all of the new AI fits in with the sort of knowledge-intensive stuff that we’ve been doing with ontologies and knowledge graphs and things like that.

Larry:
Cool. Yeah, and I’d love to talk more about that research, time permitting, but I’ll certainly link to it. I know you have a lot of cool new stuff coming up too, so I’ll be sure to keep it. I’ll try to keep the web page updated too with that.

Larry:
But hey, I want to back up just a little bit, like the dawn of this technology and this whole ecosystem we’re operating in the semantic web. That’s the first two words in your title of your book. Tell us folks a little bit more about the semantic web, its origins, what it is.

Dean:
Yeah, so one of the things I like to say about the semantic web is that the emphasis is on the final syllable; it’s about web. That’s the important insight that the semantic web brings to. Well, at the time, knowledge representation was the big thing, and the real key to the semantic web is in fact the web nature of it. So what you’re doing in the semantic web, and this is a vision that came from Tim Berners-Lee in the mid, that sort of got more popular in the late 90s when the standards started to come through. But the idea is that we have this notion of the web; cast your mind back to the nineties when the web was a new thing, and instead of just having pages linking off to each other, could you actually have bits of knowledge that refer to each other in a deep, meaningful, dare I say it, semantic way.

Dean:
So Tim Berners-Lee came up with the name semantic web, and the point of it was that we want to be sharing not just documents on the web but meaning on the web. And that was the whole idea from Tim Berners-Lee, basically in the 90s. In some sense, he often says that this is the web he always had in mind, but the document web that we know and love was the first, the crawl of the crawl walk, run of the semantic web. That’s how Tim often refers to it. But for us, this is really looking at knowledge representation, which has been around for a long time, and bringing it forward to the new information age, which is what we now call the web.

Larry:
Yeah, and one thing I want to point out is that one of the co-authors of your book is James Hendler, who famously co-wrote the paper with Tim Berners-Lee about introducing the idea that of…

Dean:
That’s right. The Scientific American article back in 2001, I think it was?

Larry:
I think it was March or May; it was one of those M months in 2001. Yes.

Dean:
Yeah, early in ’01. And that really sort of put the name on the idea of the semantic web. And so that was sort of how it all began, and the point of my book, which, as Jim and I wrote the first edition, I guess about seven years later, we were doing a course, a little corporate training with our partners, TopQuadrant, at the time, about semantic web. And we found that after four days of intense lectures and exercises and things, that we still found that a lot of very smart people were pretty tentative about answering basic questions about the semantic web. Well, we really need something to give them at the end of this course that they can take home and read. And so, one day over a beer, after doing this course, Jim and I conceived the idea of semantic web for the Working Ontologist and started to work on it. And of course, as I know you’re aware, Larry, projects like this always take much longer than you expect. And I think it was actually three years later that we finally had the manuscript ready for the publisher.

Larry:
I don’t know what you’re talking about. When I was a book editor, everything that always came in early and no… I know how it goes, but hey, I want to talk… I love the origin story because it comes right out of what you want this book to be doing. It’s like, “Okay, these people who are learning…” Because there’s new technology, well, there’s kind of two things. I guess you mentioned that knowledge representation has been around as a discipline for a while and then, but I’m going to guess it wasn’t as technical as it is now or the specific technical implementation of it changed with the advent of the semantic web.

Dean:
Yes, that’s certainly the case. Well, back in the really old days, things like what KL-ONE was a lisp-based knowledge, representation, language, and loom and all these things. They were actually very technical, indeed. What they weren’t was distributed. That’s why I say the emphasis on the final word web, and this is the thing that actually, if there’s one thing about the semantic web that I find doesn’t get through is that what we’re doing here is sharing knowledge. So if you think about the EDM Council, one of my former clients, they published an ontology called FIBO, the Financial Industry Business Ontology. What are they doing there? They’re writing down a data model. Well, anybody could write a data model, a lots of people do, and they are publishing it, and people do that as well. The OMG does a lot of that stuff. But why did the EDM Council decide to use the semantic web?

Dean:
They want people to be able to refer to parts of that independent of some document. They want to have a machine-readable way of bringing this into their system so that every last part of this great, big behemoth can be referenced on its own. And this is what documents don’t do; this is what the document web didn’t do. You could talk about a document; you couldn’t talk about the things. So, nowadays, we talk about things versus strings, which was actually something the database people are talking about, I think, as early as the 80s. But things versus strings on the web really has a lot more meaning to it; you can have a document with a bunch of words in it, but you have the concept the document is about, and you want to be able to have lots of different things. Talk about that one, that one concept.

Dean:
So this is the structure that the semantic web put together, and the whole point is that we want to be able to publish what we know about a topic. So not building applications. Yes, you can build applications with the semantic web, but that’s not what is designed for. Say, “Wow, why is this designed in this funny way?” Distributed systems are hard; the web is a distributed system, and you think, “Well, I don’t want namespaces; I’m just going to just write my program; my program will be done.” Yeah, you can write your program; we’ve been doing that for ages, and my colleague Dave McComb here, he talks about the data-centric enterprise and this sort of thing, that you need to have your enterprise go from thinking about its applications to thinking about data. Data is the thing that actually persists through an enterprise’s lifetime. Applications come and go, the data continues to be important.

Dean:
So, what you really need to do is to figure out how to share the data with today’s colleagues or colleagues in the future, who might be you. Future Dean needs to know something. How do I communicate with future Dean or future Larry? So this is what semantic web is all about; it’s about communicating across communities or across time so that data becomes the thing you’re talking about, or metadata, in the case of ontologies, and that’s what it’s all about. And so you say, “Gee, I could write my program more easily doing something else, using Postgres, using Python.” Use those things, sure. But when you want to write something down and communicate it, you’re going to need a language for that.

Larry:
Yep. Well, let’s talk a little bit about that because the foundation of this is the RDF format, which is really simple. And I think one of the interesting things to me about that as it requires just a little bit; it’s kind of like both a convenient and comfortable way to think about stuff. Just nodes, and you’re connected by little lines, which is the way people diagram a lot of stuff. But we’re all used to, as soon as we think about data, we think about tables, I guess. But can you talk about the RDF format and how that enables the semantic web?

Dean:
Yeah, yeah. So, actually, I was reminded a couple of years ago. I was talking to a colleague who’s been doing semantic web for quite a while, so I’m not going to name them because I’m going to be saying something kind of embarrassing about them. And they had a colleague with them said, “This RDF stuff is so hard, we’re going to have to figure out how to do the semantic web with something simpler than RDF.” I said, “Okay, wait a minute. You’re in a distributed system; you have a thing that one person’s talking about, a thing that someone else is talking about, and you want to say something about how they relate.” Dean’s thing relates to Larry’s thing by being more specific than. You need to talk about Dean’s thing, Larry’s thing, and the relationship. That’s three things you need, and you need these to be global references. Fortunately, the web gives us a way to talk about global references.

Dean:
There you’ve got it. IRI on the web, IRI for the thing, Dean’s thing, IRI on the web for Larry’s thing, IRI on the web for the relationship. You need those three things. You can’t get by without those. How are you going to… And that is RDF; that’s the beginning of RDF, all the content of RDF, and the end of RDF; that’s all it is. Tell me how are you going to do something simpler than that that satisfies by talk about a global thing and relate it to another global thing. That’s the basic requirement. How are you going to get simpler than that? And the person who was naysaying was dumbfounded and silent; the person who’d been doing semantic web for years disappointed me by saying, “Wow, I never thought of it that way before.” For crying out loud if you don’t think about it. So when people say, “I want to know about the semantic web,” this is what you’re doing.

Dean:
You’re connecting one thing to another; both of them are on the web, and you want your connection to be on the web. That’s called a triple. And yeah, we call it subject-predicate objects. It’s in English. That’s how we build sentences. So, Dean’s notion of customer and Larry’s notion of customer, my customer is more specific than your customer. And this is an example I like to do for a lot of enterprises. People say, “Gee, I don’t know what you mean here.” Well, and I always ask them, “When you say the word customer, do you mean the same thing as everyone else in your company?” And the answer is always no. Someone in my company has another meaning of the word customer. And if you talk to them and all you do is you say customer, customer, and you’re not aware of this, oh my God, pandemonium ensues, right?

Dean:
Supposing all we do is give you the ability to say my notion of customer and your notion of customer, and always be specific about which one you’re talking about. Oh my God, so many discussions would be so much easier. We don’t do this in everyday parlance because you don’t speak that way, but when we talk about machine-readable metadata, we can do this. And that’s what the web does, and this is not some kooky thing the semantic web’s inventing; that’s what the web does. The web gives me a way to talk about Dean’s things and Larry’s things, and all we’re doing is taking that web thing and saying, “Dean’s thing relates to Larry’s thing.” It’s as simple as that. It actually can’t be made any simpler. Some people often say, “RDF is the mathematically simplest way to talk about data.”

Dean:
And what they really mean there is that that triple is a cell in a table. If you want to talk about tables, or it’s an element in an XML document or an object; actually, it’s a line in an object, in a JSON document, and all these different ways of doing data, the triple is the tiniest piece there, and it’s a way that you can talk about it mathematically. And it’s also easy to distribute that each of these things is an IRI. It’s actually dead simple. And I think one of the things with RDF is it’s so simple that you can’t believe that it’s worthwhile specifying it. I mean, I’ve had people look at RDF and say, “This is nothing; this specifies nothing. There’s nothing to it.” Well, actually what it specifies is a distributed way to say, “You’re right.” It’s very, very little, the person who said that almost, right? RDF says very, very, very little. It’s not nothing; it’s very little.

Dean:
But the thing is that very little thing can be distributed across the web, and from that little thing, you have the cell on the table or the element in XML or whatnot, you can build up all of these other structures. And because it’s so mathematically simple, yeah, you can build XML out of triples; you can build tables out of triples. You can build JSON documents out of triples and fill in anything else you want there. You can build DDLs out of triples; whatever else you want to fit in there, it is so mathematically simple, you can build all sorts of stuff out of it.

Larry:
And that speaks to its power, the fact that is so simple and implementable across any number of encoding formats. But it seems like the real power of it is when it’s handled in a way that you can do the graph-ey stuff, that the languages above RDF. Once RDF was established, there’s all this other stuff that the OWL and… Actually …

Dean:
RDF-S, OWL, SKOS.

Larry:
RDF-S and OWL and SKOS. Those are kind of the first three, or were they?

Dean:
Certainly, RDF-S is so close to RDF that in the way the standards documents are written, they’re in the same document, which I think is actually kind of confusing, and it really shouldn’t be that way, but it’s a document thing. They’re actually in different namespaces on the web. And then OWL came along a little bit later, and then SKOS came along a little bit later than that.

Larry:
And SKOS, I think for a lot of people. I’ve talked to many people who say that many if not most ontology projects begin with a taxonomy. And that’s what SKOS helps you manage, but the thing about all of those languages you just mentioned, they’re all based on the RDF format. So once you’ve got these triples, you can do a lot of different things with them. You can categorize them in a taxonomy or classify them according to concepts in an OWL domain. I guess that’s…

Dean:
Or looking at the other way, if you build a taxonomy, say in SKOS, and you want to distribute it for governance reasons, you’re governing part of it, I’m governing part of it, we want to relate these together. RDF gives you the foundation for doing that. You don’t have to build a distributed vocabulary management language. You just build a vocabulary management language and base it on RDF, and it’s automatically distributed. And just distribution can either be in publication or, usually for that sort of thing, governance. I’m responsible for some things, you’re responsible for some things, and it’s clear to see who’s responsible for what. And when I change them, you can see what I’ve changed and vice versa.

Larry:
Yeah, that’s it. And that notion of different ways it can manifest, that seems really powerful as well, because I tend to think of it. Well, I guess, let me say a little side thing here, but I think the way a lot of people know about RDF is in the linked open data, the linked data in web documents that’s there for SEO purposes. I think a lot of people have backed into this, the awareness of RDF and that through that. But that gets at what you’re talking about. The real power of this is that it’s connectable and connected on the web. And to what you just said, though, that there’s different ways that you can do that connecting.

Dean:
Yeah. So the thing you’re talking about now, just a backbone of that, is this thing called schema.org, which is sort of a RDF-based endeavor that wasn’t part of the RDF world, which there was a lot of personal political tension there, but actually, if you look at schema.org, the fact that the folks who built schema.org disagreed with some of the folks who built RDF-S on a number of things, but nevertheless, they were able to build their schema.org structures. And the really cool thing about schema.org is, as an organization, they have some guidelines about how to expand that thing, and you’ve got various processes. If I want to build a schema.org extension about finance, I can go through a do-stop, but because it’s an RDF, I actually don’t need to talk to them at all. I can publish my own RDF or schema.org extension for finance, reference their stuff.

Dean:
It’ll be in my own namespace, and then I could promote that as much as I want and say, “Hey, you have me,” and by me, it would make more sense, would be the EDM Council, “Can do this, and yeah, this is on our authority, not on the schema.org authority.” And not to downplay schema.org skills, honestly, if we’re talking about enterprise data management in finance, what does EDM Council stand for? Enterprise Data Management. I would trust EDM Council more than schema.org, and that’s not disparaging schema.org at all. They themselves will say, “We aren’t experts in finance; we aren’t experts in manufacturing; we aren’t experts in biology. That’s not our job. We are making a backbone of everyday stuff, web things like SEO.” Like you mentioned, if you’re an expert in finance, medicine, biology, manufacturing, whatever, knock yourself out, build a schema.org chapter for that that you promote, that you have in your namespace.

Dean:
And this is Tim Berners-Lee’s vision of the semantic web. Just like the web, build your own webpage. That own webpage is your own thing that you’re saying. Similarly, build your own piece of schema.org, your own ontology, your own vocabulary, your own data. You do that stuff, and you own it, but it all connects through with everything else. Now, nowadays, Tim is doing stuff with this thing called Solid, which I’m not going to go into details of because I’m hardly an expert in it, where you try to put in a lot of the privacy and ownership stuff that was missing in the first vision of semantic web.

Larry:
Yeah, that was his third thing. But I want to go back to what you’re just saying about schema.org and FIBO: that those are both. I don’t know if you’d call schema.org an upper ontology, but it’s an attempt to organize things and permit that. But schema.org is kind of like this, just this broad, general purpose kind of thing to help identify various things on the web.

Dean:
And that was my design. Yeah.

Larry:
Yeah, exactly. Whereas FIBO is very specifically about financial industry stuff, but I guess that kind of gets into any ontology project that I’ve heard of is constrained typically by a domain that you get clear on where you’re operating, and then that leads you to do something like schema.org or FIBO or whatever ontology you’re developing.

Dean:
Well, in fact, my colleague Elisa Kendall, along with Deb McGuinness, wrote a book, which I actually do have somewhere nearby, but I don’t quite…

Larry:
I’ve got it right here. Yeah.

Dean:
You got it. Okay.

Larry:
Yeah.

Dean:
Wrote a book called Ontology Engineering, and I forget if it has a subtitle or not. And they talk a lot about methodology, and one of the things that they talk about is indeed scoping, and they actually use this notion of competency questions, which goes all the way back to the seventies or eighties or something like that. And it is basically what you’re saying there, you scope into some domain. I kind of wonder if the new AI might change some of our opinions about ontology methodology, but that’s sort of yet to be seen.

Larry:
That’s an interesting point. I do want to make sure we get to that, because I remember two years ago at the Knowledge Graph Conference going there thinking, “Oh, this is going to be great, Dean, and everybody’s going to tell me how knowledge graphs can fix these stupid LLMs.” But you were much more excited about LLMs as tools for helping to build knowledge graphs and write queries and stuff. Can you talk a little bit about how you perceive how AI is going to impact, or specifically, I guess, generative AI and the transformer technology is going to impact the knowledge graph world?

Dean:
So the impacts are obviously a two-way street, and I guess that’s what you were just alluding to. How can the knowledge graph world impact the generative AI world and vice versa? So let’s talk about the vice versa first. So there’s some work that I’ve been doing with my colleague Juan Sequeda, and this is something that I think you’re going to link in the notes here, but the idea of that is if you’ve got a whole bunch of data out there and you’re building a knowledge graph for it, or you think you might want to, the value proposition has changed because of generative AI. So the old value proposition was you got a bunch of data out there, and it’s going to be in tables because we live in a relational world, and you have a whole team of folks who are writing queries against this. You have a question, you talk to that team, they write a query, and you get an answer back.

Dean:
That’s the world that we’ve been in for a long time. Now, the value proposition of a knowledge graph was you could build a knowledge structure over this, a thing we call an ontology, that tells you the “meaning semantics of your data.” And now, instead of writing a query against all this base technical data, you write a query in business terms on your ontology, and this is a much simpler process and a much more auditable process. An example that I gave for one of my customers about the same time, you’re talking about two years ago, at the beginning of a POC, said, “At the beginning of this POC, I was dating knowledge graphs. And when I saw this query, I fell in love with knowledge graphs.” And here’s an isomorphic version. I won’t tell you my client’s industry because he’s kind of private, but supposing you have a little ontology that says, “A computer is made of a motherboard, memory card, and a disk drive,” a very simple computer model.

Dean:
And you have a business rule that says, “The cost, to me, of this whole thing is the sum of the cost to the parts,” a very simple business rule as well. And you put that in relational databases, and you’ve got a foreign key to your motherboard, a foreign key to your main memory, a foreign key to your disk drive. And each of them has a column for the cost, which might have different names. You write in an elaborate query that says, “Join this to this, and join this to this, and take the cost from here; take the cost from here; take the cost from here; add them all up. And there’s the price of your computer.” And that query is half a page long. Supposing instead, you have a knowledge graph that says, “Things have prices, your computer has parts,” and the part might be a motherboard to all sorts of things.

Dean:
“And now map that to your data.” Okay, we’re going to map one of these things to the motherboard, one of these things to the memory, and so on. And now your query says, “X has a part; parts have prices; sum price group by X.” Boom. Very simple. And that actually sounds a lot like the business rule. The price of the computer is the sum of the price of its parts; that’s almost verbatim SPARQL. So the value proposition used to be you can write queries in a much more business-oriented way, and it’s also more maintainable; if you bring in more parts, you just map the parts, and that query never changes. Wonderful value proposition. That value proposition did not make knowledge graphs take off because, lots of reasons. “Hey, I’ve got a team full of query writers. I am paying them, and I don’t want to fire them, and I have personal relationships with them. I think I want to stick with this.”

Dean:
That’s actually a perfectly valid, it makes sense, it sounds weird, but it’s actually a perfectly valid reason. Now with LLMs in the world, generative AI, it turns out the LLMs speak SPARQL and graph and things very well. Now you do all that modeling, and instead of saying, “I can write a query more easily,” now you say, “I can have an LLM write a query more easily.” And so, now I’m claiming that an LLM will write better queries over this structure than they will over the relational database. How on earth do I know that? I don’t know that. Well, I do now, because that’s the research work that Juan and I did. We actually did a neck-to-neck bake-off, and using the same data structure that was actually published by the OMG, put a knowledge graph over it and measure them, and there’s a whole bunch of research that you can read in the paper there.

Dean:
The bottom line is it works three times better, and that’s pretty cool. Now the thing about this is we published all of our data, all of our knowledge graphs, all of our stuff. And now we’ve actually had other people start to reproduce this by using our data, doing their own experiments, using their own notion of whatever fills in for knowledge gravity. We use the whole semantic web stack, OWL, RDF-S, and so on, that you and I have been talking about. Jesús Barrasa from Neo; he used the whole Neo stack. Now, he didn’t reproduce the whole experiment; he just did a little part of it, but the reaction that he got was pretty much consistent with the results that we had; he hasn’t been as scientific about it as we have, yet.

Dean:
It takes some time to build and run an experiment. This is not to disparage him at all, but he did that first work to show, “Yep, this works as well here.” And that’s the real point of this work that Juan and I did. We want to do some real science here, put our data out there like real scientists do, and encourage other people to repeat the experiment like real scientists do. And we were delighted when other people have done this and Jesús being probably the highest profile one, coming from a company like Neo.

Larry:
Yep. Hey, as you just said that you reminded me when I’m thinking, I was reflecting on your research, and you and Juan both have PhDs. You’re doing proper research there, and I know a lot of people in this world have PhDs. Do you need a PhD to be an ontologist?

Dean:
No. Well, certainly not nowadays. And because once again, and this is going to go in the other direction, the LLMs are making it really a lot easier to do this sort of thing. So one of the things that data.world is doing, and at this point I’m now talking about data.world’s future plan. So I need to be very cagey here, but I think I can speak fairly clearly that we actually believe that LLMs are going to help in this process a lot. And so some of the work that I’m… Actually, I was just chatting with one of my colleagues about this morning is along these lines. I’m not going to go into any detail there, but a lot of the research that we have going in the future is going along those lines.

Dean:
One thing I think I can talk about, because I’ve been thinking about this for ages, and I finally now have a reason to spend the time. In ontology design, things like in Deb and Elisa’s book, there’s a question that I find that a lot of ontologists will disagree about. And that is, “If I am modeling a relationship, say customers place orders, and then the other way around, an order is placed by a customer, do I need to model both of those things or should I model just one?” And if you come from a query perspective, you say, “Well, model just one because you could always in your query just swap it around and write it the other way.” And so your query language can pick up any slack that second one would use, and the second one is just going to confuse your model, add more stuff to it, and if you’re materializing things, you’re now materializing twice as much stuff, and you have all sorts of technical issues.

Dean:
From a modeling perspective, you say, “Well, actually, if I want to say all the orders placed by a customer, well, that’s a concept that’s interesting. I need to have placed by to define that concept in a logical way. And if I were to say all the customers who have placed an order, well maybe not an order, but a product, I need to go the other way around.”

Dean:
So, in order to talk about modeling concepts like customers who have ordered the same product, that’s called a market, or products ordered by a customer, that’s called an inventory. These are real concepts in business, and I want to define them, and I don’t want to have my definitions have to talk in weird, obtuse ways going the wrong way around; model them both. And I think Elisa, if I remember correctly, actually recommends always modeling both. Just don’t wait until you need it; just model them both. So these are very different viewpoints for very different reasons. Now, enter generative AI. Generative AI can read an ontology and do cool stuff with it. Will generative AI respond better to the one-way ontology or the two-way ontology?

Dean:
I don’t know the answer to that, but that’s an experimental question. And so, this is one that I actually proposed way back in March. I thought if I have some spare time, I’ll do this. But now at data.world, we’re finding our customers are starting to build models, and they’re starting to ask this question, and we’re starting to actually run it through generative AI. And it’s like, “Well, we actually don’t know what advice to give you.” And so it turns out that these very basic research questions are now things that turn into advice. And this finally, getting back to your question, no, our customers are actually out there building ontologies on their own, often with no particular training in computer science at all. There’s a lot of fussy computer science stuff, but actually, if you can do data modeling for SQL, then you can do ontology building as well.

Dean:
And one of the things that we may be working on are some tools to help you do that because a lot of fussy stuff that the machine can or at least ought to be able to do, but the conceptual part is basically the same sort of thing that you might expect from business process modeling, data modeling, business architecture modeling, those skills. Which we don’t normally think of as being PhD-requiring skills. Business architecture is actually the one that I think is the most similar to ontology modeling. If I were to pick up somebody from a different discipline and say, “I want to teach them ontology modeling,” people often ask me, “What skill set do you want people to have?” Say, “Well, it’s basically the skill set that business architects have.”

Dean:
You need to be able to think about business in an analytic way as opposed to data architects or enterprise architects. A business architect thinks about the business in an analytic, modular, recombinable way. They aren’t necessarily data modelers, but they understand the constraints that data modelers have. Business architects are pretty interesting people, and I have great respect for them, and I think they’re the ones who are sitting at pretty much the same spot where ontologists sit.

Larry:
Interesting. Yeah. Well, hopefully there’s going to be this huge recruiting problem. And they’ll be scrambling to find ontologists anywhere. But…

Dean:
Yeah, but if you find somebody who’s a good business architect who for some reason wants to re-career, that’s where I would recommend you find your points.

Larry:
That’s where you go. Cool. Hey, Dean, I can’t believe it. We’re coming up close to time already, but before we wrap up, yeah, is there anything you want to revisit from the conversation or just make sure we share?

Dean:
So let’s see. I just want to make sure we’ve gone the two ways around for generative AI and ontologies, and we’ve only gone sort of the one way I thought generative AI can help us build ontologies, help us review ontologies, help us find a way around our cataloged data, and that’s all pretty cool. But let’s go the other way around, if you think about the famous criticisms that people have for generative AI or LLMs, in particular, the language models. And that is that they hallucinate, they make stuff up, and when they do so, or whether they do so, they can’t explain themselves in a reasonable way that gives you confidence in the answers. And that’s what it comes down to. Tim Berners-Lee had this thing they called the layer-cake diagram of the web; the semantic web was the top several of them. The crowning thing at the very top of that layer cake diagram was a little block called trust.

Dean:
And I mean, this is so prescient of Sir Tim, back in the 90s, to realize that that is actually the key, you need to get answers that you trust. We had no idea what world generative AI was going to be playing in that stack back in those days, but that’s the key thing. So the LLMs will famously spout out with great confidence, pure rubbish. And then you, as the listener are now, if you do it as a chat sort of thing, are charged with the problem of sorting that out. Enter a knowledge graph. How could you do this in a more automatable or supported way? So enter the knowledge graph if you no longer think of the LLM as being your point of contact to the system. If you think of the knowledge graph as the point of contact, you go up here, and you say, “Gee, I want to know which of our customers has the biggest month-over-month return rate?” Some complicated question, and then we go to the LLM and ask it that it’ll make something up.

Dean:
If you go to the LLM and say, “Write me a query to go to my curated, supported, trusted data,” and ask that question, then you get an answer that does not come from the LLM; it comes from your database, but there’s still, of course, a flying in this ointment. How do you know the query is the right query? But now this is a constrained question; you can look at the query, you can look at the question. This is no longer looking at a whole database of gigabytes of information or, heaven forbid, the whole web. This is not an open-ended question where you have to go off and do your own research and basically answer the question for yourself. Now you’re looking and saying, “Well, yeah, in fact, yeah, here we’re looking at the customers; we’re going to this database here. We’re looking at those customers in those databases; we’re comparing them.”

Dean:
And if something went wrong, “Oh, no, no, don’t ever ask that.” In a lot of places I’ve worked, there’s always somebody who’ll say, “No, no, no, never ask this database about that thing; they don’t curate that correctly.” And I’m sure you’ve experienced this as well. “If you want to know answers to that question,” customer return rates whatever, “You go to this database, not that database.” Well, you can put that sort of stuff into the knowledge graph as well. So you can put all this advice in the knowledge graph, where you can review it, curate it, update it, steward it as it were. And when you get your query, you can look at the query. “Yeah, is that the right query?” Now, that takes some understanding of queries, but actually SPARQL is a pretty easy query language to learn if you don’t know query languages, at least to read. You go, “Okay, yeah, there’s the customer coming from that database, the customer from that database, their return rate, their return rate, subtract… Yeah, yeah, that’s what I want to do.”

Dean:
Or, “No, no, no, no, no. A new customer isn’t onboarded; a new customer is somebody who made an order. No, no, no, that’s all wrong.” You can actually review its work. And that’s the thing that LLMs lack is a way to review their work. But if you force them to put this extra step in, now you can audit this, and this is the stuff that, basically, Juan and I were doing our research on, and I’m not talking, out of school, to say, this is what data.world is doing. That’s the part that we talk about. Yeah, we do this extra step so that somebody can audit this. What we have not done yet, but this is right on our roadmap, is how do you take that structure now and turn it into an explanation affordance for various categories of users? You can imagine all sorts of categories of users that have different explanation requirements. How do you classify those users? How do you cater to each of them for the needs that they want? So that when they’re all done, trust, they can trust their answers. That is the holy grail.

Larry:
Yeah, that’s brilliant. And that meshes 100% with the stuff I’ve learned in the UX world and the conversation-design world. In interactions with these kinds of agents, or these kind of technologies, that trust is paramount, and explainability and transparency are key, and knowledge graphs give you that. Yeah, very cool. Well, hey Dean, thank you so much. Oh, hey, one very last thing before we wrap up. If folks want to follow you online, what’s the best place to connect or follow you?

Dean:
Well, I’ve been kind of remiss in keeping my Medium up to date, but I still have a lot of good stuff there, and I keep promising I’m going to write some more stuff there. So, actually, medium.com/@dallemang I think is the URL. I’ll give you the correct URL that you can put in the notes; I’m pretty sure that’s the right one. That’s where I write about all this stuff. I had lunch with a high school chum the other day and was talking about generative AI; he said, “Gee, you’ve done AI. Can you tell me some stuff about it?” And he wrote to me, saying, “I’d like something to read about this.” And of course, I can’t help but say the first thing you should read about generative AI and knowledge is my Medium, because I’ve written a lot of stuff there.

Dean:
Most of it last year. I really need to do some more stuff this year. But that sort of gets to a lot of the basics of the semantic web, the basics of how that fits in with LLMS. A couple of the early guessing experiments that I did back then before we did systematic stuff like Juan and I’ve done together, but a lot of the things that I talked about even a year ago have turned out to be not correct. But yeah, this was the right line of research to follow. And so I’m pretty pleased with how that’s all turned out.

Larry:
Cool. Yeah, I’ll definitely link to that. I love that blog. Yeah, you haven’t been as busy lately.

Dean:
I haven’t updated in a while. Yeah. I got to the point where I had so many topics I couldn’t figure out how to actually sort them out and focus them. A typical writer’s block problem, the glut of information inside of your head that you’re trying to figure out how to turn into nuggets that you can write about.

Larry:
Nice. Well, you’re really good at that, I got to say. So. Well, thank you so much, Dean. This is really fun. Always fun to talk with you.

Dean:
Yeah, likewise, Larry. And I’m glad we’ve gotten a chance to schedule this and looking forward to seeing the final product.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top