Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Mike Pool sees irony in the fact that semantic-technology practitioners struggle to use the word “semantics” in ways that meaningfully advance conversations about their knowlege-representation work.
In a recent LinkedIn post, Mike even proposed a moratorium on the use of the word.
We talked about:
- his multi-decade career in knowledge representation and ontology practice
- his opinion that we might benefit from a moratorium on the term “semantics”
- the challenges in pinning down the exact scope of semantic technology
- how semantic tech permits reusability and enables scalability
- the balance in semantic practice between 1) ascribing meaning in tech architectures independent of its use in applications and 2) considering end-use cases
- the importance of staying domain-focused as you do semantic work
- how to stay pragmatic in your choice of semantic methods
- how reification of objects is not inherently semantic but does create a framework for discovering meaning
- how to understand and capture subtle differences in meaning of seemingly clear terms like “merger” or “customer”
- how LLMs can facilitate capturing meaning
Mike’s bio
Michael Pool works in the Office of the CTO at Bloomberg, where he is working on a tool to create and deploy ontologies across the firm. Previously, he was a principal ontologist on the Amazon Product Knowledge team, and has also worked to deploy semantic technologies/approaches and enterprise knowledge graphs at a number of big banks in New York City. Michael also spent a couple of years on the famous Cyc project and has evaluated knowledge representation technologies for DARPA. He has also worked on tooling to integrate probabilistic and semantic models and oversaw development of an ontology to support a consumer-facing semantic search engine. He lives in New York City and loves to run around in circles in Central Park.
Connect with Mike online
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 22. The word “semantics” is often used imprecisely by semantic-technology practitioners. It can describe a wide array of knowledge-representation practices, from simple glossaries and taxonomies to full-blown enterprise ontologies, any of which may be summarized in a conversation as “semantics.” Mike Pool thinks that this dynamic – using a word that lacks precise meaning while assuming that it communicates a lot – may justify a moratorium on the use of the term.
Interview transcript
Larry:
Hi everyone, welcome to episode number 22 of the Knowledge Graph Insights podcast. I’m really happy today to welcome to the show Mike Pool. Mike is a longtime ontologist, a couple of decades plus. He recently took a position at Bloomberg. But he made this really provocative post on LinkedIn lately that I want to flesh out today, and we’ll talk more about that throughout the rest of the show. Welcome, Mike, tell the folks a little bit more about what you’re up to these days.
Mike:
Hey, thank you, Larry. Yeah. As you noted, I’ve just taken a position with Bloomberg and for these many years that you alluded to, I’ve been very heavily focused on building, doing knowledge representation in general. In the last let’s say decade or so I’ve been particularly focused on using ontologies and knowledge graphs in large banks, or large organizations at least, to help organize disparate data, to make it more accessible, breakdown data silos, et cetera. It’s particularly relevant in the finance industry where things can be sliced and diced in so many different ways. I find there’s a really important use case in the financial space but in large organizations in general, in my opinion, for using ontology. So that’s a lot of what I’ve been thinking about, to make that more accessible to the organization and to help them build these ontologies and utilize them effectively.
Larry:
Nice. One of the intellectual I guess foundations of that kind of practice is what we call semantics. Anyhow, I want to read part of that post you made on LinkedIn, which started a great conversation. One of the things you suggested, “I think we need to impose a moratorium on the use of the word semantics. The reason is simple, it’s ironically a term lacking any precise meaning while we assume it’s communicating a lot.” That’s brilliant, can you elaborate on that a little bit? Was there a particular, did something inspire that or has it just been on your mind?
Mike:
Yeah. I mean, it’s mostly, one term that often triggers it for me is this term I see within, let’s call it this community of practice … I see used very, very frequently, people will say, “Let’s look at the semantic meaning.” So this redundancy in terms, that we said, “Well, what in the world?” But we use it for all kinds of things. We say we need a semantic solution, we need the semantic meaning. And very often what that ends up being when we drill into that, it’s just not always clear. The term I think in some sense has become either too vague … it’s unclear of what precisely it means. Or it’s a shorthand for something else, that we’re not actually saying we’re going to capture meaning. We’re saying, we’re going to use this particular set of tools or something like that. So my concern is that it’s sort of lost. We know when we say it that we mean we’re going to use these particular set of tools, these particular set of languages, but to the people with whom we’re communicating that might remain completely unclear. So yeah, that’s my concern about the way we’re using the concept.
Larry:
Yeah. That’s really interesting that … it’s not laziness, it’s like heuristics or something like that, that people use all the time to just try to advance whatever conversation they’re in or project or whatever. It’s just like, “Oh yeah, we need a semantic thing there,” or something like that, it sounds like. Or they’re thinking of possibly 20 different things or they just say, “Oh, semantic is the closest word I know to that idea,” that we need to advance this.
Mike:
Yeah. I mean, an example is … Because I think, as I noted somewhat ironically, I herald myself as a semantic technology practitioner or something like that. After you said to me, “Well, what in the world is semantic technology?” It’s a good question. If I create a property graph, there’s part of me that says, well, that’s not really semantic but a triple store is. It’s like, well, what’s the dividing line? What precisely makes it count as semantic or not? It’s a little bit hard to pin that down.
Larry:
The way you just said it, it’s almost like there’s an on/off switch someplace. But I’ve seen, there’s a lot of representations of what various people have called something like the semantic spectrum, from just term lists to glossaries to thesauri, to the ontologies, that kind of thing. It’s easy enough to disambiguate between each of those things I just mentioned, but is there something like a spectrum in there? Is that why people are grasping for words, do you think, to describe exactly what they’re talking about in the moment?
Mike:
Yeah. As I said, I think that’s part of the problem. As I said, the people with whom I often communicate, I think we more or less mean it as a shorthand for using RDF OWL. And that might be as simple as using SKOS and creating a simple taxonomy with that, or creating a very elaborate OWL ontology. But it’s interesting because, let’s say, we create a taxonomy in SKOS. Well, is there any reason that if you just had that taxonomy and you didn’t bother to put it in SKOS, you just put it in an Excel spreadsheet with appropriate indentations? We’d say, well, how does the SKOS, or how does the RDF magically capture the meaning where the spreadsheet didn’t, right? It’s a little bit unclear. But I think other people use it differently. I think there’s lots of people who would say using a property graph is a semantic solution. We’re capturing knowledge, et cetera, in it. So I think it varies a little bit, that’s again, part of the point. But I do believe people with whom I communicate, that’s the shorthand we’re using. It’s like, this is either technology that we use to extend RDF OWL, or it’s a knowledge graph that encodes that knowledge. But that’s often what it means in my space I think.
Larry:
Also, you mentioned in that post too, I think, when you’re talking about, if your intent is to capture meaning, that there are other ways to do that technologically. You talked about just an old-fashioned ERD or a graph schema or a Python script that captures something in some project you’re working on. And even I think you also alluded to, or maybe it came up in the discussion, it could be a natural language thing that you say to an LLM that it could discern. But I guess that gets at, what is the amount of meaning you need to capture? Does that make sense? What’s the intent behind your attempt to do something semantic in a technical project?
Mike:
It sounds like a straightforward question, but it’s actually a very good one. Because as I said, that if we say, well, we’re trying to capture meaning, you could write a Python script that does it, or there’s a lot of different ways to do that. I think this whole, at least the background that I have in this, when we’re talking about capturing semantics, what we are really concerned with is really trying to say, can we get a computer to reason in the same way that we do? Can we get the computer to respond to natural language prompts in a similar way that a human can? That’s kind of what we meant by semantics. But then if we try to talk about that in the technology space, what exactly does that mean? Let’s take the Python script example. If I said, oh, I’m trying to solve this problem, I have this search engine and every time people, they’re searching for recipes, that if they’re searching for recipes with fruit in it or recipes with vegetables in them, I want it to understand that squash is a kind of vegetable or that apples are a kind of fruit or something like that. It’s like, so what if we just embedded all that in a Python script? Why isn’t that semantic? We say, in some sense it is.
Mike:
But I think what we get at is, insofar as we did that and then we had it deployed for that very specific application, it’s not very reusable. It’s not very scalable. So part of what we think when we’re trying to capture semantics, whatever that means, is, we’ve created something that we can pick up and reapply to different pieces of software to use. So if I write a Python script, even if it listed out all the different types of vegetables and it did the subtyping, we’d say, yeah, but it’s not very semantic because you’ve built it for this particular search engine. And now if I have another thing where I’m trying to reuse, where I’m trying to figure out stuff about vegetables, you’ve built that thing for that search engine, it doesn’t reapply. So we talk about the term, data-centric is another term that gets thrown around in the community a lot. But I think there’s something that we’re trying to do, is to say, we want to have some kind of conceptual capture that we can pick up and move around, that’s really part of it, and extend easily. So I think, why is our intuition that a Python script isn’t really capturing semantics?
Mike:
It’s because you built that Python script for a single application, and you can’t pick it up and use it in a different application.
Larry:
I think we’re starting to disambiguate the use of the word semantics in technology versus philosophy or linguistics. So that’s a very specific bunch of needs. Especially, I had not really thought that much about the need for reuse, if you’ve gone to that trouble of identifying the meaning of something and how you’re going to use it in that context. But that gets to, you mentioned both a specific application or use case like search. You also mentioned the notion of data-centricity, and that notion that an enterprise having all of its data in a neutral setting that could be used in any number of applications, that presumes you can have meaning independent of an end application. Is that actually doable with semantic practice? I mean, when I think about it, and when I articulate it that way, it makes me wonder where there might be details to be resolved.
Mike:
Well, I think that’s the right end goal to keep in mind, is to say … and there is a challenge with that. Insofar as, sometimes if we get too focused on just worrying about representing, capturing semantics as it were independent of the application, we end up focusing on things that might not be salient to a particular use case. So I never say people should be capturing semantics without worrying about applications. Really, we need these use cases to figure out what in that semantic capture is relevant. But the idea would be that, yeah, very definitely whatever you’re capturing, even if we were capturing it for a particular application, to be thinking all that time about, if there’s another application that’s going to need this information, can I extend this easily? And then that’s what allows you to have that independence. I think again, that’s as ontology practitioners what we should be thinking about, worrying about is to say, okay, am I creating an ad hoc model here? Because I know exactly how this application works and I know exactly the performance requirements it has, so I’ve created this ad hoc model. But now if someone else … in the banks we talk about the front office and the back office and the middle office.
Mike:
What you want to say is … Yeah, we get it that the front office is trying to work with customers who are coming in on the website, and the back office is trying to resolve accounts and things like that. But what you’d ultimately want is to say, I want to have some conception of the customer or client, that I can actually reuse that thing across all those different spaces. Even though if I were to tailor build a schema for each of those situations, they might look quite different. So if we think, yeah, we’ll build this for front office but is it going to extend easily for the back office requirements? I think that’s where you start getting into semantics. Now you’re thinking about, there’s something in common that all these different applications are talking about, let’s capture that. It doesn’t have to be a complete full representation of, but it has to be something that scales easily and extends easily I think.
Larry:
Yeah. As you were describing that front, middle and back office thing, I’m thinking of, I’ve done a lot of work in enterprises where you have customer-facing, designer and product-facing and then systems-facing at the very bottom, analogous to what you just said. But then also I think about ontology practice, and ontology is usually constrained by a domain. You’re aware of and constrain your concerns to that domain. I’m wondering, is semantics the boundary of that domain?
Mike:
I mean, I don’t think it’s necessarily the boundary of the domain. As I said, I think in some sense the semantic part of it is to say, have you created something that allows us to traverse those domains? I think good pragmatics for people doing modeling is to be domain focused. If you’re creating a model of a human being and you’re worried about it from a bank perspective versus a medical perspective, those models could be very different. But then again, I think a truly semantic model would recognize to say, oh okay, here’s how we can pick it up and extend it so that it’s reusable across both those domains. I mean, I think domains are useful, as I said, so we don’t try to model the world. I suspect that part of what happens in this space is … you think about what we tried to do in the early days of AI, is to say, well, we want to do knowledge representations, we would just represent everything. We would just say, we want to know what a human being is. Well, they’ve got a heart and they have a brain and they have 10 fingers and they have all this stuff.
Mike:
I think we started to recognize, yeah, you can do these elaborate knowledge representations but to make them useful, you really have to pick out a domain and be specific about it. Yeah. So I don’t know the extent of overlap of domain and semantics. I’m not completely sure about it.
Larry:
I mean, it occurs to me that maybe there’s just semantics in what you mean by the boundary and what you mean by the edges of that domain. That’s probably… The other thing, we’re talking mostly about ontology here … and again, I alluded earlier, I mentioned earlier that semantic spectrum of just, like a list of words. You talked about, maybe a spreadsheet would serve just as well as ontology in some situations, or things like that. I guess, is that a maturity thing or just a utility thing? That range of, sometimes you just need a list, sometimes you need a definition of it, sometimes you need to understand how it’s organized. Sometimes you need a full-blown understanding of its ontological context. Is that a spectrum or a maturity range? All of those things I just mentioned, could those be considered part of the semantic world?
Mike:
Well, like I said, I guess what I would be inclined for us to be thinking about in these situations is a little bit less about what is semantic, and a little bit more to your point, what do we need to solve this particular problem? So like a simple list of words, a simple control vocabulary. If we think about document tagging or something like that, it may be the case we don’t need to have a really elaborate schema for doing that. But we do want to say, well, we want to have a control vocabulary so that you and I, when we’re tagging documents, we’re using the same concepts. It makes it easier to find. Now that might be something, we want to make it extensible but it never has to become much more elaborate than that. Now, we may want to ultimately link it into other taxonomies or something like that, but the point would be, are we using semantics there? Well, sometimes I don’t care all that much. It’s just, it’s to say, well, how clear are you being when you’re saying this is a semantic solution? I don’t know. But if I say, I’m actually creating this control vocabulary … and by the way, I’m encoding it in SKOS because that gives me a framework with which I can actually simplify linking it to other taxonomies, et cetera, am I capturing meaning?
Mike:
I don’t know, maybe I am, maybe I’m not, or maybe I’m making it easier for the systems to interact with each other. But there’s some utility in doing that, right? Instead of saying, oh, we’re capturing semantics … in which case you’re saying, what, are you creating an elaborate ontology of documents? It’s like, no, it’s just a tagging system because that’s all we need. So we don’t have to overkill it there. And if we say, oh, and by the way, we think it makes sense to not just store this on a spreadsheet but store it in a SKOS representation because that makes it easier to store in our knowledge graph, that’s a reasonable thing to do. But again, I wouldn’t necessarily say that creating a list of document tags, you’re not capturing much meaning there, you’re just making it easier to tag.
Larry:
Right. You just mentioned the idea that you’re tagging something with those concepts. It’s that connection between the labels, the terminology you used to describe a concept, and that ‘things not strings’ notion, as Google articulated it. That’s funny, I hadn’t thought about that in the context of this. But is the meaning ‘the thing’ and ‘the string’ just application or context specific, I wonder?
Mike:
Well, I mean, I think the ‘strings not things’ thing is actually really this term everyone started to use, there’s an awful lot to it that gets to what we worry about as ontologists. It’s extremely pithy, which is the beauty of it. But that really is part of it, is that what we want to do is create objects. Now, if I reify an object rather than have a string in a database, have I captured meaning? No, obviously not, right? There’s nothing meaningful about that. But I have created something that I can link to other things or I can in the future make other assertions about, et cetera. So I haven’t necessarily captured meaning, but I’ve created a framework in which I can start to link it to other things and maybe discover meaning, et cetera. So whereas I used to think the ontologist’s job was to verily create a thing, not a string, it also meant articulating everything that we knew about those things. And now I think maybe what the realization we’ve come to is, just build that thing. Because when you build that thing, you can link other things to it, or you can say more things about the thing, et cetera.
Mike:
So there’s an awful lot of the strings … it’s ‘things not strings,’ the order matters. But yeah, that actually is just a very, very insightful three-word phrase I think.
Larry:
The way you just elaborated on it makes me picture ontology practices like building lattice works that you hang things on. So you can get them … the lattice being the domain and then overwrought with that analogy already. Does that make sense? Is that part of the practice is, okay, first let’s get clear on the thing … the concept, the entity, the object, and then semantics comes a little later in the whole process?
Mike:
Well, as I said, I don’t know if there is a magic moment where we say we’re capturing semantics. I think it’s more what we’re trying to do in the space. But sometimes capturing the thing is not nearly enough, right? My favorite example I always throw around is the concept of merger, which I discovered at one point can have very different formal meanings. It can mean two companies come together but they both dissolve and a new one is created. Or it can just simply mean, two companies came together and one survived. But those are very importantly different concepts and it’s useful in our formal systems to capture that, right? So I wouldn’t say, oh, we just have to reify this concept of merger and everything will be okay. It’s actually important in those cases to formally capture so that our computer systems can actually recognize to say, oh, there’s two different databases capturing mergers but they’re using very different concepts there. So capturing meaning is important, but it’s important use case by use case as well.
Larry:
Yeah. What’s in practice the capture of the meaning of, how much of it is just talking to people and looking at other ontologies or other expressions of things? I guess you as an ontology practitioner, how are you confident that when somebody says merger in the context of the system you developed, that they’re using it in the right way or that it’s understood in the right way?
Mike:
I mean, that ends up, as I said, I’ve thought and worried a lot about building ontologies in these large organizations. I think a key part of doing that is to not necessarily be too ambitious about capturing that meaning at the outset where you say, hey, we know there’s this important concept called the merger … We know customers, the other very famous one, where it turns out that you go across an organization and everybody’s using the term differently but they thought they meant the same thing. You really almost need in those situations to say, well, let’s create, let’s capture this notion very generically. Maybe we just have a node in our ontology for the customer, and then we want to facilitate the ability of different groups to actually extend that thing. It might turn out that, oh, when they actually capture the formal meaning of it in their space, it turns out to be very different. Now, when you’re actually capturing semantics, I found it’s sometimes important to have mechanisms in place. So you don’t try to resolve all that upfront and say, oh, the problem that plagues our company is, we don’t have a unified model of what a customer is. But in order to do that, it turns out you have a six-month workshop and there’s people throwing things at each other because they’re so upset and enforcing their meaning.
Mike:
It’s sometimes better to say, well, there’s this general notion of concept of a customer. We can all agree that it’s somebody who exists outside the company, that has some sort of relationship with the company, and then people can extend their own definitions and things like that. But capturing semantics is sometimes a very iterative process, in my experience.
Larry:
That notion of iteration … I’m thinking back, this comes up almost half the conversations I have, Jim Hendler’s notion of, “A little semantics goes a long ways.” I wonder if that’s part of what he was getting at, just think about it and put it in a little bit at the start, and then you can do more with it later. I don’t want to speak for him.
Mike:
Yeah. No, that’s what I think about that too, is that if you create, you can almost think of it as almost a category or something like that, and then people can say, oh yeah, you’re talking about customer, I’m talking about customer. And now when we actually try to see how we’re using it within our systems, we want to capture more meaning around it … because we want to validate, right? Somebody wants to bring in a new customer database or something like that, we want to say, oh, we want to capture the meaning more formally. It turns out people will be using it in different ways. Then you iterate on that and you say, okay, well, maybe we need to have different classes of customers, or maybe we need to have different named graphs in which we allow for different SHACL constraint sets or something like that. But I think that’s actually an interesting point when you’re saying that a little bit of semantics goes a long way, create your beachhead of meaning and then you can iterate on it and capture it more extensively or formally, as you iterate and try to understand it.
Larry:
I love that image of a semantics beachhead. I’ll share with you after this, and I’ll put it in the show notes, my favorite meme about language.

Larry:
Anyhow, but hey, I can’t believe it, Mike, we’re coming up close to time already. But before we wrap up, is there anything last, anything you want to revisit from the conversation or you just want to make sure we share before we wrap up?
Mike:
Well, I mean, I guess the one thing that I think about a lot when I think about what I’ve been trying to do in the semantic space, is the notion of capturing meaning. We discover and we find that these large language models, or even word embeddings and things like that, actually do a lot of the work that we thought we’d have to do with ontologies. There’s a certain kind of effort devoted to capturing semantics that has been solved for us. So there’s a new question of, what is the role of this more formal knowledge representation, the thing that we mean when we say semantics in this world? I think it turns out it actually means what we ended up doing, and that’s creating these knowledge schemas or these models that make it easy for us to interoperate and, to your point, that we extend on a use case basis, rather than trying to have a full massive capture of it. But those ontologies end up being really important, not so much for capturing the meaning but more for being able to interoperate across these different systems, and be able to then plug into these LLMs that can do some of that semantic understanding for us. An LLM can maybe disambiguate terms and then send queries over to a knowledge graph that’s been organized with respect to some ontology.
Mike:
We’ve let the LLM do some of that semantics for us, but the ontologies still play a central role for us to now pass information back to that LLM. So I think, as I said, I’m realistic enough not to suggest that we really stop using the word semantics, but it really is a word that often requires further elaboration. And I think, in a certain level of humility, that it’s not ontologies alone that are capturing these semantics.
Larry:
Nice, I love that. That’s a great … and that’s how that post came across, as a conversation starter. This little conversation was just a continuation of that, there will be plenty more to talk about. I’m sure it’s going to evolve quickly as we see people using, figuring out the true role of ontology in this new LLM-heavy world. Hey, one very last thing, Mike, if folks want to connect with you or follow you online, what’s the best place to find you?
Mike:
Yeah, probably LinkedIn is the best place. I think I’m just Mike Pool on there, LinkedIn.com/in/MikePool, I believe is the simplest way. I think there’s a disambiguation problem, there may be more than one Michael Pool. But there’s no E in my surname, so that might be one way to help you disambiguate.
Larry:
I will disambiguate it in the show notes, for sure, and include it there. Well, thank you so much, Mike, I really enjoyed the conversation.
Mike:
Yeah, my pleasure. Thank you for inviting me. It was fun to talk about this.