Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Every enterprise nowadays is awash in data, content, and knowledge, the understanding of which is all too often available only in employees’ heads.
Forward-thinking businesses are moving to knowledge graphs to capture that tacit knowledge so that they can better understand and use it.
Andreas Blumauer shows how those graphs work best when they’re accompanied by a domain knowledge model, creating a “semantic layer” that provides a vivid map of your business knowledge.
We talked about:
- the recent merger of his company, Semantic Web Company, with Ontotext to form a new company, Graphwise, which aims to help enterprises build their semantic layer
- the 20-year-old origin story behind his definition and description of a semantic layer
- the emerging trend he sees of data, content, and knowledge people coming together, often around explorations of language and content structure
- how getting domain knowledge out of people’s heads and into multimodal AI architectures can streamline business research
- how a semantic layer provides a map of your business knowledge
- the importance of a domain knowledge model in a semantic layer architecture
- the inadequacy of “desktop data integration,” the practice of calling colleagues, consulting varied sources, and otherwise searching through ad hoc enterprise knowledge sources, and the stress it can cause
- how a domain knowledge model can connect and reveal knowledge across an organization
- the surprisingly small footprint of the domain knowledge model in an enterprise knowledge graph, typically just 1% or so of the semantic layer
- how LLMs can help in the discovery of data and the construction of knowledge graphs
- the two main elements of the semantic layer: domain knowledge models (taxonomies, ontologies, etc.) and an automatically generated enterprise knowledge graph
- the different perceptions of the value of the semantic layer across data and content professionals, and how the arrival of gen AI has resulted in them talking together more to each other
- how the semantic layer can facilitate the alignment of vocabularies and the understanding of data across business divisions
- the benefits of a hybrid centralized-decentralized/global-local “glocalization” semantic layer strategy
Andreas’ bio
Andreas Blumauer is SVP Growth at Graphwise, and CEO and founder of the Semantic Web Company (SWC), provider and developer of the PoolParty Semantic Platform and leading solution provider in the field of semantic AI and RAG. For more than 20 years, he has worked with more than 200 organizations worldwide to deliver AI and semantic search solutions, knowledge platforms, content hubs and related data modeling and integration services. Most recently, Andreas has been involved in the development of AI-powered ESG solutions for companies and investors.
A globally recognized thought leader and author in the field of semantic AI and graph technologies, Andreas has helped define and implement knowledge and AI strategies for various industries and domains. He is the author of “The Knowledge Graph Cookbook: Recipes that Work”, a practical guide to building and deploying knowledge graphs in organizations. He is passionate about enabling clients to harness the power of semantic AI and graph technologies to achieve their business goals and make a social contribution.
Since 2022, Andreas has focused primarily on developing AI-based solutions that help organizations implement ESG and sustainable systems.
Connect with Andreas online
Resources mentioned in this episode
- Knowledge Graphs, LLMs and Semantic AI LinkedIn group
- From Data to Trust: Leveraging Knowledge Graphs for Enterprise AI Solutions webinar

Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 14. Data, content, and other business professionals have long known the benefits of sharing company lore in enterprise knowledge systems. Many of those systems now include a knowledge graph. Andreas Blumauer says that the thing that really lets you leverage your knowledge graph is a semantic layer that includes alongside your graph a human-crafted domain knowledge model that captures, and represents in a computer-readable way, the tacit knowledge in your enterprise.
Interview transcript
Larry:
Hi everyone. Welcome to episode number 14 of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show Andreas Blumauer. Andreas, you may know him as the CEO and the founder of the Semantic Web Company, one of the venerable companies in the Semantic space. But more recently, his company has merged with Ontotext to form Graphwise, where he’s now the SVP, the Senior Vice President of Growth and Marketing. So welcome Andreas, tell the folks a little bit more about what’s going on these days.
Andreas:
Yeah, sure. Thanks Larry. Thanks for the invitation to this great podcast series. Yeah. So yeah, exciting days. Currently experiencing the merger between Ontotext and Semantic Web Company. And actually it has come quite naturally. So we have been working together for many years and inside of our Semantic suite, PoolParty Semantic Suite, we always have been using, at least for many years, GraphDB, the core product of Ontotext. So we got closer and closer, and finally the time has come to execute on this bigger vision, which is all about helping organizations create semantic layers. And this is where we currently stand, and it’s definitely a very, very exciting time for all of us and looking forward to the next steps.
Larry:
Yeah, it’s easily the biggest news in the industry lately, but yeah, the whole point of it is to support the semantic layer, and you’re one of the first people to talk about that. I mean, if you Google it it shows up in a lot of different contexts. But in the context of these enterprise architectures that support knowledge representation tech, you’re one of the first people to talk about that. I don’t know how familiar people are or whether they need to have your notion of a semantic layer disambiguated from what they’ve got, but tell me a little bit about what you think of the semantic layer is.
Andreas:
Yeah, no, this has been a topic I always was, since the beginning of my career, quite a lot of discussions were around knowledge management and how to bring together digital assets to the actual knowledge people have in their heads, so to speak. And there was this, I would say, first wave of knowledge management discussions where it turned out to be quite clear that on the one side, at that time it was around 2005, six or so, rather let’s say new digital community where documents contained data information, and the other side of the community said, it’s not knowledge. So in a document, you will never find knowledge, it just can be translated or transformed into knowledge as soon as it has been recognized by human beings as such. And so I was always feeling between sitting between those two communities.
Andreas:
So on the one side was excited since day one. Okay, where does AI bring us? And on the other side, I was totally understanding, okay, we should not substitute human intelligence at all. So I was always in this kind of tried to bridge those two communities. And so I was thinking to myself, looking at the contemporary enterprise data architectures at those days there was a missing layer. There was something, there was the data layer, there were the document repositories, and on top there were the applications trying to represent some kind of business logics and nothing in between. So there was really a big, big hole. So I thought to myself, at that time there was topic maps around even still, and then RDF came up, semantic web standards, more and more semantic web standards got developed. So I was trying to introduce this new layer and they called it semantics layer at that time already, to the community.
Andreas:
And everybody stared at me as if I was an alien, really. So I felt like, okay, maybe I said something wrong, it doesn’t make sense to the others. And I was a little bit puzzled by that, I thought, but this is still, I mean, and that really kept me going into this direction. I always thought to myself, there is something very important. And now I think we have arrived a new era where the data people, the content people and the knowledge people come together and now they discuss how should a semantic layer look like? And I totally understand each of those have different angles and different ideas how it should be set up. But let’s take a closer look at it because it really depends on the use cases on top and the strategies of the AI, the overall AI strategy we all now want to develop driven by Gen AI. Obviously it has become even more important than ever before to have a semantic layer in place in our siloed world, in our siloed environments.
Larry:
Yeah. And that notion of the, as you were saying that you’re making me realize that the LLMs are kind of at that, to my mind, there’s emerging at the highest level that architecture, they’re really a great interface to a lot of this knowledge that we have, but they’re not the best keepers or understanders of that knowledge, whereas the RDF stack of knowledge representation stuff is. Can you maybe talk a little bit about the relationship between LLMs and other AI tech and knowledge graphs and how the semantic layer facilitates their interaction?
Andreas:
Right. I would start maybe like that. First of all, I think language is something which allows us to develop knowledge without our ability to talk and to use language. We’re not able to develop anything more complex than let’s say what we need for our survival mode typically. So obviously it’s a driver, but knowledge is something which is a bit more rigid than language obviously. So it’s really some kind of structure which has approved to be valuable under certain circumstances. So obviously it’s all about context. So the same fact or the same knowledge wouldn’t be maybe as important as it is in certain situations. So it’s all about understanding how knowledge should get embedded in the process. That’s also very important. A lot of different types of knowledge need to get handled in different modes. So just using language, and by that probably more the unstructured data side of things is just one way moving forward.
Andreas:
Highly structured data, on the other side, when you just want to describe a particular, let’s say, element of your landscape, like let’s say the product landscape, of course, it’s just one part of the entire story. So it’s in all cases important to link different aspects in a way that you have more ways to get value out of your data. So as soon as you connect your data, you increase the value of your data. So it’s really a network effect is kicking in when you start to contextualize or when you’re able to contextualize your data in a dynamic way, but you need to know how to contextualize it, and that’s nowhere written. It’s neither in your data silos, in your databases, nor in your, let’s say technical documentation, which is probably more on the unstructured side. So it’s just in our heads, and this is very important because LLMs obviously can only deliver answers to what’s written down somewhere in a database, typically in an unstructured, rather on the unstructured side.
Andreas:
Gen AI is pretty good with dealing with unstructured data, mixing in the more de facto data around let’s say products, employees, projects, competitors, whoever. This is more on the structured side. So it’s all about multimodal data first of all. So we want to link different data types first of all, and then second, we need to enrich that with domain knowledge. So this is the background knowledge we need to better understand when data, which type of data in which situations should get used. So this is something we completely miss out in our data architecture currently. Domain knowledge models are nowhere found, just in our brainware. And by that we are in strong need to do desktop data integration the whole day. So it cannot be automated yet, really, LLMs cannot do that. LLMs cannot link data across different data silos in a way it’s really usable. They cannot understand the inherent structure of structured data, they just chunk it up, it’s chunky data back to databases, really try to make chunks out of data, which has been put in in a very nice structure beforehand, and that’s just thrown away.
Andreas:
So that’s one of the biggest, let’s say disadvantages of contemporary rock architectures we currently see. And this is exactly where a knowledge graph kicks in. That’s where a knowledge graphs can help to sustain the knowledge, how content and data should get interpreted. This is really what’s written down. For instance, in an XML schema, obviously let’s say data, when you talk about technical documentation, we will far better understand the meaning of the data by the structure we have given to it. And when you just throw that away by making random chunks, all of that, then we have a problem. And this is where LLMs really fail.
Larry:
And as you’re describing it, you used a couple of minutes ago, you used the word landscape and you also said something about how to navigate all of this stuff you’re talking about. And one of the things you’ve talked about is using this semantic layer as the map for this. And you just talked also about the importance of domain knowledge. We’re just all wallowing in data. There’s plenty of that out there, but the domain knowledge, pulling that knowledge out of people’s heads, that’s been an ongoing challenge in knowledge management for decades. And now it becomes even more important. So I guess once you’ve extracted that domain knowledge out of the human beings and you’ve got it, tell me, is the map a good way to think about what you do with that knowledge?
Andreas:
Yeah, I think that’s a nice metaphor. So we all live on different data silos, which do something really good per application logic. So I think we still live in this world of having one silo per application. And as soon as you want to navigate following, let’s say the domain logics, the logics of a given domain, you start to really get lost. You need to ask around, “Hey, where do I find this slide deck? Where can I get a report on this? Who is the expert for A, B, C?” Et cetera. You really would need a map which helps you to not just look around and search around, but really navigate as you would do when you drive by car or when you want to fly, let’s say from Sofia over Vienna to New York and to San Francisco. So you would have a map for that, you would not ask around, “Hey guys, how can I get from A to B?”
Andreas:
This is currently how we work in our enterprises. So there’s no map typically, I mean data catalogs try to do their best, but still the metadata and the ways to search, for instance, data is very limited in such environments. And so I think the next big step really is semantic data fabric. So where we have all kind of possibilities to navigate the map and it’s driven by the domain knowledge model. So it’s really describing how things are connected, how the business objects are connected. Let’s say a new employee comes and you would like to find out where this person has been working before and then find out, okay, has this company where the person has worked before, does that play an important role in my supply chain? And if so, so which skills did that person acquire at the time having worked there? Can we reuse it maybe for a particular project we currently jump on?
Andreas:
And if so, who is the leader of that project and who is the customer of it? Which products are we going to ship there, et cetera, et cetera. So which training does that person need to get up to speed? And all these questions typically keep us busy. And what we do the whole day is desktop data integration. So we pull together data and information, we call, ring up colleagues. I have Zoom meetings here, put down some notes there, put something in a spreadsheet here, write something in a note there. So we do it on our desktop. It’s not collaborative at all. And the problem really is even the best enterprise search engines won’t help us because they do not connect the dots. The only thing they do is they help us to search over one, two end repositories probably, but it speeds up, it accelerates this task, but it still doesn’t give us a map at our hands to navigate the data points.
Andreas:
So this is exactly what we are going to address with the semantic layer. This is where people get better ways to navigate the landscape, the data landscape, do not waste the time with desktop data integration, but really can stay focused on the actual work, which is the knowledge work, which only comes after you have aggregated the data and linked it all together. Only then actually the really, let’s say most valuable step comes. And here we typically are completely stressed already. Now knowledge workers these days are completely stressed. They’re under pressure because they have so much of research to do first-hand before they can really get into the weeds and come to the point where they can make decisions. And it’s all about decision making at the end of the day.
Larry:
And that’s interesting, when you say that I’m reminded of one of the older forms of AI decision support systems, but this is like that on steroids, it sounds like. Once you’ve mapped your actual, all the business objects in your domain, the domain knowledge that you’ve managed to extract from folks heads and put it in some kind of an ontology and a graph that the people can work with. And one of the things, and you talked a minute ago about that, I love that analogy of the inadequacy of the desktop data integration, and you talked about the example of finding the person. Also, we talked a couple of days ago about once you’ve hired that person, the employee onboarding, that’s kind of a thing that’s like, I don’t know, every organization does it and every place I’ve been, it hasn’t been as smooth as it could have been. But can you talk about, or maybe whether it’s employee onboarding or some other kind of task, how that map can facilitate those activities?
Andreas:
Sure. Actually it’s about describing, let’s say each of the relevant business object in such a use case in a way that it gets connectors attached to it, which allows them to connect to other pieces, other business objects. Sounds abstract, but what do I mean by that? So if a person getting onboarded or working for a company, having a really rich description of skills and competence is available, let’s say, which are based on a controlled vocabulary, on a taxonomy, on an ontology, which is used also to classify other business objects, let’s say projects or products, and you can instantly connect those. So it’s very precise in a way you can connect people to projects and products and maybe also to use cases, to suppliers, to standards, to certificates, et cetera, et cetera. So if you have a domain knowledge model in place, which is going to be used as the basis to classify and annotate all the data of your data silos, you can automatically connect those to each other.
Andreas:
So you can call that a recommender system, a recommender service, which is able to understand which of the business objects belong together. And that’s not just because, let’s say person A is tagged with let’s say skill B, and project one is also tagged with the same, it goes beyond that. So here we start making use of inference mechanism, inference tagging as we call it. That means we know if a person has skill A, then we can infer from that that it automatically, for instance, can talk about as an expert while talking about a particular ISO standard, let’s say. And this is what you also just know or could do if you know about this ISO standard. And that’s typical nowhere written in the documents, it’s just background knowledge. So it’s a missing piece. So this domain knowledge model allows to do more of this kind of inference mechanisms automatically.
Andreas:
But this is in reality how we do desktop data integration because at the end when we pull together data from different silos, we look at it and try to make a holistic view of it. So we want to understand how things are connected. And here our background knowledge kicks in and I said before, you just need to start digitize that. And I mean we invest so much money in huge data repositories and content repositories, and we miss out to do the last percent. It’s just 1% of the entire knowledge graph. 1% of the entire semantic layer is the domain knowledge model.
Andreas:
When you look at the quantity, it’s probably a bit of more work than the rest because the rest you can automate. And I agree, it’s hard to get the right people into those working groups where such domain knowledge models have to be created and maintained, but it’s in reality, it’s just a matter of management should make the decision and set that up and say, “Hey, this is the core piece of our digitalization strategy, of our AI strategy.” It’s a core piece. It’s the centerpiece. So if you don’t have a domain knowledge model, you cannot do the rest on a good quality level.
Larry:
Exactly. That’s one of the things I want to talk about a little bit more is the relationship between knowledge graphs and LLMs. But one of them is that that used to be a really incredibly difficult task and a lot of the grunt work, just the digging that needs to happen to build a knowledge graph and LLM can help with that. Can you talk a little bit more about that? You’ve probably seen this how LLMs have accelerated knowledge graph development.
Andreas:
Yeah, totally. At the very first day, we all were a bit shocked in the graph community because initially the first thought was, okay, now LLMs take over completely. On the second day we already saw, no, they don’t. They actually do a really great job in bubbling up all the unstructured data in enterprises. So companies now far more aware of the value of unstructured data than there were before the LLMs came, which is great because there’s so much of interesting stuff in there, but you need to handle it in a more structured way. And this is exactly again, where domain knowledge models kick in. They can connect those pieces between the data and the content of what enterprises have. And even the people, I always have seen that there’s the content people and the data people in enterprises, they don’t know each other, they don’t talk to each other, they have different mindsets.
Andreas:
And now things come together and eventually it’s all about knowledge. It doesn’t matter if it’s data or content, it’s just about knowledge. And that’s what really helps us as human beings to hopefully have more fun with our work again pretty soon. And I think it’s the combination of LLMs in the semantic layer. But let me just make clear the knowledge graph, the semantic layer consists of two elements. It’s on the one side, the domain knowledge models, typically it’s not just one, it’s a couple of interconnected and mapped to each other. Could be taxonomies, ontologies, that’s the domain knowledge model part. That’s still, of course, and will always be done by subject meta experts together with knowledge engineers. And there’s still some manual work to be done, whereas we already have started to use LLMs also in this step to accelerate that. That’s also fun to use a tool like PoolParty, which has a LLM assistant now under the hood to help knowledge engineers and SMEs to get fast to a well-working domain knowledge model.
Andreas:
This is part one. And the second part of the semantic layer is the fully-fledged enterprise knowledge graph. The different types of that could be a product graph, a content graph, et cetera. But this enterprise knowledge graph is typically automatically generated through the usage of the domain knowledge models together with other, let’s say waste or extract and transform data from already existing data repositories. So the majority of the semantic layer is always done automatically. And again, I think it’s about maybe one or 2% of the triples or the connections you need to create manually and the rest you can do automatically. So it’s a bit a myth to say we will never be able to create a knowledge graph. There’s so many enterprises in the meantime, which has showed it’s possible. So it’s feasible. You just need to know the methodology and probably use good tools also. And the people knowledgeable about the creation of it, they’re around. So it’s a huge community in the meantime.
Larry:
And you’re reminding me of that truism that in all of this stuff, it’s about people, processes and technology, in that order that it’s mostly about people, some about processes, and the technology is just a critical enabler of course, but less important than the people, than the people knowledge. You’ve mentioned a couple of times now, that domain knowledge model, it’s such, like you said, it’s maybe 1% of your entire graph, but it’s doing, I don’t know, 80% of the work or something like that. It’s like the Pareto distribution on steroids. And I guess I think a lot of people, you mentioned earlier that the initial thought is like, “Oh great, LLMs can just do everything.” Tell me, do you have some success stories of enterprises realizing, oh, we really need this and it’s worth investing a disproportionate amount of our money into that 1% of the graph that’s going to give us all these new results. How does that conversation go?
Andreas:
Yeah, that’s also a great question. At the moment, it’s like that, there’s still the content people more aware of how important the fully fledged semantic layer would be. So they have, I would say a different definition of a semantic layer. Even that the data, people are still more on the mapping side of things. They’re not thinking of the value of enriching the existing data with the domain knowledge model. So it’s the content people introducing taxonomy, ontology discussions into the architectural, let’s say strategy decision-making processes. And here on the other side, there are, I would say AI strategies around, understanding completely the importance of rock architecture, that’s clear. But they currently stick rather with the idea of vector databases will do the magic, which is frequently not the case. And then the content people come and say, “Hey, by the way, have you ever thought about reusing our taxonomies ontologies to enrich your rock architecture? We’ve built the knowledge graph.”
Andreas:
Maybe they even transform content repositories, building a content graph, enriching them with semantic metadata, adding the domain knowledge model next to it, typically maybe because of a better search feature, et cetera. And then the AI people, the data people say, “Oh, really? We have a knowledge graph? I didn’t know that.” But still people are hesitant. Some data engineers and data scientists are hesitant to use that. They say, “No, we can do everything with algorithms and LLMs and back to database and bannings,” blah, blah, blah.
Andreas:
And then typically they try it out and fail, business at the end still is in, let’s say in the driver’s seat. Fortunately at the end of the journey they say, no, it’s not good enough. And then people start talking to each other and it truly observe, it’s the first time in my career that it frequently happens that the content and the data people start talking to each other. And they have never done that before, which is really strange. And it’s also a result of the need to finally execute on the AI promise. So Gen AI has changed the whole environment from my perspective at the moment at least to the very good.
Larry:
Yeah, I love that. And coming out of a content heritage, I love that content people are driving a lot of that. But it’s also, I’ve got a number of anecdotes. I do another podcast about content and AI and there’s a lot of content designers who’ve immediately made a lot of friends in the data science world with their language knowledge as they start collaborating around stuff. But hey, Andreas, I can’t believe it, we’re coming up close to time, but before we wrap up, is there anything last that you want to make sure we share or anything you want to revisit from the conversation?
Andreas:
Yeah, I think another aspect, what a semantic layer really does nicely brings in is ways that vocabularies get mapped to each other. In other words, we are living in an era where I think we see a lot of merger acquisition, consolidation, general more complex, let’s say cooperation models in many industries. So we are in strong, need to map out different vocabularies. And by that I also mean the vocabularies used by human beings and not just by machines. And it’s another interesting feature. The semantic layer really can also help to interpret data being produced by human beings in a way that it can be instantly understood by others. So I’m not talking about we have that somebody is going to impose an ontology on the entire company and everybody has started to use that. I really talk about mapping between different ways to look at things. And by that it’s a very important tool also to develop governance around data, which allows to fuse some degree of decentralized structures with centralized structures.
Andreas:
And I’m a big fan of these type of things. You could call it “glocalization” if you want. Yeah, it should be a mixture of globally and locally thinking organizations because those have to come together. We see it even on a political level now. So those are the two, let’s say big philosophies. We need to bring those together into one. And I think a semantic layer is even needed on the very, very big scale on a political level. So I would suggest to introduce semantic layers also on our, let’s say democratic systems on top of them, and let the different parties interact with each other in such a way and they will start to understand each other far better. Even religious discussions would ease if we would introduce a semantic layer because we start to see, hey, we are talking about the same concepts, just use different terms, use different contexts around that. But it’s the same thing. It’s the same thing we are talking about. I think humans are far closer to each other than we currently think. We are separated by a lot of words and terminology only.
Larry:
Yeah, there’s so much in there. You just introduced 10 more podcast topics, but I love the idea of glocalization. That’s brilliant. And also, we talked yesterday, and this could be a whole other conversation later on about the enterprises, is the global operating system and the role of the semantic layer at that level. But I guess in the meantime, just even supply chain stuff, it’s probably going to be very helpful. Well, thanks. One very last thing, Andreas, if folks want to connect with you or follow you online, what’s the best place to find you?
Andreas:
Quite active on LinkedIn. So I think it’s a great platform, so please try to get in touch with me via LinkedIn. First of all, I’m quite responsive there. If you write, drop me a message, there’s also a LinkedIn group I would recommend, which is the Knowledge Graph and LLM LinkedIn group. So there’s all this good stuff around and yeah, that’s the best way moving forward. It’s always webinars we do, there’s a first graph-wise webinar next week. Can also recommend to watch that if interesting for you. So semantically here is really the big topic in the next month to come.
Larry:
Excellent. I’ll put all those links in the show notes as well. Well, thanks so much Andreas. Always great to talk.
Andreas:
Thank you, Larry. It was fantastic. Thank you very much for this possibility to talk about this topic.