Lulit Tesfaye: Semantic Architectures for the AI Era – Episode 52

photo of Lulit Tesfaye, expert on semantic architectures for AI and knowledge
Lulit Tesfaye

Generative AI has prompted a flurry of experimentation in enterprises, resulting in a parade of failed pilots, unproven PoCs, and unrealized return on investment.

Lulit Tesfaye helps companies improve and optimize their AI capabilities by adding semantics to their enterprise architectures. This puts their precious knowledge assets in a context that machines can actually work with. A crucial part of that context is the interoperability that standards like RDF enable.

We talked about:

  • her work at Enterprise Knowledge as Partner & Vice President, Knowledge, Data, and AI Solutions
  • the publication of Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence, which she co-authored with colleagues at EK
  • her professional path from software engineering and applications development to her current role
  • how generative AI has increased data teams’ workloads, leading to numerous pilots and PoCs, which has exposed the need in enterprises for semantic context
  • her definition of a semantic layer
  • “knowledge assets,” a foundational concept that she uses to talk about data, content, information, expertise, and any other assets to which can append metadata
  • how she and colleagues see the distinction between the two types of semantic layers: data and analytics layers and semantically modeled layers, and the benefits of the latter
  • the framework she uses to convey the benefits of a semantic layer to executives
  • the swing of the enterprise architecture pendulum back towards monoliths, as enterprise solutions providers like ServiceNow acquire semantic tool companies
  • the importance of interoperability standards like RDF in semantic systems
  • examples of use cases that are best suited to labeled property graphs and RDF knowledge graphs and how she sees the debate about which to use going away soon
  • two common failure points in semantic layer projects:
    1. starting with tools, not competency questions
    2. insufficient skillsets and lack of leadership support
  • the use of multiple levels in a semantic layer in large, complex organizations
  • emerging trends she sees in semantic architectures:
    1. the ubiquity of AI
    2. the need to design for machines
    3. the need for investment in a context layer
    4. the disappearance of the artificial lines between structured and unstructured knowledge assets
    5. the evolution of enterprise operating models
  • the importance of staying focused on our knowledge assets and adopting solutions that are based on standards and enable interoperability

Lulit’s bio

Lulit Tesfaye is a Partner and the Vice President at Enterprise Knowledge, LLC., the largest global consultancy dedicated to knowledge and data management. Lulit brings over 15 years of experience leading global initiatives, specializing in information and data management solutions and integrations. Lulit is most recently focused on applying practical knowledge management, data governance, and semantic data foundations to optimize organizational information assets – providing the structure, context, and explainability needed to support responsible and effective enterprise AI.

Connect with Lulit online

Video

Here’s the video version of our conversation:

Podcast intro transcript

This is the Knowledge Graph Insights podcast, episode number 52. Knowledge management and data engineering were already changing even before ChatGPT arrived. Now, with the proliferation of AI, the boundaries are blurring between structured and unstructured
data and other knowledge assets. As a partner at a prominent knowledge consultancy, Lulit tess fai sees every day the implications of this change, chief among them the need for semantic layers built for interoperability on a solid foundation of industry standards.

Interview transcript

Larry:
Hi everyone. Welcome to episode number 52 of the Knowledge Graph Insights Podcast. I am really delighted today to welcome to the show Lulit Tesfaye. Lulit is currently the Partner & Vice President, Knowledge, Data, and AI Solutions at Enterprise Knowledge, the big famous consulting agency in the DC area. She’s also the co-author of the recent book, or maybe forthcoming, I’m not sure, but I’ve seen it online so I know it’s coming. Bridging Knowledge, Data, and AI: Harnessing the Semantic Layer Framework to Drive Intelligence, which she wrote with her co-authors, her colleagues at EK. Well, welcome, Lulit. Tell the folks a little bit more about what you’re doing these days.

Lulit:
Thank you, Larry. Thanks for having me. I’m happy to be here and have been looking forward to this conversation. That’s a very kind introduction, so I appreciate that. So, maybe a good way to start would be a little bit from the journey over the, without dating myself, I’ll say I’ve been in this space close to 20 years at this point and it’s been not a direct or linear path given how things are evolving, but the short version of it is started in working within software engineering and applications for data and information in that space, which led to me joining Enterprise Knowledge after I met the co-founders about a decade ago.

Lulit:
And then our focus was in knowledge management primarily, which traditionally is considered as learning and development or training or relegated to librarians, so to speak, information architecture and so forth. But over the past five years specifically, we’ve been seeing a big shift towards the silos in a way dissolving between the different types of these groups within organizations and where information or knowledge sits. So, we think in terms of structured and unstructured data, unstructured being content, PDFs, and all the things that you have in your files and folders, whereas structured data has been traditionally for data and analytics teams.

Lulit:
And the biggest swing that we have seen thanks to especially generative AI, is that unstructured data, which is really 80%, 90% of your organizations, I’ll call it knowledge asset, came into the purview of data teams, data and analytics teams because of generative AI. And that was what led to this big focus on knowledge management for data as well as handling structured data. So, I think we’ve all gone through the pilot phases for POCs and pilots for AI with the pilot purgatory, all the failed efforts. That has been the journey, especially over the past, I would say five years is working with organizations that have been experimenting and embracing AI or today trying to wrangle what we call the AI pilot sprawl, scenario that exists.

Lulit:
So, to give you an example, one of our clients recently said we just cataloged over 150 pilots, AI pilots, and we haven’t seen the return on investment as we should. Good experimentation, good learning, but we need to put an end to it and have a strategic path. So, this is where we are today. All the pilots and POCs for many different reasons that I can get to either have failed or have stalled. And at the same time, there are a lot of things in production that I would commend the organizations leading that. But the biggest failure and a big part of our journey has been the backbone, the core, the semantics or the context for AI as the lens that AI sees your organization, that has been the core of it. Just to answer maybe in a long way what we’ve been working on over the past few years.

Larry:
Yeah, I think the first time I met you, you did a talk about the semantic layer at that data summit in London, the one that Henry Stewart Events puts together. And I’ve been following your work about semantic layers since then. And what you just talked about, that’s kind of the classic sort of framework that people use to connect this, like the data stuff you’re talking about, the KM. But I’m really curious, in the knowledge graph world, everybody talks about knowledge engineering. And what you just described seems like the setting the stage for a big shift from knowledge management to knowledge engineering. Does that make sense? And if so, how does it fit in how you see with things right now?

Lulit:
It does. And I think maybe the best way to answer that would be to take a step back and really define what we mean by semantic solutions or specifically a semantic layer, because I think there’s a lot of confusion in this space. It’s funny, we start the book with it’s all about semantics, right? It is truly about semantics. And the key definition that I would make here distinction is semantic layer is not something new. In fact, we’re just chatting. It celebrated its 25th year anniversary this year. It’s come a long way I think. And what we mean by semantic layer is we are talking about the shift from the focus from the physical data itself to the data about the data, so the meta aspect of it.

Lulit:
So, this means abstracting the physical data with definitions, business glossaries, metadata, the aboutness of the data, taxonomies to control some of these vocabularies that are very unique to your organization, what fits under our product, list of product, list of services, ontologies to be able to relate those concepts, these vocabularies that you have to your real world entities. This customer owns this product, has bought this product and needs this type of marketing material or training material. How do you connect to these pieces? That’s where ontology allows you to explicitly make that machine-readable, places things.

Lulit:
And then when you apply this model, this semantic model on your assets, and I am collectively going to call going forward data content and information and expertise, anything you can append metadata to, knowledge assets, then you have your graph, your Knowledge Graph. So, this is what we mean when we talk about a semantic layer and the components that sit within semantic models. There is another definition of a semantic layer, which comes from the data side, data and analytics, that we refer to as the metrics layer because the core purpose or use of that semantic layer is really to define metrics and to connect them.

Lulit:
So, the biggest difference I would say between these two types of semantic layers, we actually have an article on our knowledge database about the tails of the two semantic layers where the metrics layer is primarily focused on data definitions, not object and concept definitions. And that’s a very important distinction. The moment you take a concept or an object customer product revenue outside of a table or your SQL table, it loses its meaning. And that’s really the key thing is we care about the things inside the cells of your spreadsheets, the objects and concepts and things, as opposed to the table IDs. And that’s the key distinction between the metrics layer and the semantic layer that I would make.

Larry:
Yeah, I love that distinction because I think a lot of people, there’s so much fuss about this in the news these days. I think there will be a lot of new people coming to this. And I think that analytics perspective on the semantic layer is way more common. What you just said makes perfect sense, but have you faced the challenge in your consulting practice of helping people make the transition from that? It almost sounds like things not strings, what you just described, the Google way of saying it. Does that make sense? Yeah.

Lulit:
It does. That’s exactly it. And I alluded to this earlier, especially with the adoption of LLMs and genAI and agentic AI, now there’s a change and shift in perspective from the data side as well because it’s the value of being able to make this business meaning and concept or being able to inject what does customer mean in this instance and giving it… not having to agree on one definition of customer because no one will agree on one definition of customer, but you would want to be able to make your knowledge about why this customer’s definition of this customer is different from the other, machine-readable in a standardized way. That’s what we are calling the context engine, the context for AI.

Lulit:
And because of that, I would say just in the past, especially three years, we are seeing a lot of investment coming from the data side, the data and AI engineering teams in semantic and the semantic layer and semantic components in semantic models. So, to give you an example, for a large retailer, they have over 40,000 stores across different parts of the world and they want to be able to do reporting to leadership about the store health or the state of the company, the state of a store, and there are multiple things that then need to be calculated within that. So, metrics, right?

Lulit:
This is a really nice neat story of how semantics supports the metric layer and where the data team is investing in the semantic layer to standardize the objects and concepts within their metric layer. So, a good example of this is how do they calculate revenue? One of their measures of store health is revenue or system or operating system up or downtime. They have multiple things that goes into consideration, but each store has their own version of it. So, when all of this is aggregated, it will take them about five to six months. The report dashboard will be wrong because this analytics team in this region did a certain way, the other did it a certain way. They couldn’t trust these decisions.

Lulit:
And where the semantic layer came into play is in having these global, well, very tangibly metadata and glossary standardization about when we say store health, what is a part of that, what is profit, what is revenue? Just making that machine-readable for their analytics team, for their AI team.

Lulit:
But larger than that, the biggest value for them was being able to connect using a graph, how each domain relates to one another, how revenue impacts products or how products are related to revenue, how operations is related to revenue, which is a unique type of relationship and making that machine-readable using a graph was one of the, we call it this delivering semantic layer as a product for them for the metrics calculations and for their team, which significantly reduced the hunting and gathering they have to do to report out from the five to six months to I would say four or five weeks.

Larry:
That’s so interesting. You probably know Veronika Heimsbakk. She was on a few episodes ago and she took something in Norway she took with a similar architecture from months to minutes, but down to weeks is still quite good. But I love that notion of the semantic layer as a product. Is that sort of how you articulate it now when you’re trying to … Because it’s not like a thing, it’s like a concept, a framework. How do you position it when you’re talking to… especially to an executive who might not care at all about the technical details under the hood?

Lulit:
Yeah, it is definitely not a product or a tool that would start there because that’s where people always like to go and fix the problem. There’s not a single product currently that is the semantic layer. It is a framework that allows you to create your semantic layer for your organization. There are some standardized things that by industry that you can use, like the financial services industry has very mature models, semantic models that you can adapt. Schema.org has some model, that version. The pharmaceutical space has some of that.

Lulit:
Those are all great starting places we always use, but when it comes to your organization, going back to revenue for that retailer is going to be very different for a coffee company versus for a pharmaceutical company, even if they both are revenue. So, you need a way, you need a framework to define that for your organization. So, the way I would describe it for an executive is think of it as this conceptual model of how you see your world and from the top down and using that knowledge that you have and making that machine-readable so that you can apply it from the bottom up to your knowledge assets and how they are related.

Lulit:
You need tools to manage this so that you can govern it and manage it and have good data quality checks and so forth. So, graph database, taxonomy, ontology, metadata manager, data catalogs, although now we’re seeing a big shift within the industry where these vendors, the large enterprise tools, solutions like ServiceNow, Salesforce, all of these are acquiring these solutions to make it a part of their suite. At the end of the day, that’s actually a whole topic where the world is going to the monolith era. The pendulum is swinging back again to the early 2000s, but needless to say, it’s showing how there is big investment in the semantic layer and having this capability.

Larry:
Yeah, no, I’ve done a lot of work with decoupled architectures and those kind of modular service oriented systems make sense, but I see this all in a lot of places, this kind of regression, or return, to the monolith. What are the implications of that when you’re building a semantic layer? Because as you just mentioned, well, I inferred from what you just said that almost every enterprise will have a unique, just everything from vocabularies to taxonomies, everything they do will be unique in some way. What are the kind of practical measures that you take in advising people about these architectures to make sure you fit it right to your organization?

Lulit:
Yeah. The number one thing I would say is interoperability. The promise of the semantic web is all about standards and being able to have to protect your data from specific solutions of today that may not exist tomorrow. So, the idea of the shift from application-centric development, application-centric organization, organizing around your CRM or your CMS or your data lake to the semantic model to the data-centric or the us asset-centric type of model. So, the biggest thing focus should be interoperability. What that would mean in reality is any organization that is a certain size would have at least at a minimum five platforms, if not more. The larger companies have 20 plus just for one domain.

Lulit:
So, if you want to be able to the good old traditional problem of, I can’t find anything, where do we store this? How do we connect this information? You need to be able to get it out of these specific applications, not physically, but at least at that metadata level, at the framework level to know that it exists. So, in order to do that, you need some way of having this system talk to this other system. And that’s really where standards, a way of building these semantic layers using RDF, the semantic web standards, resource description framework or OWL or SHACL for governance. There are well-known, well, mature ways of architecting and developing this solution.

Lulit:
That said, there are other LPG type of the Neo4j type of solutions that also could plug into this from a graph standpoint. We can come back to that later, but I would emphasize on building your semantic solutions on standards is going to pay off, not only for the issues or the problems you’re looking to solve for today, but also the problems we are not anticipating about where the vendor space would end up in the next couple of years.

Larry:
Right. I mean, I think you can lump everything you just said under vendor lock-in, which seems like a good way to get … But what you just said, I just want to do one little quick sidebar. The two episodes ago I had on Ora Lassila and Adrian Gschwend talking about RDF 1.2, and I’m wondering that if your advice to consulting clients about that choice of a labeled property graph versus an RDF-based Knowledge Graph, will that change much with the introduction of RDF 1.2?

Lulit:
It remains to be seen, I would say. Ora is a friend and a great thought leader in this space and has been actually investing a lot of his time in making the bridge between RDF and LPG. I don’t know if you remember five years ago or so there was RDF Star.

Larry:
Yes, exactly. Yeah.

Lulit:
Yeah. So, now I think it’s still developing, it’s in the process. But at the end of the day, the real world application of this is there shouldn’t be a choice. In many of our implementations, especially for large organizations, we have both because they serve very different use cases. To give you an example, for a financial services firm, the RDF defines their global domains, upper ontologies and standards definitions and relationships, whereas LPG is pulling some of those standard domains that they need to disambiguate and resolve entities within their graph but are using their analytics graph in LPG, which is highly performant and fit for purpose for those types of investigative analytics types of use cases in fraud detection and so forth.

Lulit:
So, that’s usually the scenario we see is that many organizations have both in retail is another one that we see that hop from RDF defines the standards, the facts, the relationships, whereas LPG takes that and does the analytics and the processing on the ground. So, I don’t think it’s a choice. I would love to see we have some libraries that we’ve built to do this half business hop, especially with agents now from LPG to RDF, RDF to LPG. So, I’m see promised that that’s becoming less and less challenging and is going to disappear as a debate, I think, where it’s no longer which way do we go.

Larry:
There wasn’t maybe not consensus, but there was a lot of thoughts that echoed what you just said at the Knowledge Graph conference this year that’s kind of how it feels. And one thing I’m hoping that people will tune into this podcast who are curious about a semantic layer and want to do it right. And I’ve heard you talk a few times about the cautionary tales about how it can go wrong. Can you talk a little bit about that just for folks who are thinking about it to make sure they get it right the first time?

Lulit:
That’s a great question. So, we have at this point over the last decade, I would say we’ve done over 200 engagements in this space. And if I was to do a wide scan of at very many different maturity levels, some organizations are nascent, some are more mature and advanced. I would say the common thread that I see and I would pull for failed efforts is it starts with a tool that’s a technology of sorts. And from a technical standpoint, this almost always fails because the whole point of semantics is bringing the business lens to your knowledge assets and to your data. You can always start from the top down or the bottom up through your corpus. That’s from a delivery standpoint.

Lulit:
There are different ways to slice this. However, the lens, the biggest failure rates that we see is when you start from a very technical problem. So, we find that you need to always take it up a level to what we call your competency questions. What business questions are you looking to answer? Because at the end of the day, the core value of a semantic solution, semantic layer is making your business knowledge machine-readable, contextualized and standardized across your knowledge assets. So, that’s the number one thing. So, that pocket of the organization that is very technical. That would be top of my list, which in a way leads into the second piece, which is around skillsets and leadership support.

Lulit:
Over the past few years, we see a lot of efforts fail because they were focused on a given technology maybe not executed fully. So, there’s the sentiment that it doesn’t work. It’s a lot of effort. Traditional engineering does not have the skillsets in-house, although now it’s very promising to see a lot of trainings in the semantic space. We have our Knowledge Graph university as well that is getting a lot of traction. So, there is now interest from the traditional engineers and data engineering sides to learn semantics as well. So, that I would see would be diminishing, especially with AI and assisted with agents, I think that’s going to be improving the adaption rate as well.

Lulit:
But that traditionally, especially over the past few years, has been a barrier to entry and failed one of the reasons why these efforts fail.

Larry:
Nice. Yeah, both of those are interesting. Again, at the Knowledge Graph conference this year, it came up that there’s more interest from different engineers. It’s been kind of, not insular, but it’s been a lot of us talking to each other and it was really gratifying actually to see how many new people were at the conference. Yeah, that’s really interesting. Hey, one thing I wanted to also ask, I saw you on a panel discussion online a couple months ago, I think, where they were talking about the universal semantic layer. And I wonder, is that a useful distinction to draw or was that just the topic of that conversation?

Lulit:
I would say it’s more the latter. The one thing that I would call out that distinction that it makes is it’s really hard to create, to define, especially for large organizations like one enterprise semantic layer. So, what we typically see is you start going through this T-shaped approach of what are the domains that are applicable across the board regardless of where you are within that organization. We are working with a consultancy, for instance, that has 60,000 offices across the world. So, what are those core domains that are applicable to that? Think region, location, departments, services, industry, the things that are shared across the board, that is what you start defining as your universal layer.

Lulit:
And then after that, as you go down levels of the concepts, then you would want to be able to give freedom to different groups to do their own thing. Some pull from the semantic layer as a product model comes into play, some embedded within their operations. As you go down the concepts level from universal to upper ontology to domain ontology to the schema, then that freedom, you need that freedom to these specific groups, like the group level, the Wild West for this to be adapted. So, it’s kind of a design decision, an implementation approach and an implementation decision and to make it usable for large and organizations.

Larry:
It sounds like there’s emerging design patterns in these architectures. So, you just said mirrors almost exactly what Eric Little from Accenture has said and very close to what Alex Bertails at Netflix talks about. I guess it’s unique to each organization, but it does sound like there’s some high level patterns emerging that enterprise architects can fall back on because they must be … Well, I guess I don’t want to go too much deeper into it because we don’t have a lot of time left, but that AI as a driver of all of this that all of a sudden these … You mentioned, I think right at the very start how the data folks are driving a lot of this. Is it because they all of a sudden have what, four or five times as much data because all of a sudden, they’re also getting all the unstructured data?

Lulit:
Exactly, exactly. In fact, actually this starts getting into the emerging trends that we’re seeing in architecture to your point, which is one aspect of this is AI is ubiquitous in every solution that you have, even if it’s like a legacy system, ADP has ADP Assist. They’re investing in that. Snowflake has chat with your data. Workday has chat with your HR, right? So, adaption is no longer an option. That’s number one. So, there’s a big driver here by itself that AI is happening. Even these solutions that you’ve had for 20 years are embracing it. So, now you have to work with it. That’s the biggest shift. The second thing is our design architecture is no longer just for humans.

Lulit:
We are also designing for machines. So, there’s a very different level of architecting that you would have to think about. Humans need intuitive ways to receive information, machines need granularity and context. So, your architecture needs to reflect that. So, that’s another big shift that we’re seeing within enterprise architecture. And then the third piece is it starts getting into this context. So, when you talk about machines, now you need some way to convey your organizational context and meaning and business definitions. And that’s where the semantic and context layer is becoming a big investment.

Lulit:
I think you may have seen this. Gartner earlier this year has said that every organization by 2030 is going to have a semantic layer and the way we define it, Gartner has also come around, which is great to see. So, that’s the third piece. And then the fourth piece, this is a big other change is this lines of artificial silos between data to your point about structured, unstructured is very quickly disappearing, which is why we’ve been talking about this concept of knowledge assets. It doesn’t matter if it’s data or structured or unstructured, AI really doesn’t care, especially, but it needs to consume it a certain way.

Lulit:
So, that is leading to the fifth emerging thing we’re seeing operating model changes, which is teams are actively … Many organizations are going through reorgs if you look around today and we have lived through three or four just in the past two years, the good type where the data team is merging with the engineering AI and KM team and now they’re a cross-functional team and under a COO or a managing director. Their organizations are actively changing their operating model, which is also leading to how we pre-manage security entitlements because in AI, typically we manage that at the group or application level. We’re now talking about the concept and data level, which semantics and context is all about.

Lulit:
So, now we need to upgrade our security and management entitlements to mirror the same. So, architecture, I know that was a lot, but if you talk about architecture, knowledge data, AI architecture, these are the factors that are driving the biggest change and we are seeing I would say three, maybe four types of architectural patterns emerging as a result of that. I have some articles that we published with our architects on our knowledge base. I can share the link so you can read it.

Larry:
Yeah, I’ll put that in the show notes for sure. No, that’s awesome. It’s like we saved the best for last. I’m going to leave this in as clickbait. I’m really curious about the engagement level in this one. I’m going to tease this to death that be sure to listen because really be sure to listen all the way to the end. Lulit save the best stuff. But yeah, we are coming up close to time, but before we wrap up, is there anything last, anything you want to revisit from the conversation or just want to make sure we share before we close?

Lulit:
I would say I think I’ll go back to the world as we know it is very fastly changing, right? And these are a lot of the drivers that are leading to how we think about our data, our content, our knowledge, and bringing it all together. Well, also a nod to the title of the book is really something that we’ve been championing for the past 10 years as a company, but also as practitioners throughout and it’s really neat to see it catching on and with that is also coming a lot of noise in this space. I would emphasize always ask and look for standards, interoperability as new tools and new solutions are emerging, which is really neat to see.

Lulit:
But also I think it’s really hard to be an executive and a decision maker in this space right now. I would imagine I would empathize. I live through it with many executives who are in a way going through decision paralysis because of a lot of this movement. But I think the core and anchor of this would be focus on your knowledge assets, focus on abstracting your context and your meaning of your business and a way to apply that in a standardized format to your machines and applications would be the way to go.

Larry:
I love that. That echoes something. I jotted down something you said, I can’t remember where, but you said “semantics isn’t just a layer, it’s the lens through which AI sees your world” and what you just said completely supports that. Well, thanks so much, Lulit. Oh, one very last thing. If folks want to connect with you online, what’s the best way to find you?

Lulit:
LinkedIn, my full name, I should be findable. I can also share some information in my email as well.

Larry:
Okay, great. I’ll put that in the show notes as well. Well, thank you so much, Lulit. This was an awesome conversation.

Lulit:
Thank you, Larry. It was really fun. I’m glad we got to do this.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top