Jacobus Geluk: Use-Case Trees for the Data-Product Marketplace – Episode 26

photo of Jacobus Geluk, expert on building use-case trees for the data-product marketplace
Jacobus Geluk

The arrival of AI agents creates urgency around the need to guide and govern them.

Drawing on his 15-year history in building reliable AI solutions for banks and other enterprises, Jacobus Geluk sees a standards-based data-product marketplace as the key to creating the thriving data economy that will enable AI agents to succeed at scale.

Jacobus launched the effort to create the DPROD data-product description specification, creating the supply side of the data market. He’s now forming a working group to document the demand side, a “use-case tree” specification to articulate the business needs that data products address.

We talked about:

  • his work as CEO at Agnos.ai, an enterprise knowledge graph and AI consultancy
  • the working group he founded in 2023 which resulted in the DPROD specification to describe data products
  • an overview of the data-product marketplace and the data economy
  • the need to account for the demand side of the data marketplace
  • the intent of his current work on to address the disconnect between tech activities and business use cases
  • how the capabilities of LLMs and knowledge graphs complement each other
  • the origins of his “use-case tree” model in a huge banking enterprise knowledge graph he built ten years ago
  • how use case trees improve LLM-driven multi-agent architectures
  • some examples of the persona-driven, tech-agnostic solutions in agent architectures that use-case trees support
  • the importance of constraining LLM action with a control layer that governs agent activities, accounting for security, data sourcing, and issues like data lineage and provenance
  • the new Use Case Tree Work Group he is forming
  • the paradox in the semantic technology industry now of a lack of standards in a field with its roots in W3C standards

Jacobus’ bio

Jacobus Geluk is a Dutch Semantic Technology Architect and CEO of agnos.ai, a UK-based consulting firm with a global team of experts specializing in GraphAI — the combination of Enterprise Knowledge Graphs (EKG) with Generative AI (GenAI). Jacobus has over 20 years of experience in data management and semantic technologies, previously serving as a Senior Data Architect at Bloomberg and Fellow Architect at BNY Mellon, where he led the first large-scale production EKG in the financial industry.

As a founding member and current co-chair of the Enterprise Knowledge Graph Forum (EKGF), Jacobus initiated the Data Product Workgroup, which developed the Data Product Ontology (DPROD) — a proposed OMG standard for consistent data product management across platforms. Jacobus can claim to have coined the term “Enterprise Knowledge Graph (EKG)” more than 10 years ago, and his work has been instrumental in advancing semantic technologies in financial services and other information-intensive industries.

Connect with Jacobus online

Resources mentioned in this podcast

Video

Here’s the video version of our conversation:

Podcast intro transcript

This is the Knowledge Graph Insights podcast, episode number 26. In an AI landscape that will soon include huge groups of independent software agents acting on behalf of humans, we’ll need solid mechanisms to guide the actions of those agents. Jacobus Geluk looks at this situation from the perspective of the data economy, specifically the data-products marketplace. He helped develop the DPROD specification that describes data products and is now focused on developing use-case trees that describe the business needs that they address.

Interview transcript

Larry:
Okay. Hi everyone. Welcome to episode number 26 of the Knowledge Graph Insights podcast. I am really happy today to welcome to the show, Jacobus Geluk. Sorry, I try to speak Dutch, do my best. His last name means happiness in Dutch, which you’ll see a lovely bit of serendipity here. Jacobus is the CEO at Agnos.ai, which is a really prominent enterprise knowledge graph consultancy based in London. So welcome Jacobus, tell the folks a little bit more about what you’re doing these days.

Jacobus:
Oh, thank you very much, Larry, for the opportunity. Well, we are a small, let’s say, boutique consulting company, but focusing on the enterprise knowledge graph in combination with gen AI, LLM, of course. I think that’s a match made in Heaven, these two technologies, they need each other and we jump on that completely all in because I think that’s the future. And yeah, would love to talk about this topic-

Larry:
Yeah, well there’s one topic in particular we’ve been talking about that I really want to air out today is the notion of both of those, all of AI, it runs on data and there’s this emerging thing, the data product that is like, I don’t know how clearly articulated it is, but you’ve worked on the DPROD initiative working on articulating an ontology for what a data product is. So you know as much about it as anybody. So what I’d love to do is just talk about, well first, what is a data product? For folks who aren’t familiar or discovering this for the first time, how would you describe a data product?

Jacobus:
Yeah, well, we started talking about this, I set this up in September 2023, and I started talking to Tony Seale, who’s a very famous, he’s the The Knowledge Graph Guy. You’ll find him on LinkedIn. He writes fantastic articles. And at that time, he worked on the largest knowledge graph project in the financial industry, I think at the moment, which is one of the largest which was at UBS. And I asked him to become the chair of a new work group that they wanted to set up in the context of the Enterprise Knowledge Graph Forum, which is part of the Object Management Group, which is a standards organization like W3C. And to my surprise, almost it went so well that within a year we basically hammered out this standard, or it’s now an official OMG standard called DPROD that stands for the data product ontology.

Jacobus:
And we work with many, many people mostly from banks like JP Morgan and some other people from UBS, Credit Suisse was involved, London Stock Exchange Group, Bloomberg, you name it. Like a whole range of people, Amazon, British Telecom. So there’s multiple different types of people working on it: specialists, data architects, et cetera. And the idea was we want to see the world moves towards data marketplaces, like large companies like London Stock Exchange Group for example, or Azure, Google, et cetera. They host data or they sell data. They are data vendors in a sense. So if you sell something, what is your product? Your data is a product or access to that data is a product. So rather than modeling things in terms of data sets, et cetera, which is what we did so far, that is the state-of-the-art data, cataloging, et cetera, using a standard called DCAT that is basically the data catalog standard.

Jacobus:
So we thought, “Okay, let’s elevate that a little bit higher.” We don’t want to think in terms of pure data sets anymore. That’s kind of more like an internal thing. The customer doesn’t really care. The customer just wants to know, “What is your product, how can I access it, how can I buy it? What are the terms and conditions? What is the purpose? What are the use cases that we can support,” et cetera. So I’m not claiming that we have all of that, but we have at least created a one step up from DCAT, literally an extension to DCAT that basically says, “Okay, now we define what the data product is, what are the inputs and the outputs. Input ports, output ports, basically how does it fit into a larger supply chain of data?” And that is already a step forward and companies are using it.

Jacobus:
Like there’s already several companies that have told us that they are actually using it like JP Morgan and London Stock Exchange Group, UBS, they are all using it. I’m not sure if it’s already in production or not, but they worked with us to create this and there’s apparently a need to model the world in terms of your data products.

Jacobus:
But my own story to that is, okay, you have products, but you would say instead of a data marketplace, you could also say it’s a data economy. Basically you want to look at the world as a data economy, your own enterprise, but also beyond the enterprise. It’s a larger data economy where you have supply and demand. So we have now defined all these data products on the supply side of the data economy, but what is the demand side? That there’s no real standard yet, I think, that defines what the use cases are.

Jacobus:
Like if you want to talk to the business, the business has a problem, has a budget, and they want us to build something every time. So what is that? That is a use case. That’s the term that they use. The business is talking about use cases or apps or systems or whatever you call it, but most of the time they use the term use case. So we want to create a new thing, basically defining the stuff on the demand side of that data economy called use cases. And these use cases can cannot only serve the purpose of defining very precisely what the business wants and needs, but also how that links to data products with data contracts in between et cetera.

Jacobus:
And last but not least, before I stop talking, also the LLM, the gen AI basically needs to know what is the use case I am supposed to operate in. So that’s the most exciting angle to this story basically is we want to control the LLM agents, of course, and make them useful, productive, producing high quality output in production, in mission-critical use cases. But what are those use cases? We want those use cases to be data and it’s data as code almost. But let’s stop talking about this now.

Larry:
No, this is amazing because a lot of things are clicking into place. We’ve already talked about this and I’m getting even more insights just listening to what you said right now. But that notion that articulating the demand side, and like you said, it’s always use cases and then instances of those use cases would be like anything from an application generally to this specific kind of application. So I’m starting to get my head around what you mean. So the DPROD spec, that’s sort of like an ontological description of what’s in a product, right? And so you’re kind of building, it’s like a complement to the DPROD spec it sounds like. So is it mostly about the business use cases and articulating those or is there data stuff built into this demand side thing you’re picturing as well?

Jacobus:
Yeah, because if you define what business wants more and more precisely with more and more depth to it, basically the idea is we start with talking to the business about, “Okay, what is your problem? What do you want to do? Describe what your outcomes are, describe how do you call the use case? What is your budget? What is your timeline?” Budget being very important, of course. Everything is driven from the money that comes from the business anyway, so let’s not talk around that. Money is an important driver, but then we can add more and more information to it. And what traditionally happens is that people, business analysts, et cetera, they talk to the business, write it all down, and then basically throw it over the wall to IT and IT built something and six months later or a year later, they deliver something that the business barely understands anymore because things change basically.

Jacobus:
So there’s kind of a disconnect traditionally between what the business wants and what eventually gets delivered. So we want to create this continuum basically. Like everything that you ever do in IT and in data management should always be related to a use case. So you should be able to say, “Okay, I changed this thing in here and that relates to this use case,” and maybe multiple other use cases, that’s fine, but we want to have that map basically where all the use cases are there, all your data products are there, and behind those data products, we have data sets and data elements and databases, et cetera. That’s one level deeper in detail, but at the biggest picture we want to see these are the use cases and these are the data sets, the data products, and this is how they relate to each other so that you can manage your data economy. That’s the idea.

Larry:
Manage your data economy because everybody’s aware, I don’t know what, it’s the new oil or whatever, but it’s the engine that drives everything now. And so to have a way to manage your data economy, that seems really… I’m fascinated now with it’s almost the fractal of the granularity of this demand-side specification that you’re talking about because you talk about this use case and then you talked about both the specificity of the use case, like the business objectives and all those things you mentioned. “What’s the problem you’re trying to solve? What are your goals? How are you going to measure this? What’s your budget?” So many of these kinds of models seem very technical, that they’re about a technical specification for something. This seems like a pragmatic distillation of a bunch of business stuff in a product. Am I hearing that correctly or-

Jacobus:
Yeah. Okay, first of all, the LLM doesn’t know anything until we tell it. And maybe, of course, it has been trained on data sets, but they are generic data sets, not all your data is in the LLM. So you have to explain things to the LLM, you have to instruct it. So where do you get those instructions from? So what is the context basically, which is very much that use case. The context is let’s say I want an LLM to help me checking whether a director should be added to the board of this legal entity, for example. Well, usually that’s done by accountant type personas, let’s say. But an LLM can do a lot there. So we have to explain it then. Okay, what is the legal entity? What type of legal entity? Is it in dissolution state or not? Like what is the jurisdiction? There’s lots and lots of detail there that you can feed. That’s all part of that use case description where you say, “Okay, these are the concepts.”Dissolution is a concept, for example, board of directors is a concept, and these concepts have to be explained in all the possible detail as well.

Jacobus:
Usually a lot of that knowledge is not captured anywhere, it’s in the heads of people basically. So people do that themselves, specialists, accountants, et cetera. So we want to kind of take that knowledge and capture it in data and all around that use case. So the use case itself is in the middle, but then you have your outcomes, your personas, your stories, your workflows, and especially the concepts is important because concepts change per context. What if someone says, I call this a position here that could be a transaction there or let’s say there’s multiple different names for the same thing in different contexts. So we want to understand what are all these concepts per context, per use case, and then link it up in the right way to the various data elements and eventually ontologies and schemas, et cetera.

Jacobus:
But all that linkage is a very complex graph of information with all these hardcore effects basically. And hardcore effect is not something that LLM is very good at actually. So that’s why the combination is so good. LLM is good in many things in understanding language, but it’s not very good in dealing with the hardcore effects. Like Tony Seale actually draw a very nice picture in one of his LinkedIn posts with a brain with two halves. You have the logical half and emotional half. It’s like almost the LLM is the emotional half, I would say, and the knowledge graph is your logical half and they complement each other.

Larry:
Yeah, I had Tony on the podcast a while back, and we talked about… He used that and he also used the analogy of the yin and yang symbol and kind of neuro-symbolic AI is this kind of rotating, that’s at least the feeling I had, this rotating yin-yang symbol of them following each other. So that’s really interesting. And you talked at the start about your enthusiasm for integrating LLMs into these knowledge architectures. So is that sort of what drove this, that need to better understand all this, both in sort of the traditional way that an enterprise knowledge graph would solve it, but also to have it ready for LLMs to both… Is it both for training data or is it more for the hybrid AI architectures that are crafting prompts or whatever the mechanism is?

Jacobus:
No, no. I can’t say that LLM drove this because this use case thing, let me drop another term here. I call it the use case tree because basically it’s about how high does your helicopter fly, basically. Like if you’re the CEO of a company, then your helicopter is all the way at the top and you look down on your data economy, let’s say, your enterprise, for you, a use case is maybe client 360. That’s the arts use case that every CEO wants to have, “Give me the holistic view of my customer. I want to know everything about that customer. But if you then go down a little bit, then you break that down into all kinds of sub use cases. For example, KYC is one, know your customer,

Jacobus:
Customer onboarding, et cetera. And if you go down there, you see, oh, there’s a section about sanctions or very important person or very wealthy person. There’s lots of different sub areas in KYC. So you create a tree structure, a simple breakdown, basically, of your use cases. And at certain levels you might call it business capability instead of use case, and that’s fine, there’s all kinds of different terms for that, but it’s still a tree structure. And at each level of that tree structure, you can define the concepts and the outcomes, et cetera. So it’s like managing your objectives, for example, and break them down to the lowest level all the way to personal objectives perhaps. That’s one topic that you could think of. There’s all kinds of relationships between use cases as well. So if you have that whole use case tree, then you have all your contexts defined and that’s where LLM comes in.

Jacobus:
But the idea for the use case tree came up back in 2015 when I had my team in Bank of New York where we built a very large knowledge graph with billions of triples and onboarding more than 30 data sources and doing 25 different use cases was running a production. I can even claim that this is the first enterprise knowledge graph in the financial industry that worked in production. So that environment taught me you cannot just put billions of triples in a triple store and then at some point it becomes too much of a chaos. You have to partition it basically. So that’s where that idea started. We need these cases basically that points precisely for these are the triples let’s say, or this is the data in the knowledge graph that belongs to this use case and this is the owner, who owns it, that’s the business owner, who pays for it for additional work, adding more quality, et cetera.

Jacobus:
So it’s an old idea, but it turns out to be extremely useful to rein in the LLM. And I think with LLM, we are starting to work on multi-agent architectures where you kind of orchestrate the work across multiple specialist agents. So just like you do with humans, you could say, “I have a big task to do, but this piece of the task is done by that agent.” And you could run those agents in parallel, like there’s all asynchronous architecture, you name it, and most of it is IO because you’re either waiting for the OpenAI APIs, for example, and Stardog APIs or whatever APIs you use. It’s all IO basically. So you could run all these agent conversations at the same time and agents can talk to each other. You can make them talk through tool functions, you can make them talk to each other.

Jacobus:
So you could basically say, “I found something that needs to be checked by the regulator.” So there’s a regulator agent that knows everything about some regulation, “Please check whether this is compliant or not.” And each of those agents has to be instructed. It’s like a little job description for the temporary term that they are actually doing some work. The agent only exists for the duration of that job, but in the beginning, when you start it up, you have to instruct it. You are the regulator, you need to know everything about BCBS 239 regulation, we want you to operate in this way, et cetera, just like you do with humans basically. And then this is the use case that you are assigned to. You play this role as this persona. So your personas can be humans and agents or other systems. So you play this persona for this persona we have this job description, explain exactly from this is what that persona is supposed to do.

Jacobus:
And these are the stories. These are basically the function calls, the tool functions. So the APIs, so you will, but let’s call them stories because we are agnostic to whatever technology is eventually used to implement those stories. And in most cases, we don’t even have to implement it. They can just run, they will be executed by generic software. So there is zero code use cases in that sense. You can define the stories. As a regulator, I want to approve a given conclusion of another agent, let’s say. So you have the concept there, regulators, conclusion, et cetera, as all described. Okay, the agent gets that information, gets, for example, “Here’s my conclusion, can you approve this? Yes or no?”

Jacobus:
Then the agent can instruct, let’s say, the code that calls the agent for, “Please invoke this story for me,” accessing your local internal data source or data product. So reach out to the data product, get that information, factual information, and give it to me so that I can then approve this or not. So the agent is in control. The agent can say, “I want to fetch this information, that information,” et cetera, and can decide to drill down on its own. So the agent can drill down into something that it finds of interest, but then we want to rein in what that agent can do. We don’t want that agent to just willy-nilly all hallucinating fetching data that is not relevant for the context and for the topic at hand. So we have that control layer, that’s all these stories and the implementation of those stories, that can make sure that the agent is not going to just fetch data that it’s not appropriate because it would be a nightmare scenario, I think, if the agent would’ve direct access to your database endpoints.

Jacobus:
I don’t think that anyone who runs a production system wants some agent to generate some sequel statement or sparkle statement or whatever, willy-nilly, potentially not only getting inappropriate information that’s not relevant or confidential information that should not be exposed, but it could actually also bring down the system, basically. It is possible to bring down a system creating SQL statements that are just too massive basically causing production issues, et cetera. I had that situation multiple times, actually. So I don’t think it’s a-

Larry:
So that begs the question, is there a SPARQL expert agent that says, “Dude, that query is crazy. You can’t do that.”? Is that one of the agents in this architecture?

Jacobus:
Oh, yeah you have to do your proper homework. Like if you put a particular query in production, then it has to be tested and so you have… with each of those stories, we can have all the X case tests also defined. So the system tests itself in that sense. You can test, if it really takes five minutes instead of five seconds, then maybe you should look at it. There’s a problem. So you want to know what types of queries are being productionized and don’t get any surprises whilst in production basically. But it’s not just that, but it’s also the security and, for example, doing all kinds of other advanced things like lineage or provenance, like registering what happens, basically, or pricing even, like did you pay for that data product? Do you have enough budget for getting all this information? All those things…

Larry:
Yeah, let me ask you about that because one of the things I’ve done in big enterprise architectures was a system that involved offer pricing based on the customer knowledge and supply of a thing that you’re selling. And it comes together in this sort of ephemeral offer that like, yes, right now you can buy this thing for 20 bucks and here you go. When you say pricing, is that the kind of thing that would be an example of that, like the ability to know the supply, know the pricing criteria, know the customer, and then put together an appropriate offer?

Jacobus:
Yes. Yeah, or just register what your usage was and what the price is going to be. That can be settled later maybe. But there’s all kinds of scenarios there, but pricing is not really an issue inside an enterprise. But if your data economy is bigger than just your local division and your local enterprise, then you may have to, for example, you get data from Bloomberg, Bloomberg knows how to price and they have restrictions. You can only use the data in this context, in this use case, in this division. So those pricing contracts can be quite complex and I think your data product should be able to explain, “These are the terms and conditions, these are my pricing conditions,” et cetera. And you don’t want LLM agents to just get data from all kinds of places when it’s not really necessary. So yes, when it’s necessary, you can get it, but if someone just ask all kinds of questions in the prompt and it’s not their remit to even get that data or let alone let the company pay for it, then what is going to control that basically.

Jacobus:
Like you don’t want an LLM agent to go straight to the Bloomberg APIs basically and drive the price up for the company. There has to be some sort of layer in between that does this, and that layer is, I think, should be completely use case driven, where all the conditions, like in the ranges for budgets, for example, are defined and all your governance policies, et cetera. Policies need to be enforced somewhere and it’s all about context. Context is king. We need to have the context in the knowledge graph.

Larry:
Yeah, this is really super fascinating. I can’t wait to see where this goes because I was just having a nightmare the other night about agents running amok and the exact kind of scenarios you’re describing, but you’re describing sort of a knowledge-based human… The knowledge graph foundation of this combined with that use case tree structure, this seems like a really… Is this a unique architecture that you’ve crafted here or is this put together from other…

Jacobus:
Yeah, well yeah, because I’m doing projects with knowledge graphs since 2010. I started JP Morgan and worked at Bloomberg for a couple of years in New York on a large knowledge graph project, Bank of New York, various other organizations, British Telecom, et cetera. And the people that in Agnos have worked in different companies as well. So yeah, this came up over the years basically. And actually we want to now start in the OMG Enterprise Knowledge Graph Forum, we want to start a new work group and we call that the Use Case Tree Work Group. To make it very specific, we want to talk about the use case tree and not all the other topics and that will start hopefully in the next couple of weeks. Maybe let’s say that in the next month you want to start, we are preparing for mailing list, et cetera. Please let me know if you want to be on the mailing list through LinkedIn.

Larry:
Absolutely. Is that the best way to… Because I can imagine a number of people who might be interested in participating or following that.

Jacobus:
Yes. Well, yeah, LinkedIn is maybe the most practical way of doing it, yes.

Larry:
Cool. Yeah. Hey Jacobus, I can’t believe it, we’re coming up close to time. I’ll definitely include it. But hey, before we wrap up, is there anything last, anything you want to make sure you share about this new data, product and economy and use trees and stuff? Any last thing you want to share or revisit from the conversation?

Jacobus:
What I see so far in this little semantic technology industry is that all products are trying to do it their own way and there’s not much standardization. And the strange thing is that the whole semantic industry started with standards, RDF, SPARQL, Owl, SHACL, Linked Data, and a whole range of other standards. So it’s very much came up out of standards. But now we see all these vendors and they all go in their own direction. They all do their own thing. The standard, it’s almost like standards are forgotten because there’s not much progress in those standards. So I hope that, I think for something like this data economy thing where LLM agents are doing all kinds of work that is currently done by humans, freeing humans up to do more interesting work, I would say, I hope they won’t get fired. And that’s not the goal actually. It’s about getting humans to do more interesting… You don’t want to check regulations all the time, basically do something more interesting, I would say.

Jacobus:
And I think everyone will become more productive in that sense. You could create your own department or virtual employees in that sense that do your work and you just oversee it and you can therefore be much more productive. But I think for something like that to happen, we need more standards basically. So that’s why we do this in OMG. This is all volunteer work. No one gets paid, so we just dream about a little future in that sense where we can get this thing going.

Larry:
Nice. I love that. Coming back to the roots in standards of your practice, kind of revisiting that and doing it, and this sounds like a really great approach to that. So hey, one very last thing, Jacobus, if people want to follow you or stay in touch about the use case tree working group. You said LinkedIn, is there other places to follow you or?

Jacobus:
Yeah, LinkedIn. Yeah, the Agnos website, of course, Agnos.ai. In a couple of weeks we will have a new one there actually. Then we have also, if you go to ekgf.org, also an old website, I’m a little bit embarrassed about, it should be much better. Google for the words OMG and EKGF, those two words. And then you see more information about this. And there’s an old website from a previous work group, the method work group, method.ekgf.org, which describes a little bit more about these use cases and the objectives behind it.

Larry:
Excellent. Yeah, I’ll include links to all that in the show notes as well. Well, thank you so much, Jacobus. This is really fascinating stuff. I’m really glad we got to talk about it.

Jacobus:
Thank you so much for this opportunity. It’s very cool. Thank you.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top