Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Graph RAG is all the rage right now in the AI world. Paco Nathan is uniquely positioned to help the industry understand and contextualize this new technology.
Paco currently leads a knowledge graph practice at an AI startup, and he has been immersed in the AI community for more than 40 years.
His broad and deep understanding of the tech and business terrain, along with his “graph thinking” approach, provides executives and other decision makers a clear view of terrain that is often obfuscated by less experienced and knowledgeable advisors.
We talked about:
- his work building out the knowledge graph practice at Senzing, and their focus on entity resolution
- the importance of entity resolution in knowledge graph use cases like fraud detection
- the high percentage of knowledge graph projects that we never hear about because of their sensitive or proprietary nature
- his take on the concept of “graph thinking” and how he and colleagues illustrate it with a simple graph model of a medieval village
- how graphs add structure and context to our understanding of the world
- the importance of embracing complexity and the Cynefin framework in which he grounds various types of business challenges: simple, complicated, complex, and chaotic
- how to apply insights discerned from a Cynefin framing in management
- how knowledge graphs can help oranizations understand the complex environments in which they operate
- the wide range of industries and government entities that are applying knowledge graphs to concerns like supply chains, ESG, etc.
- his overview of RAG – retrieval augmented generation and graph RAG
- the wide variety of uses of the term “graph” in the current technology landscape
- Microsoft’s graph RAG which uses NetworkX inside their graph RAG library, not a graph database
- Neo4j’s approach which creates a “lexical graph” based an an NLP analysis of text
- “embedding graphs”
- ontology-based graphs
- Google’s approach to RAG, using graph neural networks
- graphs that do reasoning over LLM-created facts assertions
- “graph of thought” graphs based on chain-of-prompt thinking
- “causal graphs” that permit causal reasoning
- “graph analytics” graphs that re-rank possible answers
- the evolution of graph RAG libraries and the variety of design patterns they employ
- the shift in discovery dominance from search to recommender systems, most of which use knowledge graphs
- examples of graph RAG from LlamaIndex and LangChain, in addition to Microsoft’s graph RAG
- his prediction that we’ll see more reinforcement learning, graph tech, and advanced math capabilities like causality in addition to LLMs in AI systems
- his reflection on his efforts to advance graph thinking over the past 4 years and the current state of LLMs, graphs, graph RAG, and the open-source software community
- the need for a shift in thinking in the industry, in particular the need for cross-pollination across tech proficiencies and enterprise teams
- the “10:1 ratio for the number of graph RAG experts versus the number of people we’ve actually worked with a library”
Paco’s bio
Paco Nathan leads DevRel for the Entity Resolved Knowledge Graph practice area at Senzing.com and is a computer scientist with +40 years of tech industry experience and core expertise in data science, natural language, graph technologies, and cloud computing. He’s the author of numerous books, videos, and tutorials about these topics.
Paco advises Kurve.ai, EmergentMethods.ai, KungFu.ai, DataSpartan, and Argilla.io (acq. Hugging Face), and is lead committer for the pytextrank and kglab open source projects. Formerly: Director of Learning Group at O’Reilly Media; and Director of Community Evangelism at Databricks.
Connect with Paco online
Resources mentioned in this interview
- Connected Data London conference
- Knowledge Graph Conference
- GraphGeeks community
- REALM: Retrieval-Augmented Language Model Pre-Training, Guu, et al.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis, et al.
- NebulaGraph Launches Industry-First Graph RAG: Retrieval-Augmented Generation with LLM Based on Knowledge Graphs
- Graph Retrieval-Augmented Generation: A Survey, Peng, et al.
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 10. As enterprises and tech companies have looked to ground in factual knowledge the answers that their LLMs deliver, graph RAG architectures and products have sprung to the fore. With his deep background in Silicon Valley culture, the open-source software community, artificial intelligence practice, and knowledge graphs and semantic technology, Paco Nathan is one of the best-positioned people in the industry to help us understand the current state of graph RAG.
Interview transcript
xLarry:
Hi, everyone. Welcome to episode number 10 of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show Paco Nathan. Paco is the principal DevRel engineer for knowledge graphs at Senzing, the big company that does entity resolution for a large-scale mission-critical applications, really fancy high-end graph stuff.
So, welcome, Paco. Tell the folks a little bit more about what you’re up to these days.
Paco:
Thank you very kindly, Larry. I appreciate. Yeah, I’m over at Senzing. Actually, I was presenting a master class about Senzing integrations at the Knowledge Graph conference last time we saw each other in Manhattan and then joined the company shortly thereafter.
Paco:
I’m building out the knowledge graph practice area because we do… I’m with this team that has been doing work for many years in entity resolution and most people have probably never heard of it, but most people have probably used it. So, the idea is, say you have a bunch of different tables or data sets and you want to try to find what are the consistent entities inside these tables.
Paco:
So, you might have Bob R. Smith and Bob Smith Jr. and they’re both at 101 Main, but one of them is spelled 101 Main Street and the other is maybe spelled a different way or abbreviated a different way. And if you can think about that kind of problem, but spanning across billions of records in a lot of different data sources, how can you pull out the consistent entities?
Paco:
And it sounds like a trivial data science problem. We could just use string distance, Levenshtein distance, which is a typical thing. But when you take into account the fact of, well, what if you’ve got Bob R. Smith at 101 Main, Bob R. Smith Jr., but then you get Bob R. Smith Sr. at 101 Main, and they’ve both got voter registration. Is that the same person?
Paco:
Because your Levenshtein distance will tell you it is. If you set a threshold on string distance, they’ll tell you they’re the same person. So, when they try to register for vote, one of them will be denied voting rights. And so, this problem becomes very much complicated when you’re working in a world where there are companies that have offshore subsidiaries and maybe you don’t know the actual owners.
Paco:
You might know some of the directors and you get a very tangled web of some very bad people who are moving a lot of money around to do very bad things offshore, sorry, illegal fishing, illegal lumber, overthrowing democracies in Asia or in North America, for that matter. Basically, when you get the problem of trying to understand who’s who and what’s what, and a lot of different people or companies or ships that might have a registry somewhere, but you don’t know exactly in a given business context who they are, how can you triangulate on them?
Paco:
And so, it’s typically not a matter of just a string distance, it’s a matter of, well, I have enough elements of their address that are in common even though there are five different ways to represent this address in Singapore. I can tell the difference between a company at the same address or a hundred companies that are in the same shopping mall, which actually in Singapore is really a hard problem to understand.
Paco:
And same thing for tax records or passport control. There’s an area called UBO, which is ultimate beneficial owner, has a lot to do with sanctions compliance and catching oligarchs and understanding who is trying to do money laundering in an offshore tax haven, who is funneling billions of dollars out of Kremlin assets to try to influence a campaign somewhere. These are the kind of problems we work with.
Paco:
And so, the long and short is that these are… If you look at any episode of Homeland or The Wire or NCIS, any crime drama, inevitably, the protagonist goes up to a wall and they’ve got pincushion, they’ve got all these clippings and photos and notes, and they take yarn and draw a graph between them. And the thing is, the people who do that real work, if you’re in the US, you’re talking about three-letter agencies. If you’re in the UK, you’re talking about four-letter agencies.
Paco:
But the people who really do that work 24/7, they actually use knowledge graphs. They use collaborative knowledge graph tools like Aptitude Global, SiReN, GraphAware, Linkurious, Esri, ArcGIS Knowledge, Kineviz. There’s a bunch of different tools that allow people to collaborate on building knowledge graphs to catch bad guys.
Paco:
In finance, we have acronyms like AML, anti-money laundering, or UBO, ultimate beneficial owner, or PEP, politically exposed persons. All of these things have to do with the fact that somebody has committed very large-scale crimes and governments have reacted by saying, “Okay, regulatory, we will not allow this to happen again.” So, you end up having data sets like LIFE, was a multi-government response to the problems of 2009 global financial crisis.
Paco:
And so, now, when companies want to engage in certain types of derivative trading, they have to have a unique identifier because you have to be able to understand who that company is in the context of a knowledge graph to be able to do due diligence. And so, yeah, that’s where we work. It’s a lot of knowledge graph work that doesn’t see the light of day because some people can’t talk about it.
Larry:
Exactly. Yeah, no. And as you were talking about that notion of string proximity, I think is the phrase you used, and that reminds me of that old saw now about things, not strings, which is sort of the transition… not transition, but a bridge from not simplistic, but less contextually aware understanding of a thing versus like, no, I understand this entity. I know what this thing is.
Larry:
I know which ship that is. I know which oligarch this is. And that’s where graphs come in. One of the many things inspired me to reach out was I re-watched your 2021 video on graph thinking.
Paco:
Yeah, cool.
Larry:
And I think that the… It seems like that notion of thinking going from a… because that sounds like a machine learning kind of thing, like the proximity of strings to each other versus the actual entity associated with those strings, which gets into the graph world. So, I would love to revisit just for people who are maybe, if not new to, at least newly, more importantly engaged in graphs to get just your overview of that graph thinking and how your thinking has evolved. Because three years is a long time.
Paco:
I know, I know. And definitely shout out to my colleague Jürgen Müller at BASF. I wrote the piece about graph thinking and some other articles and presentations after. And also shout out to other friends of ours who’d actually come up with the term “graph thinking”. We didn’t invent it, but we can give footnotes there about where this comes from.
Paco:
I think that probably one of the best personifications is in the real world, if you look at search and you look at that whole trajectory of going from AltaVista, which was a mess of keywords and good luck, you almost had to have a graduate degree in how to use AltaVista to use AltaVista. If you go from that world of strings to the current world of You.com, Perplexity, et cetera, of AI tooling, not just for search, but actually for productivity, for research and developing things, it’s a transit that in the middle of this was Google saying “things not strings”.
Paco:
We have to build more structure and context into our understanding of the world. And we do this by using graphs. And of course, Google had made the big change. The big split from AltaVista to Google search engine was the fact that they had a graph algorithm called PageRank and other kinds of graph understanding. And then, from there, the next big split was they said, “Well, it’s not just strings. We’re actually identifying entities and the relations between them.” And that was 2014, whatever, knowledge graph project did at Google.
Larry:
’12, yeah.
Paco:
2012. Yeah, 2012. So, you see these 12-year intervals of where this has been evolving into a plateau and then it reaches another. And now we’re at the stage of, well, actually we’ve got large language models, we’ve got RAG, we’ve got some other elements that are coming in here, but at the end of the day, we’re using graphs to organize it. There is a notion of graph thinking.
Paco:
I’ve given a talk at PyData Global and it was based off of an article we’d written about this. And in this, we took the example of a medieval village that had seven people in the village and everyone has a craft business. So, there’s the person who grows the grain, there’s the person who mills the grain, there’s the person who bakes the bread, there’s the person who ferments the beer.
Paco:
We did this little village with seven little businesses and then we started looking at the ties between them, because the person who grows the grain sells it in a couple of different directions. The person who grows the beer needs to have grain and water and hops and yeast. And so, you get this circular economy with seven nodes in a small graph.
Paco:
You get the circular economy, and then you start to understand with graph analytics, if you actually run a centrality algorithm on that little medieval village off in the Black Forest, you realize that one of these people has a really killer business and they’re going to be central to everything and they’re going to clean up. So, we built out this idea of just with a super simple data example of seven nodes and the graph and the relations between them.
Paco:
How can we represent a graph to understand the economics of this village? And somebody comes along who is, I don’t know, making pizza or they’re making sausages, something like that. They’re doing something else that hadn’t been introduced. How do they fit into that graph? What are the ties? How does that change the dynamics of the economics in this village?
Paco:
And we just use it to explore the fact of it’s something that anybody should be able to pick up on, is there’s an importance of understanding graph relations here. And then, we built out from that. The other part of the talk was really exploring sort of embracing complexity. Actually, we had used the example of something called Cynefin, which is, it’s a Welsh word that describes… I believe it translates to habitat, but it describes about context.
Paco:
So, if you have a really simple problem, and it’s something that anybody can go through a minimal amount of training, here are the best practices, you get A, then you do B and you end up with C. Entry level order processing, something like that. If you have that kind of problem, there’s a certain level of math, if you will, a certain level of understanding, and you can control the chaos, so to speak, with just some best practices.
Paco:
And they’re like, “I’m going to teach you how to do your job. If you have this problem, you do this thing. If you have that problem, you do that thing.” And that’s great for a lot of problems in the world. It’s relatively simple. Gosh, spreadsheets are almost overkill. You can just have a checklist. But if the business problem gets more complex, then you need to have leadership, you need to have analysts.
Paco:
You have a more complicated workflow where, okay, well, we can do some analytics and we can have some different tables to describe things. And this is the level of chaos that would be controlled by having a relational data warehouse. And you’ve got the hierarchy of business executives, but then you’ve got another cadre of analysts who can manipulate the relational databases and come up with answers.
Paco:
Here’s your year-over-year yield, therefore make the following decision. You’re kind of guesstimating, but you’re guesstimating in a very educated way and you’re using certain level of math to do that based on the data. But that’s where, okay, you’ve got a thousand spreadsheets to be able to come up with your tax reports at the end of the day, which is, unfortunately, the case for 95% of the global 2000. But that’s tabular thinking and we can drill into that.
Larry:
Yeah, I would love to just back up because, to me, that was… I’ve talked to a ton of people about this. But that transition at the foundation of this, that because we’ve been stuck in relational databases – and we talked about this before we went on the air – for 60 years-
Paco:
Yeah, exactly.
Larry:
… just because of some IBM propaganda back in the day. Just kidding. And then, breaking out of tables and tabular thinking and piles of spreadsheets to get to this kind of graph. It seems it ties into that Cynefin framework of-
Paco:
Exactly. Well, just to complete about Cynefin, so the idea was that when you get into more of a complex realm where business leaders don’t actually know the answer and they can’t hire analysts to come up with the answer, you have to explore the phenomenology of what’s happening before you can start to put together a plan to how to respond.
Paco:
And increasingly in global scale business… Well, increasingly in business, it’s complex. It’s not complicated. Your analyst wizards can’t give you the answer. Instead, as a leader, you have to try to understand and discern, where are we? Because if you don’t, then it goes into chaos. So, the idea of Cynefin was this four different domains of simple, complicated, complex and chaos.
Paco:
And so, you have the known knowns in the simple case and you have the known unknowns in the complicated case, but you have the unknown unknowns in the complex case. And then, you don’t actually know what the hell you’re doing in the chaos case, you just sort of run screaming away from it. And so, it’s something that came out of the, I think 1999 out of an IBM consultant who came up with this term.
Paco:
And oddly enough, certain high-level officials in the US government were given a briefing. So, there was the famous quote about unknown unknowns, and it does trace back to a Cynefin framework. So, we were using this to try to describe that there are certain kinds of analytic tools. If you have a simple case of known knowns, you can give an entry-level person a checklist and just say, “Do the following.”
Paco:
That level of inference or reasoning, it dates back to Code of Hammurabi. The Sumerians have checklists of do the following, which is brilliant. But when you get into needing to have a lot of data driving your decisions and a lot of analytics and having an analytic process, a data science practice coming out of data warehouses, data lakes, that’s the complicated case.
Paco:
But when it becomes more complicated than that, guess what, you end up with a graph because our world is interconnected. And so, the way to address these kinds of complicated problems, to understand the phenomenology of it is to understand what are the entities out there that I’m dealing with? What are the relationships among them? What are the different properties that I can attribute? What are the semantics of all of the above that I can try to have some sort of meta understanding about it, which is an ontology? How can I build a phenomenology of an increasingly complex world and be able to operate in it in a smart way? And this is where we’re seeing increasingly large firms are turning to knowledge graphs.
Paco:
I was just at the K1st World IAC conference at Stanford, our third edition, and it was interesting because this is heavy industry, Patronus, Hitachi, Panasonic, Samsung, et cetera, talking about their AI practices. And the second year that we held the conference, almost every AI lead referenced their knowledge graph practice. And then, the third year, I mean this is just part and parcel, everybody’s talking about graph RAG.
Paco:
So, it’s interesting that when you get into the kind of government cases that we talked about with Senzing or finance, where even the practitioners can’t usually talk about it, the fact is they’re using knowledge graphs at scale. You talk to the heavy industry people where they’re using AI to confront really complex world of global supply chain, ESG, et cetera. They’re using knowledge graphs.
Paco:
So, it’s because of these really complex problems that people are having to address that we’re seeing graphs sort of mimic that arc of going from AltaVista to You.com of where we go through a realization of things not strings, and then we go through a realization of how do we really leverage the graph analytics and graph ML and all these other aspects.
And the thing that they have in common is they have the word “graph” in them, but we can break that down because it also means many different things.
Larry:
Another thing we talked about before we went on the air is disambiguating all the various meanings of the term “graph RAG”. And maybe this is a good time to talk about that because you just mentioned that, it sounds like there’s these varying degrees of complexity that the organizations and other entities have to deal with.
Larry:
And then, there’s the architectures and the other implementations of this stuff. Can you talk a little bit about, and I don’t know if it’s exactly analogous to AltaVista to Google, to Google with knowledge graph to You.com, but is there something in the LLMs to RAG to graph RAG to whatever’s next that you see coming?
Paco:
Yeah, it’s really super interesting and we just got a good dose of it because we had IAC, but then we followed up with AI conference in San Francisco, and there are other events also. So, it’s been a really, really compressed last week and a half.
Paco:
Yeah, so one of the first things I think people should really understand when they’re trying to approach this and really understand it is that the word “graph”, just even talking about graph RAG, if you look at the different practices that have been published, the word “graph” means at least six different things, and wildly contrasting or complimentary things.
Paco:
But if you look at the Microsoft graph RAG, which, by the way, wasn’t the first but the loudest, if you will. Microsoft made a big splash in February about that. And when they say graph, they’re not using graph database, they’re using NetworkX inside their graph RAG library.
Paco:
And they’re talking about… First and foremost, the main thing that they were talking about is when you have a RAG solution, which is retrieval-augmented generation, what that means is you take your source content, like here’s PDFs or some sort of text documents, and I have a hundred million of them and I want to use them to come up with intelligent answers.
Paco:
So, when I get a prompt, I’m going to go and run that prompt through my embedding model, project it out into a vector space, and I’m going to find the other text chunks of all my input data. I’m going to find out which ones of those are the closest neighbors, and I’m going to put them into a list that’s ranked by priority, and I’m going to feed that short list to my LLM.
Paco:
The LLM is going to sequence it into a very nice answer that sounds like human language. That’s retrieval-augmented generation. That’s basically saying LLMs can be very creative. We want to ground them. In fact, we want to, instead of paying tens of millions of dollars to retrain a large language model, instead, why don’t we just throw in our own facts so we can update these things with our own context, our own business context.
Paco:
RAG is a way of basically co-opting recommender system technology to try to make LLMs grounded. One of the things, though, you find is that RAG relies on, with vector databases, you’re doing this nearest neighbor similarity search. And it’s great, it’s really useful for RecSys if you’re doing e-commerce, but if you’re trying to get answers for something that’s more crucial, it has fairly poor recall.
Paco:
It’s a fairly simplistic way of understanding how the world is connected. You’re only going to find the nearest neighbors. You’re not going to find the neighbors that are one or two hops out. So, the non-trivial connections, like if I say kitten, you and I both know kitten is related to house cat, and house cat is a kind of feline. It’s related to a tiger and a leopard.
Paco:
None of those words have anything to do with each other lexically. So, if all I do is think of them as strings, I’m not going to understand that they’re connected. So, if I’m just doing a naive search in a vector database based off of the strings, the embeddings will only go so far. But if I start to tie it together with a graph, I can say what is a very large grownup kitten as a prompt, for example, to an LLM.
Paco:
I can run that through, bring in some text chunks that are relevant to the word “kitten” and larger. And then, I can start to traverse a graph that’s an overlay on them to say, “Well, a cat is a kind of… sorry, a kitten is a kind of cat and a big cat.” An example of that might be a tiger or a lion or a leopard.
Paco:
So, I can come up with more intelligent answers and maybe go out and retrieve some chunks that are non-trivial and put those into my priority list of chunks to feed to the LLM to produce the end answer. That’s what graph RAG is in a nutshell. The trouble is the graph part of it might mean a half dozen different things.
Paco:
You might have a graph, if you have a bunch of text sources and you chunk them by paragraph or whatever, 5, 12 character, whatever, you come up with these little chunks of text and you put each chunk through an embedding model, you’ve got a big vector describing it. You can look at what is nearby given chunk, what’s the distance from its relative neighbors. You can construct a graph based off those distance metrics.
Paco:
And that’s what Microsoft talks about in the graph RAG paper is just building a graph based off of the embedding metrics, not even having the semantics. So, that’s one. Another one that Neo4j talks about is, well, what if you actually take that text and do some analysis, some NLP analysis of the actual text that’s in the chunk? What if you parse that paragraph and build a lexical graph?
Paco:
Here’s my parse tree of the text, the five sentences that were in that chunk. I’ve got a little tiny graph of it. If I parse all the chunks of a hundred million documents, I’ve got a pretty interesting lexical graph to tie them together. And this is very different than the embedding kind of graph. They’re complementary. And if you use them together, they’ll catch things that would’ve been lost if you’re only using one.
Paco:
But you can also come in and say, “Well, what if I actually have more semantics in there? What if I have an ontology?” I know about my world. I am RHI Magnesita. I’ll shout out to a friend, Sebastian Kukla, who’s leading AI at RHI Magnesita. They do a lot of work in steel mining and steel mills. Very, very complex area. And what if I actually know the terms of art that have to be used if somebody’s creating an invoice?
Paco:
Because steel mills, shipping things, it’s kind of complicated and there’s export controls and tariffs, and on and on. What if I actually have an ontology to guide the answers that should come out of my LLM? And so, that’s another kind of graph. So, I could have a lexical graph, I could have an embedding graph, I could have an ontology that shapes this, on and on.
Paco:
I can also do graph neural networks to try to do node prediction. That’s something Google has been working on for their use of graph RAG. And so, we can build up a number of different definitions just for the word “graph”.
Paco:
And there’s more. You can also use an LLM to create facts out of your prompt plus your data. What are some assertions? Tie those together in a graph, and then instead of using the LLM to do reasoning, go from that graph to do reasoning. And so, there are projects like… oh, gosh, there’s a great paper called “Barack’s Wife Hillary”, which exemplifies the problems if you lean too heavily on strings, what kind of reasoning do you get? Because Barack Obama and Hillary Clinton, of course, because of the 2016 election were so adjacent.
Paco:
So, that’s one where they actually used models to create a graph and then do the reasoning based off of the graph that was generated. That’s another form. There’s also, if you’ve heard of chain of prompt thinking where you take a prompt going into an LLM and you decompose it into different parts and then use different agents or sessions in LLMs to go and solve for each part, and then you combine them together.
Paco:
Well, you can do that with a tree or you can do it with a graph. And so, there’s a lot of graph of thoughts now too, where let’s take a prompt, deconstruct it in different parts, we’ll start going down the different pathways to solve each part. We’ll keep track of the cost as we’re going down the pathway. If we get into a dead end, we can backtrack. If we’re starting to surpass our budget, we can go to a cheaper option. So, Graph of Thoughts is yet another way of coming in and handling the reasoning.
Paco:
And then, I think one that I will add from the K1st World IAC at Stanford a couple of weeks ago, I saw this great project from Urbashi Mitra at USC where they were taking and constructing causal graphs.
Paco:
And then, they were looking at subgraphs to be able to optimize them in terms of causality and be able to understand counterfactuals and do all the things you would do with causal graph analysis. But then, they were also doing reinforcement learning to try to optimize this over time. And the reason I’m mentioning it is because now we can get into some actual really good reasoning based on graphs that has all the evidence and attributions that you would want by virtue of being a causal graph.
Paco:
And we can do what-if testing and we can start to put together a plan and run counterfactuals against it. And that is sophisticated reasoning. That is the pinnacle of what judges do when they’re doing reasoning in a complex case. So, I’m mentioning a lot of different things in rapid fire here, but the point is graph means a lot of different things. There’s a lot of different kinds of graph technology.
Paco:
One that I didn’t even mention was the fact that once you get this multiple graphs together with your text chunks in a RAG rec system, you can run graph analytics to try to understand centrality and then re-rank possible answers. So, there’s a breadth of different technologies that relate to graphs and knowledge graph practice, and they come from many different sources, but we can start to put them together in the context of AI applications to come up with something that’s much, much smarter at the end of the day.
Larry:
I think if I counted right, you actually mentioned eight things.
Paco:
Okay, good.
Larry:
Yeah. So, a lot of disambiguation … Do we just need adjectives or there’s more? There’s a lot to that. But I guess I’d love to talk just a little bit about how all of those different conceptions of graph can be. Are there graph RAG things that use each of those? Or when we say graph RAG, how precise is that term at this point?
Paco:
Yeah, that’s a really great question. When you look at the open source tools, if you haven’t worked with graph RAG before, there’s some great resources. There’s actually a Discord that’s all about graph RAG. The lead committers for the different popular open source projects are on the Discord. You can ask some questions directly. Neo4j has been hosting that.
Paco:
I’ll shout out to Andreas Kollegger who put that together. There’s some great resources, but when you look at the libraries, they tend to be a little bit opinionated. Also, the libraries like LangChain, LlamaIndex, Haystack, and Microsoft are the four ones in the Python world. There’s another one in Java from Spring. You look at these libraries and they’re evolving. There’s a lot that’s happened.
Paco:
They’re not the most stable APIs, so to speak, because they’re rapidly evolving. So, the design patterns in them are a little iffy. And to understand, you might need some help on that, but there’s some good tutorials. They tend to include everything but the kitchen sink, which I think will change over time. So, the short answer to your question there is they’re trying to put everything in and see what sticks.
Paco:
If you look at the papers coming out of, say, Google in this area, they are using a plurality of methods. They’re going after it more formally. And I think that… I mean, it’s not public, but if you look under the hood at Google search or Bing, I imagine that seems to be what’s going on. Probably also the case if you look at what’s going on at Anthropic and OpenAI and others, they have these really sophisticated chatbots.
Paco:
We mentioned about You.com, Perplexity and others, Andi, all of these I think are having some different elements of graph to try to make sense of the answers that they’re providing.
Larry:
Interesting. So, it’s more, I’m just picturing, and one of the other things we talked about before we went on the air was this pretty apt analogy to Kahneman’s fast and slow thinking, that LLM were kind of the Systems 1, the graph world kind of Systems 2. And then, these graph RAG and similar RAG implementations are the attempt to get 1% of the way to a human brain or something, I guess.
Paco:
Yeah, it’s really interesting. When we look at discovery and how do you understand content, how do you monetize content, certainly search was early with the success of Google and all. But discovery, in general, I think had shifted more toward like RecSys.
Paco:
So, when you look at practices like Amazon in e-commerce or Google search, Bing search, Meta with social networks, Twitter, X with social networks, Pinterest, on and on, a lot of these huge eCommerce success stories were really grounded on the fact that recommender systems were the workhorse. And each one of those has very large knowledge graph practices associated with their recommender system.
Paco:
And so, it’s natural now that we have a bump in language model technology and we’re seeing all these applications come from it, even though it’s only a couple of years old. You can look back to the Gu paper and the Lewis paper from, I think, 2020 and 2021 about RAG, and then you start to see where, I think it was Nebula, the one out of China. They were one of the first ones in, was it 2020?
Paco:
Really, it was less than a year ago when they came out with one of the first graph RAG demos. And Microsoft, of course, made a lot of noise in February of this year, but there were other examples of it. LlamaIndex and LangChain had examples of it before Microsoft came out. So, it’s fairly recent, but the idea is that we have been borrowing from what we knew worked really well with recommender systems and applying that to AI applications now in general.
Paco:
And vector databases were really there for recommender systems. So, it’s natural you would see this through line. But I think what we’re hearing from the groundswell of conferences recently and what the leaders for a lot of these projects are talking about their pain points and their near-term outlook, I think that we’ll see more reinforcement learning coming in.
Paco:
We’ll see a lot of graph technologies coming in. We’ll see a lot more advanced math, like I mentioned, causality. There are other areas of using fairly advanced math in this area. So, it’s not all just about language models. And in fact, when you talk to people about production systems, we’re going back to what we learned from the machine learning tech debt paper years ago, was that LLMs were 10% of the problem. They get all the headlines, but the real system is actually much more complex and graphs play an enormous part of that.
Larry:
Right. That’s why I wanted you on the show was to get that point across. But hey, Paco, I can’t believe that we’re coming up close to time. But before we wrap up, is there anything last, anything you want to revisit or reiterate from the conversation or anything you just want to make sure we leave with the folks before we wrap up?
Paco:
It’s interesting, we were trying to present about graph thinking back in 2021, and the world has changed so much as we’ve discussed. But it’s interesting because it does provide a really good rubric. Back in the day, IBM, even when databases were a new thing, IBM had graph databases.
Paco:
Going back to the late ’60s, and System/370 and all that on mainframes, they had graph databases, but they really understood from market research that graph databases back then were really going to only be used by the best of the best. And so, they made a very deliberate choice. I used to work for one of the executives who’d been in the room at the time at IBM.
Paco:
They made a deliberate choice to dumb it down and put up guardrails, and the result was something that we now know and love called SQL. So, over the years, we’ve been brainwashed because Microsoft wanted to expand their market share of workforce. We’ve been brainwashed that everything should be a table. The reality is out in the world, yeah, tabular data is important.
Paco:
I work in an area of sanctions compliance where there are watch lists, there are tables, they’re very important. Your CFO’s spreadsheets are very important to the IRS. But when you look at the world, there’s kind of a Pareto ratio of an 80:20 rule, where there’s 20% structured data, 80% unstructured data.
Paco:
And I think that there’s a problem right now with graph RAG where the lead committers are making a deliberate choice of, “Well, we can’t do everything. What’s our minimum viable product?” So, there’s been a very conscious choice amongst the graph RAG people to say, “Well, let’s use an LLM to just pull in some unstructured data sources, build a graph, and we’ll just take care of that under the hood.”
Paco:
“You don’t have to worry about it. The graph’s not important. We use it as a means to an end to be able to ground the LLM using graph RAG.” And when you look at these open source projects, that’s kind of how they started out. They’re slowly but surely starting to make some affordances, where if you have a knowledge graph that you have curated because your business depends on it, maybe we’ll let you use parts of it.
Paco:
But I think that this is a shift in thinking that has to happen, because frankly, the LLM people don’t really understand the graph space very well. This is what we’re finding on the ground. And so, there needs to be more cross-pollination between the communities to tell them, “Hey, look, we are using graphs. And by the way, if you are involved in chasing after oligarchs who are doing really bad things in a particular country, you have a graph that you’re using to put together evidence that you’ll take them to trial on. And you don’t want an LLM to just make up stuff because I’m pretty sure the judge will throw the case out of court.” So, we have to go from the mission-critical enterprise world, the adults in the room who use graphs 24/7 for really important things that keep the world running. We have to go from their needs to where the graph RAG developers are right now.
Paco:
And I think what I’m seeing in industry is a lot of teams that have more mission-critical needs, they’re building graph RAG on their own. They go out and they use something like, I don’t know, LanceDB and KùzuDB. And they throw a few things together, Llama and whatnot, and they build their own graph RAG because you can. And then, you can start to do the right thing.
Paco:
So, that’s a snapshot of what I’m seeing right now is this transition from different camps that don’t know each other, trying to cross-pollinate, and enterprise teams trying to do the right thing. And I’m hopeful that it will get better. But that’s circa Q3 of 2024.
Larry:
Okay. Yeah. And just to timestamp this, we’re talking on September 16th, because that seems really relevant now, how quickly this might unfold. Well, thanks. I think this is really going to help people get oriented to the… because I get the feeling, I don’t know. I’m not as sophisticated a consumer of this as you, but when I look at LinkedIn, I think, “Wow, where did all these graph RAG experts come from?” And I just want to ground them in your personal knowledge graph about this stuff. So, yeah.
Paco:
It’s definitely picked up. I’m glad to see how much it’s picked up. I’m a little concerned though because there’s probably a 10:1 ratio for the number of graph RAG experts versus the number of people we’ve actually worked with a library.
Larry:
That’s where I was getting at and you wrote some of those libraries. Anyhow, hey, one very last thing, Paco, if folks want to connect with you online or follow you, what’s the best way to find you on?
Paco:
Great. I’m on LinkedIn. If you look for Paco Nathan, and my tagline is Evil Mad Scientist, you’ll probably find me on LinkedIn. I have a public profile on Sessionize. We can provide it in the show notes. And my portfolio is DerwenAI/Paco. But yeah, no, please connect up with me on LinkedIn. We’re really trying to put together more of a community.
Paco:
We’ve got some great conferences. Connected Data London will be coming up in December in London. And we’re both working on that, so we hope to see you there. Knowledge Graph Conference will be coming up next May in Manhattan, and there’s other communities like GraphGeeks that really involved with. So, if you want to learn more about this space, there’s a lot of great community resources.
Larry:
Yeah, I’ll link all that stuff in the show notes. And if you can send me, you mentioned a couple of papers, I’ll send you a note afterwards and try to get it. But anyhow, well, thank you so much, Paco. Always fun to talk.
Paco:
Thank you, Larry. I really appreciate it.