Robert Sanderson: Building Yale’s Cultural Heritage Knowledge Graph – Episode 46

photo of Robert Sanderson, expert on cultural heritage knowledge graphs and ontology engineering design patterns
Robert Sanderson

Yale University manages huge collections of precious cultural heritage artifacts housed in multiple museums, libraries, and other collections.

Using knowledge graph and ontology engineering design patterns that he has developed over his career, Robert Sanderson helps scholars, researchers, and the general public access information about — and make connections across — millions of unique items in Yale’s collections

We talked about:

  • his work as Senior Director for Digital Cultural Heritage at Yale University
  • the knowledge graph and ontology engineering design patterns that guide his work
  • the scope of his work — improving discoverability of Yale’s extensive collections of artifacts, facilitating the management of collection information, and even collecting data on physical artifact storage facilities
  • how their linked data approach lets researchers easily connect information about artifacts and information housed in multiple museums, libraries, and collections
  • how the growth of LLMs has affected their KG user interfaces
  • how AI is accelerating their ability to add to their knowledge graph the millions of artifacts in their collections that aren’t yet accounted for
  • the compact nature of their three-billion-triple KG ontology, just 10 classes and 50 relationships
  • the extensive vocabularies and taxonomies they use
  • how they handle the need to reconcile the identity of lesser-known people who don’t have a Wikipedia page or other authoritative references available
  • how they balance the competing needs of comprehensiveness and usability as they build their knowledge graph
  • how knowledge graphs facilitate discoveries that other search tools can’t
  • current opportunities for post-docs to join his team to work on leading-edge AI projects

Robert’s bio

Dr. Robert Sanderson is the Senior Director for Digital Cultural Heritage at Yale University, where he works with the libraries, archives, and museums to ensure that data and other digital efforts are coherent and connected. He is the principal architect for Yale’s cross-collection discovery system, LUX, which is built on the Linked Art specifications, for which he is an editor. He is also an editor for the IIIF specifications, was the co-chair and editor for JSON-LD and the Web Annotation data model in the W3C. He has previously worked at the Getty in Los Angeles, Stanford University, Los Alamos National Laboratory, and the University of Liverpool. His current areas of work and research are at the intersections of cultural heritage, knowledge graphs, data usability, and generative AI.

Connect with Rob online

  • LinkedIn
  • email: robert dot sanderson at yale dot edu

Rob’s LinkedIn post series on KG and ontology design patterns

  1. The 10 Design Principles to Live By
  2. Ontology Design Patterns
  3. Naming Things
  4. Avoiding Reification
  5. Foundational Ontologies
  6. Multiple Inheritance, Not Multiple Instantiation
  7. Predicate Reuse… Meh
  8. Document your ABCs
  9. Separate Query and Description Semantics
  10. Usable vs Complete
  11. acknowledgements

Video

Here’s the video version of our conversation:

Podcast intro transcript

This is the Knowledge Graph Insights podcast, episode number 46. When your job is to help scholars and the public discover information about millions of cultural heritage artifacts that are housed in multiple museums, libraries, and other collections, you need a powerful — but also manageable — knowledge graph. That’s Rob Sanderson’s role at Yale University. He and his team apply time-tested ontology and knowledge engineering design patterns to help people discover — and see the connections between — these precious human artifacts.

Interview transcript

Larry:
Hi everyone. Welcome to episode number 46 of the Knowledge Graph Insights Podcast. I am really delighted today to welcome to the show Robert Sanderson. Rob is a professor and the senior director of Digital Cultural Heritage at Yale University, the Ivy League School in Connecticut. Welcome to the show, Rob. Tell the folks a little bit more about what you’re up to these days.

Rob:
Hi, Larry. Thank you so much for inviting me to be part of the illustrious lineup of guests on your podcast. So yeah, I’m Rob Sanderson, as you said, Senior Director for Digital Cultural Heritage at Yale. So I work with the libraries, the archives, and the museums and other collecting organizations at Yale to help them to be more connected with linked data organizationally and more coherent in the way that we do things digitally. So our projects really focus on discovery and access to the collections in service of the university mission, which of course is teaching and learning, research, and preparing our students to be the next generation of leaders in the world.

Rob:
So for that, the university invests very heavily in the collections, which is fantastic. We are super proud of the 300 years of collecting that we’ve done. But we want to make sure that if you can’t come to New Haven, you still have as good access to those collections as possible. And the ability to find amongst the many millions of objects that we steward exactly what it is that you need. So a lot of our projects focus on describing the collections in a more computationally tractable way so that that discovery can be better. And also how to manage the information that’s associated with the collection, but isn’t a museum object or a archival object itself. For example, I have two postdocs that are openly available. So if you are a few years out of your PhD or just about to graduate, do get in touch to work on how to use AI to extract the ownership history or the provenance of particular museum objects from the archival content that we also manage. Equally, how can we align research data sets with the collections? So we also have a natural history museum as well as two art museums. How can we align the environmental datasets that are out there on the web with the natural history specimens that could have been impacted by those environments?

Rob:
Yeah. And then equally, we look at the environment of Yale. So we have a large project at the moment to set up environmental monitoring with sensors for light, for humidity, temperature, and so on, to be able to generate a large data warehouse aligned with linked data with the collections so that we can have evidence of what the effects of the environment are on the collection items themselves.

Larry:
Interesting. That is so fascinating. What a fascinating remit. One quick thing about what you just said. Is that about humidity and temperature and all the things that might affect the endurance of these physical artifacts?

Rob:
Yep. Yes. That’s right.

Larry:
Yeah.

Rob:
We have about 200 sensors around the place monitoring every five minutes a new data point, which if you think about it, it’s actually not that much data.

Larry:
Yeah. I have to say, I just love that you’re doing data stuff along with it. That you’re not just sitting in a dusty old room collecting things. You’re doing cool modern stuff too. But hey, I want to quickly interject how we met, and I just want to put this in because we won’t have time to talk about it today, but I want people to know about this fantastic series you did. That’s how we met was somebody drew to my attention the series you’ve done on ontology design and on knowledge engineering design patterns. And I’ll point to that in the show notes, but I just wanted to mention. And the more I think about what you just said, because I didn’t know all of this background before we started recording, I’m like, “Oh, this is even better than I thought.” So I’ll point to that in the show notes.

Larry:
But the main thing I wanted to talk about today is what you were just talking about. This amazing cultural heritage operation that you’re running there, especially the knowledge graph component of it and the AI, of course, because we’re in the 21st century, and that’s all anybody talks about. One of the things we talked about before we went on the air was how AI is accelerating the ability for you to build your knowledge graphs of these cultural heritage artifacts and data. Can you talk a little bit about that, how AI is helping in that?

Rob:
Yeah. Of course. Absolutely. So just a little bit of a background about the knowledge graph itself first before I get to the AI part. So over the past five years, we’ve built without AI, a very large scale knowledge graph, well, in cultural heritage terms of very large scale, which has about three billion triples in it. And it follows the principles and the design patterns that you mentioned in those posts on Linked Art. It then aligns the people, places, concepts, events, objects, works, collections that we manage here at Yale across the two art museums, Natural History Museum, the dozen or so libraries. There’s also a collection of musical instruments, the Institute for the Preservation of Cultural Heritage, and we even have a little outpost in London, in England for art history research that we include. So that work uses the linked art ontology, which is based on the foundational site CRM ontology and is publicly available both in terms of the data, you can just download it. But also in terms of the graph queries, we don’t force you to learn SPARQL. We have a user interface on top of it, which allows you to generate queries and find the objects that you are looking for.

Rob:
So one of the things that we noticed first about the user interface is that only about 5% of searches are actually using the graph affordances. Mostly, 95% of the time, people just put in keywords because that’s what they’re used to. You go to Google, you type in your five favorite keywords that you think might match and you scroll through the results. However, now in 2026, people are more used to typing in full sentences and then having AI take that natural language and process it somehow. So we are in the user interface phase of building out not a chatbot, but I don’t think there needs yet another chatbot, but instead a natural language processor that takes the user’s research query, a research question and translates it into the graph query that will take them to the objects that might help them to answer that particular question.

Rob:
So that’s our front end work. But in terms of actually building the knowledge graph, that’s where we’re focused for many reasons. One, we have linear miles of archives and not if you laid them edge to edge, but in a regular folder or books of which only about two to 3% have been digitized. Of that two to 3%, most of them have not been described beyond just, here is a folder and it’s got some letters in it or here is a box and it has a whole bunch of photographs, but that’s it. You don’t know what’s on the photographs, you don’t know who the letters are from, who they’re to, what they were talking about or anything.

Rob:
So we did an experiment, which was to use our high performance computing cluster to do handwritten text recognition using large language models on 650,000 images taken from the archives. That then cost us about $1,200, including the replacement cost of the H200, because it took about 120 hours. And if you divide up the cost of the H200, then yeah, $1,200. That would cost about $12 million if we were to pay people to do it because that’s about 14, 15 people for a decade just going through every single day transcribing page after page after page.

Rob:
So once we have the full text of the information, that’s where we can start to process what we have in a way that makes it more discoverable, more computationally tractable. Scott Weingart a couple of weeks ago who was at the National Endowment for the Humanities and is now the CTO at the University of Virginia libraries said essentially that people will understand history through what is discoverable, not through what actually exists. And this is just a rephrasing of the old adage of if it’s not on the web, then it doesn’t exist. So how can we get without spending that $12 million to the point where those archives are part of the knowledge graph? So the handwritten text, recognition is the first step, of course, along with PII detection and other … We don’t want to expose people’s personal information through some overzealous AI transcribing little footnotes and so on.

Rob:
But then we can use the language models to go through, understand the ontology, which uses the same techniques as we do for the AI querying. And then produce from, paragraph by paragraph, here are the triples which are, according to the ontology, described in the paragraph of text. So for example, imagine a letter from someone in Germany just after World War II to Alfred Stieglitz in New York City, and they’re talking about the conditions in Germany and how … “Did you receive my last letter?” So now we know that there was a letter. So we can say Fritz Gertz corresponded with Albert Stieglitz at this time. We know that he also corresponded on these dates, even if we, Yale, don’t have those particular letters because maybe they were intercepted by the German authorities, maybe they got read and thrown out by Alfred Stieglitz. We don’t know, but we do know now that they did exist.

Rob:
And he also talks about correspondence with other people that we don’t have the archives for, but other famous contemporaries such as Edward Steichen or Georgia Engelhard we know that they were corresponding even if we don’t know what they were talking about. So that social network, then you can’t get from, there is a folder of letters, which is the only description of that particular collection that we have. Instead, by using the handwritten text recognition, we can expand the knowledge graph out from the objects into history, which then enables research and understanding of the past.

Larry:
Yeah. There is so much in there. I have so many questions. The first thing that I want to mention though is that it occurs to me at the very highest level, this ontology is probably fairly simple. You only mentioned a half a dozen core concepts, events, people, was it like collections, works? So is it that simple or am I missing something there?

Rob:
No. No. We’ve tried to keep it as simple as possible, and this falls back to those design principles where there is about 10 classes that we use out of maybe 30 or 40 higher level abstract classes. So we don’t directly instantiate things like conceptual object or abstract physical thing. Instead, there are human-made objects, there are physical objects, there are intellectual works and people groups and so on. So from those 10 classes, there’s maybe 50 or so relationships between them. But yeah, the fewer things that we need to manage over the long-term and with cultural heritage, long-term is really long-term, the more sustainable it is and the easier it is for us to work with over that time. And by us, I include AIs as well. So the better described, the more consistent and the fewer overlapping relationships, the clearer it is for human developers and for the AI reading those same letters to be able to determine which class do I need to use, which relationship do I use and so on.

Rob:
So it is relatively simple, based on our experience over the past 15 or so years building services in this space, but we also try to be extensible. So instead of trying to capture a hundred percent of everything that anyone might ever want to talk about, we try to capture the 90% that everyone does need to talk about and leave that last 10% for extensions. So at the moment, for example, we are looking at the archives of performances, so in theater or at festivals and so on, what does it take to talk about a performance as opposed to a painting that hangs on the wall? So we’re using that extension model to say, and here is how within the foundations of the ontology, you can use the same consistent patterns to talk about this new domain. And because it’s built on a robust foundational ontology, we have all of the concepts needed. It’s not that they don’t exist, we just need to select them and say, okay, this is the patterns that we use. Here is the new predicates that we’re going to introduce between these few additional classes and then rely on the good work of 30 years of ontology design, both in the cultural heritage sector and beyond.

Larry:
I’m going to pull that out as a high level benefit of building a good knowledge graph. But anyway, there’s again, so many questions. But one thing I wanted to ask about that, you were just talking about the extension of the model, and as you described, the quite simple, the highest level conceptual concepts that you’re working with, but millions of artifacts. I’m wondering, do you know Dave McComb’s notion of the CBox? Basically vocabulary and taxonomy management. Is that a huge part of this knowledge graph? Because you just have works or collections, but you’re across all these different disciplines and kinds of artifacts. Is that all managed with vocabularies?

Rob:
Yes. Yep. Absolutely. So it would be completely unwieldy if we had to, for example, say, here are the class of painting versus, here’s the class of book versus class of newspaper and so on. Particularly because we also have the Natural History Museum, the Peabody Museum. So then you have all of the taxonomic hierarchy of the animal and plant classes, and we’re not going to create a new class for every single species, so we have to use taxonomy and vocabulary. Also, it means that the developers and AIs have that ease of use of the data because at the end of the day, unless there’s a new relationship between the classes, and hence you need them for domain and range, the fact that it’s a painting versus a watercolor versus a print versus a sketch, the humans care about that and we need to be able to display them. So we need a concept to say, “This is a painting, find me all of the paintings and this is a drawing, find all the drawings.”

Rob:
But it doesn’t need to be a class in the ontology. We can just use that concept notion with the vocabulary to be able to say, this is a one of these or has a classification rather than the is a of painting.

Rob:
One of the things which is very weird to non-art historians, but I’ll try and explain it, is there is a widespread disagreement about watercolors. So watercolor, it’s a painting. Well, no, is it a drawing? So yeah, there is a widespread disagreement about whether watercolors are paintings or drawings. So we don’t want to get involved at that level in the ontology that we need to keep around and be able to use. Instead, when they figure it out, it’ll be part of the vocabularies and it will have the right hierarchy of narrow and broader terms using scarce and it will just work.

Larry:
And the engineers-

Rob:
They can argue about it all they want.

Larry:
Yeah. So the subject matter experts can get their say in, which is the whole point of this thing, but the engineers are just like, “Ah, just another vocabulary.” Yeah, that’s awesome. Hey, one thing I wanted to come back to, and I wanted to connect it to something else I wanted to talk to you about, was that notion that things don’t exist until they’re discoverable. That’s so powerful, and it sounds like the opportunity here is so huge. But another thing you talked about in relation to that before we went on the air was about this notion of people who don’t have a Wikipedia page, how do you resolve their entity in this because people must be a huge part of this, artists and all this stuff is human created. Can you talk a little bit about that, discovering people who aren’t well known and how to make them discoverable?

Rob:
Yeah. This is a huge part of history, particularly in this era. The joke, what did Watson and Crick discover? Rosalind Franklin’s notes – is something that drives us. How can we make sure that we are being responsible to what actually happened rather than just following the famous names, the middle-aged famous white guys, he says as a not very famous middle-aged, not very wealthy white guy. And instead reveal the people who deserve the credit for what they did. And in the archives and in museums, that’s where the information is managed for history. So the more that we can use AI to extract the knowledge and then make it discoverable, the more we can bring to light those hidden stories.

Rob:
So one of the things that we’ve been experimenting with is using Wikidata and the links from Wikidata to Wikipedia to be able to then build a social network of people who are mentioned in either of those two data sources. So imagine in the letter from Frederick Gertz to Alfred Stieglitz, he says, “Give my regards to Ida.” Without the social network, there’s no way of knowing who Ida is. It could be one of a gajillion Idas. However, we know that Alfred Stieglitz was at the time married to O’Keeffe, the famous painter from New Mexico, and Georgia O’Keefe’s sister is called Ida. So it’s almost certain that when Gertz refers to Ida, it’s Ida O’Keeffe.

Rob:
So Ida does have a Wiki data entry. I forget whether she’s notable in Wikipedia terms enough to have a page, but certainly the mother and father of Stieglitz or of the O’Keeffes, they are not famous enough to have Wikipedia pages, but they are mentioned by name and there’s information in the full text of the Wikipedia articles about them, just they don’t have their own article. So mining the encyclopedic knowledge of as many open data sources as we can lay our hands on, not for the famous people, but for the people related to the famous people. So we don’t collect every single archive of every single person that we’d fill up many, many warehouses for the stuff if we tried to do that. But for the archives that we do collect, how can we make sure that the people who contributed to the success of the person to make them archive worthy, how do their stories get told along with the famous white guys?

Rob:
So yeah. We are relying somewhat on AI again to do all of that processing, to extract the linked data graph from the Wikipedia texts, and we use Wikidata as a way to then ensure that the records are reconciled. So if we have, for example, a person who is a painter in the union, our list of artists names that the Getty managers, and maybe we also have a book about them, that the Library of Congress manages the identity. And then the Library of Congress often refers to OCLC’s VIAF, the virtual and national authority file, and VF is then reconciled with Wikidata. Wikidata is reconciled with. Wikidata also has the links to the Wikipedia page, and now we’ve connected the circle without having to rely on labels or other potentials for mismatch with polysemy and the same name of different people or the same person, but with different names over time.

Larry:
That’s so interesting. So there’s enough various sources and they’re all connected, it sounds like fairly well, that you can disambiguate most of these, not second-tier, but these less well-known people. Does a need ever arise to create some canonical, like a Wikidata-like thing about them that might not clear … like you said, there’s that requirement for notoriety to get into Wikipedia. For people who don’t clear that hurdle, is there enough like that Library of Congress and other connections you just mentioned, is that enough or do you sometimes have to kludge together some way to resolve these entities?

Rob:
Yeah. No. We frequently create new identifiers for people who are not notable enough to be in any of the authority systems. And indeed the authors of books that are not published in America, we have trouble reconciling as well for the library work or the subjects. So if you have a book about someone, then we still need an identity for that person so that we can have the relationship over from the work to the person that the work is about, as opposed to the author who wrote the work. So we have more than five million people in group records of which I don’t know the percentage, but it’s not that high. I reconciled directly with one of the major authorities. We try our best, but we also accept on every single page in the user interface feedback. There’s a very prominent blue submit feedback button to then get people to write in and say, okay. No. You’ve misaligned this person or you’ve merged two people who are really separate, or you haven’t merged these people together when they should be merged and so on. And we get feedback every day from people who often are searching for themselves and find their record in our knowledge graph and then submit to say, oh, hey, can you update this or you’ve merged me with this other person with the same name and so on.

Larry:
Hey, now as you’re reminding me there of the scale of this, and you mentioned earlier the need to reconcile, the completeness that any archivist … I’m a pack rat, I want to know everything and keep everything, versus the semantic tidiness that really makes something like this hum. How do you balance those, the volume you’re dealing with and there must be competing human interests as well, people who want to know everything about a certain niche subject area, but you just don’t have the resources. And yeah, there must be other issues as well. How do you address that?

Rob:
Yeah. And this is a big challenge. It comes back, of course, to what data we have. So a lot of the time in historical bibliographic records from the past decades of library cataloging, all that was written down was the name of the person. So that’s all we have. So we do our best to try to figure out who that was algorithmically. But with millions of people being talked about and more records being added every day, we don’t have the person time to do that by hand. So we do rely on external authorities a lot to try to enhance and enrich the records rather than us, Yale managing everything. So that then gives us a baseline for the work. But in terms of the sorts of information that we manage, that’s when the link back to the ontology comes in where there needs to be some relationship that we can express between two entities that we have in the knowledge graph.

Rob:
So for example, we know where Rembrandt lived. So we have the place down to the street address and we have Rembrandt as a person, not because we have that many Rembrandts, but because we have a lot of books about Rembrandt as a famous artist. So there is a challenge of boiling the ocean. How far do you go? Do you then try to find all of the other people who lived in that same house? No. Clearly no.

Rob:
But equally, we need to work with our stakeholders, and I think this is really the best practice. Understand what your system needs to manage and manage that, and then maybe go one step further, but don’t try and have a single system that manages all knowledge from anywhere. Before we went to linked data for the discovery system, we tried to use the more traditional full-text database type queries of Solr or Elasticsearch or similar. And we did so many demos to our stakeholders that the curator of American art in the art gallery said, “I never want to see another Solr demo until you can answer this question.” The question was, which paintings do we have that were created by people from Europe, but depict a place in America? You can’t do that without a knowledge graph. It’s just impossible because you have too many joins. We didn’t manage all of the information about the artists. We just had their names for some of them. So who is European? Well, what does European mean? Well, it was born somewhere in Europe or has a European country as a nationality. Somewhere in America, well, that’s an awful lot of places. And if you’re managing the identities of the places, you need to then be able to do transitive queries to say, is this anywhere within hierarchically the United States of America?

Rob:
So at that point, we then went back and said, “Okay. Let’s try this as a knowledge graph.” Which of course suited me to no end, given my background. And we can now answer that sort of question. And when the AI front end becomes available, you can just write that in and you get the answer.

Rob:
Yeah. It really comes down then to what do you need to do? What would you like to do and what data do you have in order to do those things? And then making sure that you’ve got the right ontology at the backend in order to be able to answer the questions that are needed. But yeah, the trade-off between usability and completeness really is the dark art of linked data. So if you have every single piece of knowledge that you know managed in the ontology, it’s almost certain that it will be completely unusable. So for example, if we tracked the provenance, the data provenance of every triple … So did the birthplace of Rembrandt, it doesn’t come from us. So did it come from Wikidata? Did it come from ULAN? Did it come from the Library of Congress or the National Library of France or Germany or any of the other 25 different sources that we use?

Rob:
That would then multiply our dataset by 10 because every single triple needs to have where did it come from, which could be multiple places, when did we get it and so on. So that was one of the questions that we had to decide on in terms of scope pretty early on because that was requested of us. We want to know where all this information comes from so that if something is wrong, we can go back and see if we can get it fixed. That led to the submit button as the halfway house. So we track the records that it comes from, but we don’t track which triples for each entity come from which sources.

Larry:
I love that. That’s such a great pragmatic example of the tough trade-offs in this kind of big project. I also love that example you gave of the ability to discern which paintings depict something painted in America by a European painter. I’m going to steal that because it’s hard to come up with those kind of examples sometimes, so that’s a great one. But hey, I can’t believe it, Rob. We’re coming up close to time. But before we wrap up, is there anything you want to revisit from the conversation or just make sure that you share before we wrap up?

Rob:
Yeah. No. Just another call out that if you would like to come work with us at Yale, we have some open positions and for postdoctoral researchers, we’re applying for some money at the moment from an external source, so keep your fingers crossed for us, which would open up another one. That would be to work again at this intersection of AI and link data. But for particular research questions around how …

Rob:
Once we have all of this knowledge built out, so we have hundreds of thousands of images of text that we can extract out into a knowledge graph, now what? Now we have this abundance of knowledge as opposed to the current scarcity. So how can we use modern cutting edge technology in order to make that knowledge more accessible so that researchers can do research so that people can interact with the archives and enjoy the content there, perhaps using customized on demand generated interfaces built by Claude Code or some other equivalent that specifically understands the ontology, understands the data, understands what the researcher needs conversationally, and then can then generate this on demand interface to be able to present the information from the graph to the user in a way that the researcher can interpret it rather than just being presented with the hairball graph of here are all the connections. Well, that really doesn’t help anyone, but a timeline would or a map or the list of all of the archival objects that refer to all of the components of the researcher’s question.

Larry:
I love this vision. I’m going to have to circle back in a year or two to see what kind of really cool architectures you’ve built, because everything you said is completely plausible with current technology. You just need to hire a few postdocs to do the grunt work and there you go. And that sounds like an amazing opportunity. I’ll mention that in the show notes as well. Well, hey, one very last thing, Rob, if folks want to follow you or connect online, what’s the best place to find you?

Rob:
Yep. So feel free to email me, robert.sanderson@yale.edu or LinkedIn. I’m happy to accept connections from all and sundry on LinkedIn. So yeah, Robert Sanderson on LinkedIn.

Larry:
Excellent. Great. Well, thank you so much, Rob. I’ll put those in the show notes as well, but thank you so much for the awesome conversation.

Rob:
Great. Yeah. Thanks, Larry. Thanks for having me.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top