Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Skills that Rebecca Schneider learned in library science school – taxonomy, ontology, and semantic modeling – have only become more valuable with the arrival of AI technologies like LLMs and the growing interest in knowledge graphs.
Two things have stayed constant across her library and enterprise content strategy work: organizational rigor and the need to always focus on people and their needs.
We talked about:
- her work as Co-Founder and Executive Director at AvenueCX, an enterprise content strategy consultancy
- her background as a “recovering librarian” and her focus on taxonomies, metadata, and structured content
- the importance of structured content in LLMs and other AI applications
- how she balances the capabilities of AI architectures and the needs of the humans that contribute to them
- the need to disambiguate the terms that describe the span of the semantic spectrum
- the crucial role of organization in her work and how you don’t to have formally studied library science to do it
- the role of a service mentality in knowledge graph work
- how she measures the efficiency and other benefits of well-organized information
- how domain modeling and content modeling work together in her work
- her tech-agnostic approach to consulting
- the role of metadata strategy into her work
- how new AI tools permit easier content tagging and better governance
- the importance of “knowing your collection,” not becoming a true subject matter expert but at least getting familiar with the content you are working with
- the need to clean up your content and data to build successful AI applications
Rebecca’s bio
Rebecca is co-founder of AvenueCX, an enterprise content strategy consultancy. Her areas of expertise include content strategy, taxonomy development, and structured content. She has guided content strategy in a variety of industries: automotive, semiconductors, telecommunications, retail, and financial services.
Connect with Rebecca online
- email: rschneider at avenuecx dot com
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 25. If you’ve ever visited the reference desk at your local library, you’ve seen the service mentality that librarians bring to their work. Rebecca Schneider brings that same sensibility to her content and knowledge graph consulting. Like all digital practitioners, her projects now include a lot more AI, but her work remains grounded in the fundamentals she learned studying library science: organizational rigor and a focus on people and their needs.
Interview transcript
Larry:
Hi, everyone. Welcome to episode number 25 of the Knowledge Graph Insights podcast. I am really excited today to welcome to the show Rebecca Schneider. Rebecca is the co-founder and the executive director at AvenueCX, a consultancy in the Boston area. Welcome, Rebecca. Tell the folks a little bit more about what you’re up to these days.
Rebecca:
Hi, Larry. Thanks for having me on your show. Hello, everyone. My name is Rebecca Schneider. I am a recovering librarian. I was a trained librarian, worked in a library with actual books, but for most of my career, I have been focusing on enterprise content strategy. Furthermore, I typically focus on taxonomies, metadata, structured content, and all of that wonderful world that we live in.
Larry:
Yeah, and we both come out of that content background and have sort of converged on the knowledge graph background together kind of over the same time period. And it’s really interesting, like those skills that you mentioned, the library science skills of taxonomy, metadata, structured, and then the application of that in structured content in the content world, how, as you’ve got in more and more into knowledge graph stuff, how has that background, I guess… what’s been the transition like as you start to consider knowledge graphs in your content work? How’s that been going?
Rebecca:
Well, I mean, librarians, we’re all about organizing things, and we like to organize stuff, and this is just sort of the next step in helping organize stuff so people can find things. I mean, that’s what we’re all about, right? We want to help people find things. So moving into more and more sophisticated mechanisms that help people not only find things, but leverage content in new and different ways and useful ways is, I think, just a logical progression.
Larry:
Yeah, that’s really… that content discovery, like finding things, discovery, and uncovering things, that’s sort of always been the information architecture-y part of content practice. But how is that changing? We were talking a little bit before we went on the air about the emergence of LLMs and then the ensuing interest in knowledge graphs associated with that. Are you finding different challenges in helping machines discover stuff as opposed to humans?
Rebecca:
Well, okay, there’s a couple of different aspects to that. One is that, as we know, structured content is so very important for the success of LLMs, AI applications, et cetera. And the thing is, I have to convince my clients to take the time to clean their house, so to speak, and say, “Okay, you need to clean up your data. You need to clean up your content. You need to get rid of the old, outdated, trivial, extraneous stuff so you have a clean basis to start from.” And convincing clients to take that time is a bit of a struggle because they see all the fancy LLMs and all different wonderful things people are doing, and they just want to jump in with both feet, which is great, but you need to have a solid basis first. So it takes some convincing to have people take the time to do that.
Rebecca:
And then on the LLM side, we have to acknowledge that there are many different kinds of LLMs with different pros and cons depending on what the use case is, and you need to not only understand the LLMs, but also the organization’s business drivers, what are their goals, what are their objectives, what are they trying to do with this? Because it’s not just to write me a poem in the style of e.e. cummings about apples or something. There are definite business drivers and goals because a lot of money is putting a lot of funds into these kinds of technologies, and you got to make sure that you’re getting your bang for your buck.
Larry:
Yeah. As you say that, you’re reminding me that everybody’s wrestling, you’re not wrestling, but figuring out how to do work with RAG architectures and graph RAG and all these hybrid AI architectures. And I love the way that it sounds like your work… you start with the business goals and kind of back up from there, which seems like a good approach. Are you discovering any patterns or approaches as you figure out how do you balance… because that’s an interesting triangle of needs, the business part of the client needs. There’s sort of content needs, and then the capabilities of the LLMs and the other tooling and these architectures. That’s got to be a lot, it sounds like.
Rebecca:
It is, and you also have to think about how much work am I going to make people do. You can’t train everybody to be prompt engineers, right? And so you need to think about from their perspective, what they’re trying to get out of whatever application you’re creating, and also their interaction with it to the extent to which, yes, the information is in there, but you have to ask better questions. Does that mean I have to retrain people on how to ask better questions? To what extent is that necessary, or are there other ways of leveraging how we use an LLM in this sort of hybrid structure? How can we use ontologies, et cetera, informing the knowledge graph to help the user so they don’t have to be prompt engineers, that they can get what they need with the minimum of fuss.
Larry:
Yeah, that’s really interesting. As you say that, using ontologies to inform the knowledge graph, I was just kind of assumed in my head that ontologies were a prerequisite for knowledge graphs, but I had Jessica Talisman on a while back, and she makes the case that you can judge a lot just with a SKOS-based thesaurus. Do you consider that kind of continuum of semantic sophistication from just term lists to taxonomies to thesauri to full-blown ontologies, are you playing across that spectrum in your work?
Rebecca:
Yeah, absolutely, absolutely. And everybody… not everybody. A lot of people say, well, taxonomy, and they use it to refer to controlled vocabularies, ontologies, et cetera. To them, it’s all taxonomy. It’s not, actually. It is a spectrum. It is a progression of, okay, I’ve got my control vocabulary, then I have hierarchical structure, and then I have synonyms and use-for and all of that kind of stuff. And then you have the ontology. So it’s definitely to my mind, a spectrum. And sometimes you don’t need the full-blown… I agree. You don’t need the full-blown ontology. You can do a lot with a well-structured, well-thought-out taxonomy in an SKOS architecture, and you can really leverage that and you might not need to go the full ontology route. Baby steps.
Larry:
Yeah. Well, I think baby steps, that’s a lesson that comes up everywhere. And that’s it. And some of those baby steps, so how… I’ve talked to a lot of people about this from different backgrounds, and this seems to come really naturally to librarians, or people out of a library science background, that comfort with that spectrum of options for how to organize things is that’s kind of a… I didn’t go to library science school. I’ve hung out with you all a lot. But is that sort of just a part of the mindset of a librarian?
Rebecca:
Yeah, I mean, a lot of people actually become librarians as a second career. When I went to school back in the day, there were people who were lawyers or were in the biological sciences or something like that, but typically, they all had the same sense of wanting to organize things, and also a sense of service, a sense of wanting to help people, in particular. So it’s helping them find things, get what they need, things like that, that tended to be kind of the… I won’t call it the central factor, but the predilections are certainly there.
Rebecca:
And you don’t have to be… being a former librarian or going to library school or iSchool or et cetera, you don’t need to do that to really get into this work, but you have to have kind of be willing to take on that role of the organizer. You have to want to like it. If you have a really carefully organized record collection, then you might want to consider yourself a library. Having that tendency, if you organize all your stuff in your pantry just so, and if all of your spices are… the cuts and organized, that might be a flag there.
Larry:
I love that. Flags like that, that you might be a redneck meme, now there you might be a librarian.
Rebecca:
Librarian is… yes, yeah. Absolutely.
Larry:
Hey, what you just said, I love what you just said about especially the librarians who come in as a second career, that one of the considerations in that is this notion of service. And I’m thinking of every super-helpful reference librarian I’ve ever encountered as the tip of the iceberg on that. But the other thing about that reminds me that I’ve also… I know enough librarians and have hung out with y’all enough to know that there’s a lot going on in the background too. And when you said earlier, that notion of the other ways to use LLMs, besides that conversational interface stuff, is that part of what you were talking about there is this sort of the machinery under the surface that’s kind of helping or structure things? Am I reading that right?
Rebecca:
Yeah, absolutely. Because for example, somebody may want an answer to a very technical problem, but they don’t want to plow through thousands of pages of documentation in order to answer a specific question. But if the user can just ask the LLM, “Hey, what’s the heat tolerance for this particular part in these environmental conditions?” Then you can get an answer, you can get references to that answer, so you can build, so the user knows that, “Okay, that’s where they pulled it from. Okay, that makes sense. Yes, I didn’t have to go through all this documentation.” It saves them time.
Rebecca:
And for many people who are very busy, time is a precious commodity. So looking to save time, all other ways of making things more efficient. Again, back to those business goals, what are we trying to do? How can we help people work smarter and not have to do rote things that don’t need to… that just waste their time, but they have to do it in order to get to an answer.
Larry:
Yeah, that’s it. That one, the efficiency one, that’s an interesting one because as a consultant, I want to get your take on this because the two main benefits of any kind of automation and technology are reduced costs or improved bottom-line revenue. And it’s a lot easier to, when you can show the latter, it’s like, wow, golden. But it’s sometimes hard to show that efficiency gains. Do you have sort of metrics or heuristics or benchmarks that you look at in your practice to illustrate the efficiency?
Rebecca:
Yeah, yeah, we do. We definitely look at amount of time saved, looking at, so for example, if you’re in a sales organization, and I’ve just saved you 10 hours, let’s just say, a week, then you can use that 10 hours to have face-to-face with clients, with customers, et cetera, and sell more. So it’s definitely that time-save, the efficiency, and again, helping people get to that place where they can do their jobs without, again, having to spend a lot of extra time, and honestly things they shouldn’t have to do. They should be used more strategically. Because it’s not about taking away someone’s job, it’s not about that. It’s like, let me help you find better uses for your time as opposed to looking to make you redundant or something. That’s not the point. The point is you’re a smart person and I want to put you in a role where you can do more interesting things with your time, essentially.
Larry:
Yeah. Yeah. And it is occurring to me as you were talking that a lot of this stuff, a lot of this efficiency, a lot of the things you’re talking about trace back to that structured content and the fact that you’ve done that and the fact that there’s metadata associated with it. One of the things that struck me is, because I come out of a similar non-library, but very contenty background, and I’m kind of coming to the knowledge graph world that way, and one of the things I’ve found is the similarities between content modeling and domain modeling and ontology design. Are you finding the same thing for you?
Rebecca:
Well, again, I think it’s, yeah, in essence that, one, it’s a spectrum. It’s not an either-or. It’s a spectrum. So first you need to figure out the basic lay of the land. That’s kind of more of your domain modeling, how do I understand this at a higher level? But I can’t implement based on this domain model. I’ve had some rough-out definitions, I’ve identified key entities, key interactions, but it’s not the lurid detail that I would need for implementation, and that’s where that content, much more specific content model, comes into play. Okay, I know I need to have a structured document. Well, the content model says, “Well, we need a title, and we need a byline, and we need…” and all of that stuff that goes into the detail of it all. Because you need that to implement in order to have that really good structured content. And then you have tools like DITA XML, things like that, that can help you do that along the way.
Larry:
Yeah. Do you work mostly in CCMSs and DITA XML or do you do other kinds of content systems as well?
Rebecca:
It’s a mix. Yeah, it’s a mix. Some DITA-based, some… not standard, but some just content management systems, digital asset management systems. People ask me, well, what’s the best one to work with? And I always say, well, they’re all really good, or they’re all really, really bad. It depends on your requirements. So I can’t give you an answer. You have to figure out what your requirements are, and then pick the tool that’s the best fit for purpose.
Larry:
Yeah, “it depends” is always the consultant’s answer. It’s pretty handy. I use it all the time.
Rebecca:
It’s true.
Larry:
No, it is. But the interesting thing to me about DITA is a particularly good fit for this, it seems like. I know Michael Iantosca has talked a lot about his DOM Graph RAG stuff, and the fact that it’s a W3C standard and fits in nicely. But like you said a minute ago, the model is just the start and the implementation, whatever. There’s so many factors that go into that, the technical capabilities of the org and all that stuff. As you’re modeling, are you sometimes modeling to a specific technical implementation or sometimes… is that almost always the case, or do you sometimes model and then decide on the implementation?
Rebecca:
Yes.
Larry:
It depends.
Rebecca:
It depends.
Larry:
Yeah. Yeah, no, I get it.
Rebecca:
I tried not to say it depends.
Larry:
Thanks. Yeah.
Rebecca:
Yeah. For some clients, they’ve already selected the tool and then I like work… I dislike working with that to a certain extent because sometimes I feel like it’s a round peg into a square hole, because they’ve selected something, they thought it was great, but I’m digging under the… I’m in the belly of the beast going, “Okay, how am I going to make this work?” as opposed to… So ideally I’d like to design things like metadata, schema, et cetera, that sort of… technology agnostic. What does the client really need, and then go from there. But I don’t always have that luxury, to be honest.
Larry:
And I know I’ve seen architectures where the metadata, there’s a semantic tooling adjacent to content systems or something. Sometimes they’re built in. I guess so I can see how that would pan out that way. But underlying that, and I don’t see this, I know that places have it, but I don’t know that they articulate it, the idea of a metadata strategy. That’s always in the back of my mind as I go into these things, it just seems it’s so germane to so many of these activities. How do you think about, because you’re going to need that metadata down the road, how do you approach a project? Do you have something that you call a metadata strategy, or how do you make sure you are… that all the metadata needs in these systems are considered as you proceed?
Rebecca:
Yeah, that’s a really good question. So honestly, I interview stakeholders. I talk to the business. I talk to current content creators about their pain points. I talk to folks about, “Okay, what are we trying to solve? What’s the problem we’re trying to solve here?” Because there are a lot of problems that don’t need machine learning or knowledge graphs. They just need a really good metadata structure if they’re only trying to do certain things.
Rebecca:
And so I talked to them about their needs, what are they going to be using this for? And then it’s reviewing their content or that they’ve currently have plan to create, et cetera, and then pull together, “Okay, this is the metadata I think you need.” So here’s the administrative metadata, which is pretty straightforward, and here’s some descriptive metadata, that the whole thing is to make sure that it’s not overly burdensome, that they… Sure, I have developed extensive metadata schemas for people, and if they promise on a stack of Bibles to use it, but I’m always very concerned about overdoing it, making sure you just have, again, that fit for purpose, the right amount of metadata to get the job done without being overly burdensome because as things go through the tagging workflow and things like that.
Larry:
Yeah, and as you said that it’s just enough metadata, it sounds like is the goal.
Rebecca:
Yes. Mm-hmm.
Larry:
And a lot of that is driven by… Well, it’s interesting. I know I’ve seen some workflows now where they’re using… This has been going on for a while with just old-fashioned machine learning, but sort of are things getting maybe not easier, but are there more things you can do as you’re designing content workflows to facilitate tagging or describing other kinds of metadata?
Rebecca:
Yes. Mm-hmm. Yes. And because you have these tools now that can do text analysis and recommend tags, and then people can say, yep, that’s right, or no, that’s not right. Also, you can enforce governance a little bit more easily now. For example, I had a client years and years and years ago, and they had a knowledge base, and for one of their more very popular knowledge-base articles, which was very, very short, but it had key pieces of information to help the customer support people, they over-tagged so much, there were more tags than the words in the actual article. And then the tags got so noisy that they just stopped using them and started the people entering in the information would fashion their titles in such a way so they knew they could easily bring it up quickly. So yeah, you want to get assistance with automated tagging and stuff like that, which is great, but also you need to enforce governance. You don’t want to create noise that’s going to create problems down the road.
Larry:
As you’re saying that it’s kind of like that thing that a little knowledge can be a dangerous thing. You emphasize the importance of tagging, and people start to perceive that as saying like, oh, I’ve got to get these things everywhere, and then you end up with more tags than content, but also you end up with people doing keyword stuffing in headlines and things like that, it sounds like?
Rebecca:
Oh, yeah. Yeah, yeah, yeah.
Larry:
So that seems like a really germane issue from a number of perspectives, but especially from a usability and readability, legibility, and all the content things. But from a semantic perspective, do you run the risk of diluting or perverting the meaning of that document or that content artifact?
Rebecca:
Yeah, yeah. Yeah, yeah. And then you’ll also have people trying to accommodate for search that might not be properly tuned, so you’ll have, like, you’re not how people are going to search for it. So is it e-commerce, is it ecommerce, one word? Things like that. So they’ll pull in all these variations that in a well developed taxonomy and thesaurus, you won’t have to do that. It’ll just accommodate for you. You have a controlled list of values, things like that, so people aren’t trying to game the system, so to speak.
Larry:
Yeah, that’s the classic place for the combination of a controlled vocabulary and a good thesaurus there and boom, you’re good to go, and that kind of thing. We were talking before we went on the air about… When I talked to Jessica Talisman, she’s talked about the 3000-year history of library science and information architecture versus data science is maybe 50 years old. So are you finding, as you work more, because it sounds like you must be working more with data scientists, data engineers, machine-learning engineers and things like that, what’s the back-and-forth there in terms of understanding the benefits of controled vocabularies and taxonomic schemes?
Rebecca:
Yeah. One thing that’s important, I think, that people might not realize is, at least for when you’re focusing on a particular topic area, like electrical engineering, I’ll say, or even some sort of clothing retail or something, I often use the phrase, “know your collection”, which means I’m not expecting someone to be an electrical engineer. I’m not expecting someone to know how to build oil rigs, but what I’m expecting is that they’ll understand the basic concepts of a particular subject area. And because in terms of data analysis and things like that, the things that are coming back, you have to be able to critically ascertain, okay, does this really make sense in this context?
Rebecca:
I’m not just saying that you have to go out and get a master’s in something or a bachelor’s in something, but go to YouTube, read a few books, learn a little bit about what you’re working with so you can more effectively evaluate what’s coming back. As a taxonomist, as a consultant, I’ve worked with financial services, electrical engineering, wine and liquor sales, jewelry, home improvement, hardware. And I’m not saying that I learned all that there was to know about a particular topic area. What I’m saying is I do my best to ask good questions and try and really understand what I’m looking at, right?
Larry:
Yeah, that’s really-
Rebecca:
And I think that’s super important.
Larry:
I agree. And it’s reminding me that we tend in our heads to just think subject matter experts here, ontologists and taxonomists and engineers over here. But in fact, they can work really better together if they each learn a little bit about the adjacent domain, get a little bit of that subject matter expertise. I’m thinking now in the ontology world lately, there’s been a lot of talk about I need tooling to help me better work with my subject matter experts to help them see how the classes and concepts and things are… how their subject matter expertise is represented in the system, to kind of verify that, yeah, that’s correct. That’s really it.
Larry:
Is there tooling that you like for that? Like you just said, it’s old-fashioned tooling. You go read a couple of books, watch some YouTube videos, but in terms of the flip side of that, of helping to get the subject matter experts at least a little bit acquainted with how you’re operating, like you said earlier, they think of it’s taxonomy, that’s just like they think all of this stuff. Do you have to tease out the different practices and…
Rebecca:
Yeah. Yeah. And actually, with one of my current clients, when we do subject matter review, and we take three slides, I think it’s three, yeah, three slides, and we make them a taxonomist. So three, four slides, we walk them through the conceptual, this is our approach, this is how we think about things, this is how we think about taxonomy, this is how we think about attributes, and here’s some examples of how you can leverage this, here are the top kind of guiding rules for category creation, these are the guiding rules for attribute creation, we talk a little bit about hierarchy. And oftentimes they’ll ask really good questions. So for example, we use the example of a very small world of men’s clothing, and some intrepid soul said, well, what about jumpsuits? Where do they go? And things, so they’re really starting to think about this stuff. And then we can dive into the details of their particular category of content or products.
Larry:
Nice. If those three slides are available anywhere, I’d love to put those in the show notes. If it’s proprietary, no worries at all. But yeah. Hey, Rebecca, I can’t believe it. We’re coming up close to time already, but before we wrap up, I want to make sure, is there anything we haven’t got to yet that you want to make sure we share or something you want to revisit from the conversation, or…
Rebecca:
That’s a good question. I think the big thing is be willing to take the time to clean up your content, clean up your data, in order to make a successful AI application or whatever you’re trying to do. Understand why you’re trying to do it, and don’t be afraid to talk to people. People like to talk about their work. People like to talk about what they do. And ask questions. Be willing to learn a bit.
Larry:
Nice. Yeah. As the host of an interview-format podcast, I wholeheartedly applaud that last version. Hey, one very last thing, Rebecca. If folks want to connect or follow you online, what’s the best place to find you?
Rebecca:
LinkedIn, Rebecca Schneider. Feel free to contact me through email, which is rschneider, R-S-C-H-N-E-I-D-E-R, at avenuecx.com.
Larry:
Great. I’ll put those in the show notes also.
Rebecca:
Okay, great.
Larry:
Well, thank you so much, Rebecca. Really enjoyed the conversation.
Rebecca:
Oh, thanks for having me.