Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

As a 10-year-old photographer, Margaret Warren would jot down on the back of each printed photo metadata about who took the picture, who was in it, and where it was taken.
Her interest in image metadata continued into her adult life, culminating the creation of ImageSnippets, a service that lets anyone add linked open data descriptions to their images.
We talked about:
- her work to make images more discoverable with metadata connected via a knowledge graph
- how her early childhood history as a metadata strategist, her background in computing technology, and her personal interest in art and a photography shows up in her product, ImageSnippets
- her takes on the basics of metadata strategy and practice
- the many types of metadata: descriptive, administrative, technical, etc.
- the role of metadata in the new AI world
- some of the good and bad reasons that social media platforms might remove metadata from images
- privacy implications of metadata in social media
- the linked data principles that she applies in ImageSnippets and how they’re managed in the product’s workflow
- her wish that CMSs and social media platforms would not strip the metadata from images as they ingest them
- the lightweight image ontology that underlies her ImageSnippets product
- her prediction that the importance of metadata that supports provenance, demonstrates originality, and sets context will continue to grow in the future
Margaret’s bio
Margaret Warren is a technologist, researcher and artist/content creator. She is the founder and CEO of Metadata Authoring Systems whose mission is to make the most obscure images on the web findable, and easily accessible by describing and preserving them in the most precise ways possible.
To assist with this mission, she is the creator of a system called, ImageSnippets which can be used by anyone to build linked data descriptions of images into graphs. She is also a research associate with the Florida Institute of Human and Machine Cognition, one of the primary organizers of a group called The Dataworthy Collective and is a member of the IPTC (International Press and Telecommunications Council) photo-metadata working group and the Research Data Alliance charter on Collections as Data.
As a researcher, Margaret’s primary focus is at the intersection of semantics, metadata, knowledge representation and information science particularly around visual content, search and findability. She is deeply interested in how people describe what they experience visually and how to capture and formalize this knowledge into machine readable structures. She creates tools and processes for humans but augmented by machine intelligence. Many of these tools are useful for unifying the many types of metadata and descriptions of images – including the very important context element – into ontology infused knowledge graphs. Her tools can be used for tasks as advanced as complex domain modeling but can also facilitate image content to be shared and published while staying linked to it’s metadata across workflows.
Learn more and connect with Margaret online
IPTC links
- IPTC Photo Metadata
- Software that supports IPTC Photo Metadata
- Get IPTC Photo Metadata
- Browser extensions for IPTC Photo Metadata
Resource not mentioned in podcast
(but very useful for examining structured metadata in web pages)
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 21. Nowadays, we are all immersed in a deluge of information and media, especially images. The real value of these images is captured in the metadata about them. Without information about the history of an image, its technical details, and the context in which it originated, it might as well be a piece of random clip art. Margaret Warren is the creator of ImageSnippets, a tool that helps you create and manage the metadata that imbues your images with meaning.
Interview transcript
Larry:
Hi, everyone. Welcome to episode number 21 of the Knowledge Graph Insights podcast. I am really happy today to welcome to the show my friend, Margaret Warren. Margaret and I helped co-organize a weekly event called The Dataworthy Collective, which is a collaborative learning environment, but she’s better known as the CEO and the CTO of a company called Metadata Authoring Systems, which is best known for the product, ImageSnippets, which does metadata management for photos and other kinds of images. Welcome, Margaret. Tell the folks a little bit more about what you’re up to these days.
Margaret:
Hello! How are you doing? Yeah, it’s a good introduction, thanks. Well, what I’d like to say about Metadata Authoring Systems is just that we want the most obscure images to be findable and easily accessible by describing and preserving them in the most precise ways possible. That’s kind of what we try to do. One of the tools that we created is called ImageSnippets. It’s been around for quite some time. ImageSnippets is, I think it can best be described as a way to take image description data itself and parse that into knowledge graphs, which is not something that people necessarily think of that much. It adds a whole new dimension to that content. You have an image, and what is the metadata that’s attached to that image? How do you turn that metadata into a graph? Then, what can you do with that graph once you have it created?
Larry:
Yeah, no, that’s super interesting too. A lot of people would back up from that. Like the use case, the end use case, it would be like, “Well, I want that picture of my daughter’s 4th birthday party,” or this specific product piece, or that famous Porsche engine manifold, the thing that you’ve example-
Margaret:
That I show that a lot.
Larry:
Exactly, yeah. Well, it sort of gets in, but I guess, I definitely want to focus on the use cases, but I think there’s a few people, and even people who are versed in metadata from one perspective or another, may benefit from a bigger picture view of how you see the world of metadata. You come at it from image management in particular, but you’re really a metadata strategist and management person at heart. Can you talk a little bit about what it is and why it’s important?
Margaret:
Yeah. Well, I’ll lead into this by talking a little bit about my background, because it helps to set the stage. I grew up with a photographer, my father was a photojournalist. In the 70s, he was working in a newspaper and I was growing up with this, and I was sort of the metadata person in the family without really understanding that exactly what it would lead to. I would go out and make photos. We were using Nikon film cameras then and everything, and we would go back to the dark room and develop the prints and everything. I would take these, I have a classic example that I used visually, but it’s like I have the metadata written on the back of the photo. It’s really cute. I was 10 years old, and it’s like, Who’s in the car? Where we were? Who took the picture? Who developed the picture? Who printed the picture? I even got into the administrative metadata. All this seemed important to me.
Margaret:
My mother used to write a little bit onto the slides and stuff a little bit, but somehow I just had this knack and I was very interested in this. Now, that continued on over the years because I’ve always also been an artist and a photographer and a technologist. I was in the Coast Guard and I was an electronics technician, and then I started working on computing systems and doing programming, and ultimately, had my own company doing hardware and networking support and building networks and deploying software solutions, lots of things. Then, I was trying to publish my images on the web and it’s super interesting to me that the problem that I was trying to solve about publishing my images on the web almost 30 years ago is still, well, actually, I created a solution that solves it with ImageSnippets, but it’s still something that is not really mainstream.
Margaret:
The entire vision that had to do with metadata and everything was sort of hijacked by the centralized companies that were just so data hungry and so … well, engagement hungry rather like we want you to engage constantly, so we’re going to make it increasingly easy for you to share your images as fast as you possibly can without thinking very carefully about the metadata aspect of what is in these images, why you’re sharing it, why it’s there. This is such a deep, deep, deep subject.
Margaret:
To answer your question about what is metadata, I often say that metadata really ultimately is data about anything that you’re talking … I can hold up this coffee cup and say, “Oh, I can describe the coffee cup. It’s got flowers on it, the size of it, the shape of it, how warm it is, the temperature,” anything that I really say about anything is metadata. But when it comes to, say, images that you’re going to put on the web, and I like to narrow it down to images because that’s really my forte, still images, but you can have, of course, metadata on PDF files and audio files and video files and everything can have metadata attached to it.
Margaret:
Inside of images, you can add headers and other types of information that would contain things like the EXIF data. A lot of people are very, very familiar with EXIF data, which is, that is the data that usually comes from hardware. That will come from a camera or it will come from a scanner or some piece of hardware that will say what the settings were. Sometimes it will give the GPS data, if you take an image with your smartphone, and it will include the GPS data unless you turn it off, of course. There’s reasons why people want to do that, which are understandable, but we’ll come back to that in a minute.
Margaret:
There’s lots of different types of metadata. There’s descriptive metadata, there’s administrative metadata, there’s technical metadata. You can break metadata down in so many different ways, in so many different facets, but it really just comes down to what is it? Maybe, where was it made? Then, you can get into more descriptive metadata, like all kinds of things. The IPTC, the International Press and Telecommunications Council, of which I’m a member of the photo metadata working group, is responsible for setting a standard around photo metadata that is generally followed by many organizations to look at this as an open standard.
Margaret:
There’s a lot of potential fields in that you can get very, very deep into describing metadata. Most recently, we just added a field for data mining, which is something to attempt to put something out there that potentially could be the starting point for conversations to be had around AI training, when you might train images and when you want to prohibit training, opt out strategies. This also gets very deep into another whole thing. We also have things like digital source type, which is, describes the source of the image. Is it generative AI or was it made from scanning a negative or is it a digital capture? That sort of thing. Anyhow, this data gets very complex. I am mostly and often concerned with the actual description of what you see in the image when you look at it.
Margaret:
That is the information … Actually, we parse everything into triples, into RDF triples that are stored in a triple store and also shared with the image as well. I’m very concerned with, say, taking the context of an image and taking that description and turning that into the knowledge graph. Let me give you a really good example of this, because a lot of people think, “Oh, we can do everything with just AI, can look at an image, describe it, give it captions, give it everything.” One example I’ve been sharing a lot lately is that there’s a terrific collection of images from the Stennis Space Center, NASA’s Stennis Space Center, which is down in Mississippi, and those are in the internet archive. We can import images from the internet archive and then describe those images and create metadata about them and linked back to those images.
Margaret:
A lot of those images have incredible metadata on them. The person who uploaded this collection was an awesome curator. There would be a picture, for example, of a groundbreaking ceremony, I’ve shared this before on LinkedIn. There’s a ceremony of a groundbreaking, there’s all these people there, and a lot of them are looking down. They have on their hard hats, the ceremonial hard hats and the ceremonial shovels, and they are standing in front of this construction site. I can give that image to ChatGPT, to a multimodal AI system, and it can tell you a ton of things about that. It can tell you that it’s a groundbreaking ceremony. It knows that, it can tell you that there’s construction in the background and it can identify the type of day it is and the fact that there’s people standing around. But of course, it doesn’t know any of the people who are in it. Now, a facial recognition system might, if it could see the faces real clearly, but it can’t.
Margaret:
But in this case, a lot of them are covered a little bit by hard hats and stuff, but the metadata of that image contained embedded in it an entire description of exactly who the people were in that image. Some of them were Congress people, and it tells you what the site was being built for, that it was for A3 test stand, there were the leaders of the companies that were building it. It gives you a ton of context, and in my mind, that’s what’s really important about that image. That’s what’s really going to be important about finding that image, not that you can identify that it’s groundbreaking. It’s great that a multimodal system can identify this generic information so easily and so well right now with just using pixel data and all of the combinations that have come out of machine learning over the past two decades, but I still feel that the provenance, the real context, that’s the metadata that I really think is important. I know I went off on a little bit-
Larry:
No, no. I got to say, I love that because there’s so much, you interwove all the stuff I asked the first couple of questions I had about what it is, why it’s important, and then got into some of the technical manifestation of it. I’m so curious about, you mentioned the EXIF format and the fact that those NASA images had metadata associated with them. How much of this travels with the image natively? I’ve heard you talk in other contexts about how some social media platforms strip out metadata and leave it there, and then ImageSnippets is a way to ascribe metadata after the fact, or you get a photograph and then you add. Are you extracting some of that technical metadata out of the image itself with ImageSnippets? Is it all like a human-tagging endeavor?
Margaret:
Yeah. Well, let me answer one thing at a time. Okay. Yes, I do want to want to emphasize one of my pet peeves in the entire world that I’m very passionate about is that social media, and I think I touched on this a second ago, but I got distracted by my rabbit hole. The data-hungry, engagement-hungry machines, they immediately started stripping metadata. Now, you can argue about all the reasons behind this. There are multiple factors. Sometimes you can sit there and look at, well, why would Facebook or anybody delete all the metadata? Well, some of it could have to do legitimately with file size at the time. When you have the billions of images that they have, those extra bytes, of course, can take up a lot more disk space. But also, there’s this whole possible deniability factor.
Margaret:
There are bad actors who will do things with metadata. There are stalkers, there are people who are having domestic issues. There’s just bad people. There’s pedophiles. They want to use, they want to look at embedded metadata if they can, to try to find clues of where this image was made, so they can go find somebody. I find this curious because we are in this age where people want both anonymity and transparency at the same time, which makes it really difficult.
Margaret:
I like to use this example. You can ask people, “Do you want us to remove your metadata from your image?” And they’ll say, “Oh, yeah, absolutely. Take it out. I don’t want anything about where I’m at, where my actual location is, left in my image.” But then, in their post, they will tell people exactly where it is and like, “Oh, here we are hanging out at the hotel in San Antonio. We were having cocktails last night.” It’s like, well, okay, you want it both ways, but you don’t really want to think about it.
Margaret:
I do feel this is a mindset that the engagement-hungry, centralized platforms have nurtured, is this idea that you can share anything that you want and as fast as you want without having to think about the responsibility of why you’re sharing it. Maybe it sounds really old school to say this, but I always felt, and I’ve said this in other places before, I’ve always felt I typically don’t share things online ever unless I feel I was comfortable standing in the middle of a mall or somewhere in a shopping plaza on a Saturday afternoon telling anybody that would walk up to me those things. That’s what I share online.
Margaret:
I feel these companies have stripped away the personal responsibility that we could have for just being more aware, having more awareness of what we’re sharing and why and being a little more careful about it. But of course, that’s not what everybody wants. Everybody wants to put more and more imagery, more and more details of their lives, more and more stuff out there so that they can get engagement. I have to say, I am sure this is … maybe it’s a little bit of a boomer kind of take. I’m still blown away by just the amount of personal information that people share on TikTok and YouTube all the time.
Larry:
As you’re talking about that, it’s really interesting to me like that example of you take a picture of a hotel lobby, and in one context, you’re posting it on social media and saying, “Boy, we had a great time in San Diego last weekend.” But a social media platform is probably harvesting that because it’s tagged as that location and using it in some way with that hotel and somehow using it as user-generated content for that. I’m curious about, you mentioned earlier different types of metadata, but I’m wondering if there’s different, and you’ve mentioned context, and now I’m thinking of that convergence of types of metadata and context like in one situation, you want to know what the place is, but in another you don’t. How do people manage that? How do you manage the appropriate contextualization for the use case you’re in the middle of while at the same time keeping that knowledge-ey connectability of a generically-tagged photo? Does that make sense, that question?
Margaret:
I think so. I think so. I’m not sure.
Larry:
Because the photo is this one thing, and it’s got some kind of metadata attached to it, and then, the hotel might use it in one way, the social media platform in another way, and you might use it in another way in your social media feed.
Margaret:
Yeah. I think I get what you’re saying, and I realized I got off on the social media tangent there for a second. In our case, with the tool, and I think this is what you were asking a minute ago as well. As an architecture, ImageSnippets is built around linked data principles, meaning that all of the data itself is stored as URIs, and then we can import images using their URIs. Then, when we import the images, we import all the metadata that was embedded in that image if it was there, and of course, there’s a lot. When you talk about workflows, this is the one thing that ImageSnippets has been striving to do, is to have this consistent travel through workflow, where the metadata stays attached to the image in multiple ways, and it also gets written into the knowledge graph, so that’s all consistent.
Margaret:
We import and we can import from a URI or you can upload into it, but once you do that, whatever metadata was there, so let’s just suppose you were a photographer who is very metadata-aware and you use the metadata, the file info panels and things like Adobe Photoshop or other utilities that let you deal with IPTC metadata. Just as a side note, there’s a wonderful site, IPT site that lists softwares that really conform to the standards and how much they use the standards. If you’re curious about finding a whole range of software and how they interoperate with each other, there’s a whole set of software tests that show the flow from one piece of software to another piece of software to another piece of software. If you need your images to travel through different workflows, which ones will try to keep that the IPTC and the EXIF data consistent without deleting it. There’s a lot of tools that you could do to test that.
Margaret:
But anyhow, we import that, all the metadata we can. Then, we are very keen on, of course, as I said, taking that context, which is what’s maybe written in the description, and there’s a lot of different descriptive fields that you can work with. There’s also alt text for accessibility, for example. That is very handy to have that travel with images from one environment to another. There’s constantly initiatives to try to get some of the social media companies to do better, maybe like Bluesky or, I’m trying to think of another one. Well, even WordPress, like trying to go from other content management systems. When you import an image into there have, if there was existing alt text data that maybe got put in a data asset management system on a person’s local computer, that it would actually travel all the way online and through all the way to whatever website that it has gone to.
Margaret:
Almost everybody strips it, it’s ridiculous. LinkedIn strips it, Substack strips it, I could go on and on and on. That’s another thing you could see on the IPTC site. You could see all of the social media companies that strip it. The last one that actually used to do really well at this, in my opinion, a big one, it would be Flickr, because Flickr tended to respect whatever was in an image. But anyhow, except for ImageSnippets, of course, which we keep it all in there.
Margaret:
Now, you can also edit that metadata. Let’s suppose you spelled a name wrong or something, and then you get it online and you find out, “Oh, okay, I spelled a name wrong.” Okay, you can go in and edit that and re-save it back to the image, as well as saving it in the triple store. Talking about that context again and how we turn that into triples, we have a lot of different tools that go along with ImageSnippets that you can call to assist a user in creating the metadata as fast as possible, and soon to be conversational type agents that can help, or not agents. This is going to be a controversial word in 2025.
Larry:
That’s right. We got to be careful with the buzzwords, yeah.
Margaret:
Let me not say that word. No, no. Okay. Because that’s already being really misused already. We also can call a utility that will look through the text and find in the data sets that when you’re using ImageSnippets, you can look up terms that you would typically use as your keywords, but you can look those up and match those, so you can entity match to a term in a linked data set, like in DBpedia or Wikidata, or you can create your own, if you can’t find it. If it’s somebody that you know personally that isn’t in Wikidata, that you’re going to talk about frequently in your images, you can create an entity for that person or that place or what have you. Then, now you’re able to write everything as triples as just a very simple subject and predicate and object. Basically, this image depicts the Eiffel Tower, and then each one of those parts of that becomes a URI.
Margaret:
This image is the JPEG itself. There’s a relationship between the image and the keyword, and then you have the keyword, which is Eiffel Tower. We have a small set of terms that was worked out between myself and my colleague, Pat Hayes and some other people that did a lot of testing in between as we were developing it, but it’s a lightweight image ontology. That is the small ontology that we use to do relationship between what would normally be the keyword and the image itself. This is how we’re building name graphs around images. Then, they’re all part of the big overall image graph, and so then, you can do so much with that. Like we said, everybody’s always quoting Jim Hendler. I love quoting Jim, he’s such a great person. When he said “a little semantics goes a long way,” because it really does.
Margaret:
When you’re able to describe an image using some term like, oh, let’s say if I say a lighthouse. Because of the SCS concepts, the subject concepts in DBpedia, I can also find that lighthouse by doing a search now for Egyptian inventions and all kinds of other things that I didn’t have to know. Let me just rephrase this. As a keyworder, you may or may not know certain things or know all the potential subclasses or subjects that something might be related to. You may not think to put those into your keywords that you would be attaching to that like normal keywords. Like Egyptian inventions, I would never think to put Egyptians inventions on looms or forks or lighthouses, but they were all Egyptian inventions. That’s a perfect example of Jim’s saying, “a little semantics goes a long way.” You get a whole lot of machine inference and machine reasoning from just using one term from a linked data vocabulary.
Larry:
It’s so funny, I’ve heard Jim’s quote a thousand times over the last, I don’t know, 20 years, but the way you just said it, that really ties it up nicely that a little bit of semantics, we identify this thing as a lighthouse, and because there’s an ontology that knows that lighthouses were invented in Egypt 4,000 years ago or whatever, you can attach that. But hey, Margaret, I can’t believe it. We’re coming up close to time already.
Margaret:
Yes.
Larry:
These things always go way too fast. But hey, before we wrap though, I want to make sure, is there anything last that you want to talk about? Is there anything you want to revisit from the conversation or any last little tidbit you want to leave us with?
Margaret:
Well, I guess the main thing is just that what I believe is in the future, a lot of people think, “Well, how much does it cost to put all this extra metadata on all these images?” Here’s what I have to say about that, which is that I think that in the future, especially where we’re headed with generative AI images, and there’s so much more that could be said about that, but I really truly believe that original, accurate, provenance-supported, long-tailed data, long-tailed content is really going to get increasingly more valuable, because it’s going to be harder to find, and so the items, any sort of resources that you have that have more of that type of metadata that supports the provenance, that supports the originality, supports the context of what is in that content, I think that’s going to just become increasingly important, because it’s just too hard to find otherwise, and those are the things that are going to be valuable.
Margaret:
Even now, it’s hard to find. Suppose you want to find a part for your tractor or something, or your overhead door in your garage or something like a little part to fix it, it’s just a real struggle because everything you’re presented in Google searches are for products that most of them don’t necessarily match what you’re looking for and so on. I just think the more metadata the better.
Larry:
I love that you ground this thing that a lot of people think it was this highfalutin, like the super-complicated technical thing, it’s like, “Nah, that’s the thing that’s going to help you find that little hinge for your garage door opener or whatever.”
Margaret:
Exactly.
Larry:
Yeah. No, that’s perfect.
Margaret:
Exactly.
Larry:
Hey, one very last thing, Margaret. If folks want to stay in touch or connect with you online, what’s the best place to find you?
Margaret:
Yeah. Well, I’m on LinkedIn a lot, so Margaret Warren, ImageSnippets, I think it’s Margaret Warren IS or something, MargaretWarren.us. Of course, you can go to imagesnippets.com. My fabulous domain name is Metadata.rocks, so metadata rocks. But yeah, that’s a lot of the places probably,
Larry:
Okay. I’ll put all those in the show notes as well.
Margaret:
I’m on Bluesky. Yeah, I’m on Blue Sky.
Larry:
Oh yeah, Blue Sky, yeah. Okay, because we’ve all got our-
Margaret:
I actually have accounts in Facebook and Twitter, but I try to use those as little as possible. I’m not real happy with those services at the moment.
Larry:
Understandable, yeah.
Margaret:
Yeah. They’re just too, it’s a lot of garbage no matter how you, yeah.
Larry:
Well, thanks so much, Margaret. We talk all the time, but this was like, I love that we got to just focus for a half hour on metadata. It’s a treat, so thank you so much.
Margaret:
Well, thank you. I hope I didn’t go too off into too many rabbit holes. Sometimes I have a tendency to do that.
Larry:
Here, it’s all about the rabbit holes, so thank you.
Margaret:
That’s good. Okay. All right. Thank you.