Never mind clean data. Annotate as you collect it.
CIO.com
ne of the themes I keep coming back to is how much of the tedious, rigorous work that we should have been doing all along but didn't always have the incentive or the budget for is now necessary for AI to work properly. Quotas, documentation, lineage, authority, sensitivity labelling: they might be nice to have for humans who will be responsible enough to probably figure things out if they're not there, but they're imperative if you want what you spend on the flashier bits of AI to be worthwhile.
Data for AI normally gets annotated (manually or automatically) when you make a training dataset; while I was researching this, I wondered how many of the AI pilots that failed in production were failing because that extra information just wasn't there in the data they get fed in production. Here's another thing we should shift left, automate as much as possible and turn into a virtuous cycle.
Alas, not everyone I talked to for this piece is on Bluesky; the founder of SurrealDB talked to me about how they attach so much extra information to unstructured data that it gets some structure. there's one school of though that says any document is at least semi-structured and XML creator Jean Paoli told me about the DGML spec Docugami is open sourcing, which lets you tag not just a document but an object inside a document with metadata. And Ulik Hansen gave me some great examples of where this is already routine, because it's so useful.
Never mind clean data. Annotate as you collect it.
Leaving a breadcrumb trail from the original context of the data you use for AI might let you trace a single bad prediction to the source.
If you find this piece interesting, I've written about the dangers of overcleaning data and removing context before.
Article
I've also looked at synthetic data: often useful, but sometimes just too clean and sometimes not clean at all.
Article
- data cleaning
- shift left
- context
- metadata governance
- annotation and labelling
- data provenance and lineage
- data quality
- IoT digital twins
- DataHub
- DGML
- context engineering
- fine tuning
Did you enjoy this article?
Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.