Synthetic data’s fine line between reward and disaster
CIO.com
t can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. What’s the way forward to ensure the path of least resistance?
I've been saying for a while that synthetic data is useful and in some situations maybe irreplaceable.
For this piece though, I wanted to dig into what can go wrong.
Issues with synthetic data turned out to be quite topical.
Synthetic data’s fine line between reward and disaster
It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. What’s the way forward to ensure the path of least resistance?
- synthetic data
- data bias
- privacy
- simulation engines
- oversampling
- differential privacy
- iterative validation
- storage
Did you enjoy this article?
Recommend it — Standard Reader surfaces well-loved writing to more readers across the network.