It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. What’s the way forward to ensure the path of least resistance?
I've been saying for a while that synthetic data is useful and in some situations maybe irreplaceable.
For this piece though, I wanted to dig into what can go wrong.
Me on synthetic data for @cio.com: It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. Already widely used in medicine, it's coming to business...
Issues with synthetic data turned out to be quite topical.
yeah; I just did a whole piece on what can go wrong with synthetic data and simulation like this; if you get the population representation right, it can be *less biased* and way less intrusive than instrumenting humans but the expertise to get is right isn't well distributed
synthetic data
data bias
privacy
simulation engines
oversampling
differential privacy
iterative validation
storage