It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. What’s the way forward to ensure the path of least resistance?

I've been saying for a while that synthetic data is useful and in some situations maybe irreplaceable.

Mary Branscombe's avatar

you can't give a language model actual invoices from a real company because they contain what can be either personal or confidential information but you want it to know what a really wide range of invoice types look like. synthetic data is how you do that (Microsoft did this for Phi3 for example)

Mary Branscombe's avatar

Also check out the Microsoft research papers from a couple of years back about how they used synthetic data to train the HoloLens models on a wide range of human body shapes and motions, they put out a lot of details and released tools to do it

For this piece though, I wanted to dig into what can go wrong.

Mary Branscombe's avatar

Me on synthetic data for @cio.com: It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. Already widely used in medicine, it's coming to business...

Issues with synthetic data turned out to be quite topical.

Mary Branscombe's avatar

Synthetic data is *really good* when it’s good (representative etc) but it has a little curl right in the middle of its forehead…

Mary Branscombe's avatar

yeah; I just did a whole piece on what can go wrong with synthetic data and simulation like this; if you get the population representation right, it can be *less biased* and way less intrusive than instrumenting humans but the expertise to get is right isn't well distributed

Mary Branscombe's avatar

I was just saying in a piece last month that fear of data breaches will drive use of synthetic data (so we had better get good at making it truly representative of actual data)

Synthetic data’s fine line between reward and disaster
It can be faster, easier, cheaper, more representative, and more privacy preserving to generate data than to collect it. But done wrong, synthetic data can double down on everything you’re trying to avoid. What’s the way forward to ensure the path of least resistance?
https://www.cio.com/article/3986687/synthetic-datas-fine-line-between-reward-and-disaster.html
  • synthetic data

  • data bias

  • privacy

  • simulation engines

  • oversampling

  • differential privacy

  • iterative validation

  • storage