If you’ve posted anything online, chances are OpenAI might have used your words to train ChatGPT. The chatbot has had a meteoric rise in the tech world and has proven helpful for just about anyone. But its undeniable prowess has begged the question of where it learned all of its tricks.
As it turns out, even OpenAI doesn’t fully know what exactly ChatGPT was trained on. With so much uncertainty swirling around the artificially intelligent chatbot, a group of anonymous individuals has just issued a class action lawsuit claiming OpenAI has violated privacy laws.
The filing stated that 300 billion words from the internet had been scraped to teach the machine. So yes, everything from Wikipedia pages to your Facebook post from 2009 may have played a role in training the AI.
“Despite established protocols for the purchase and use of personal information, Defendants took a different approach: theft,” reads the filing. “They systematically scraped 300 billion words from the internet, ‘books, articles, websites, and posts—including personal information obtained without consent.”
The lawsuit alleges that OpenAI has failed to adhere to the correct procurement guidelines before taking content from those who have created it.
While this violates privacy against users’ intellectual property to some degree, the case could sway either way, given the finicky nature of the internet. As it stands, whatever someone posts online becomes part of the platform it is hosted on. Whether or not everyone gets paid their due for providing free teaching materials to a chatbot is still very unclear, as governments and individuals are only now starting to wade through the complications of AI and privacy.