Salesforce CEO Marc Benioff has shared his concerns over how GenAI companies leverage data to train their large language models (LLMs).
During an interview with Bloomberg at the World Economic Forum, he stated:
"I own TIME, then there is Bloomberg and the New York Times. We’re all finding our intellectual property, your stories, your work… surfacing in these results because all of the training data has been stolen."
The CEO then noted that this data is one of three critical tiers in developing GenAI apps, including those created by Salesforce partner OpenAI. Indeed, Benioff stated:
Download the OpenAI or Copilot app. [There is a] commodity UI on the front-end, in the middle you have what is becoming somewhat highly commoditized large language models… [and] then you have this broad set of training data, which is the third tier of it, which has been basically ripped off.
While Benioff caveated that by highlighting how nobody knows the fair price for that training data, he suggested AI companies shouldn't just continue doing what they want.
Instead, he proposed there is an opportunity to build a standardized set of training data that allows these companies to play a fair game and ensures content creators receive fair compensation for their work.
"I think that bridge hasn't yet been crossed, and that's a mistake by the AI companies," he concluded.
The views of Sam Altman, CEO of OpenAI, unsurprisingly clashed with Benioff’s remarks. Instead, Altman played down the significance of training data within its applications during his own session with Bloomberg.
"There is this belief held by some people that you need all of my training data, and my training data is so valuable, he said. "Actually, that is generally not the case.
"We do not want to train on the New York Times data, for example, and - more generally - we're getting to a world where… you're going to run out of that data at some point anyway.
So, a lot of our research has been [focused on] how can we learn more from smaller amounts of very high-quality data? And I think the world is going to figure that out.
Altman also painted a picture of how ChatGPT may soon respond to users by citing content from various publications and providing snippets or “probably something cooler.”
Benioff on AI Regulators and the Evolution of Contact Centers
Alongside his critique of AI companies, Benioff applauded the enthusiasm of international governments to regulate the technology – citing the UK Safety Summit.




