OpenAI Models Used Public Census, SEC Data, Bloomberg Reports
Fintech

OpenAI Models Used Public Census, SEC Data, Bloomberg Reports

5h ago

OpenAI's artificial intelligence models made use of publicly available information from two major US government sources, according to a Bloomberg News report carried by Reuters. The material involved records published by the Census Bureau and the Securities and Exchange Commission.

Both agencies release substantial volumes of information to the public. The Census Bureau publishes statistics on population, households and business activity, while the SEC maintains filings and disclosures from listed companies. That material is freely accessible and is routinely used by researchers, journalists and analysts.

Its presence in AI training data reflects a broader industry practice. Developers of large language models typically assemble datasets from web crawls, licensed archives and other repositories. Government records are attractive because they are structured, authoritative and openly available, and because they can be downloaded at no cost.

The report nonetheless points to questions regulators and rights advocates have raised about how AI systems are built, including whether the origins of training data are clear enough and whether public records deserve different treatment from other content. How companies document and handle such sources is likely to remain under scrutiny as the technology spreads.

Source: Reuters