A new study produced by Stanford University has revealed that most major AI models do not comply with new regulations in the EU AI Act which is set to pass through it’s final legislative stage soon.
The European Parliament has adopted a draft AI Act which will then be deliberated by national bodies before it can be formalised into law.
Researchers at Stanford University investigated some of the most popular AI models, from OpenAI’s ChatGPT to Meta’s LLaMA, to see if they complied with the proposed legislation.
Across the board, the ten investigated models all failed to some degree to be fully compliant, meaning the EU legislation would lead to major changes across the use of genrative AI large language model creation and rollout.
The Research
The ten models investigated included: GPT-4 (OpenAI), Cohere Command (Cohere), Stable Diffusion v2 (Stability.ai), Claude (Anthropic), PaLM 2 (Google), Bloom (Big Science), LLaMA (Meta), Jurassic-2 (AI21labs), Luminous (Alpha Alpha), and GPT-NeoX (Eleutherai).
Researchers gave each model a point score based on their compliance under several key categories, including data, compute, model, and deployment.
Under data, AI models would have to describe the data sources used to train their foundation model, use data that is subject to data governance measures like suitability, bias, and appropriate mitigation to train the model, and summarize any copyrighted data used to train the foundation model.
For compute, the companies would have to disclose the model size, computer power, and time used to train the foundation model, as well as measure the energy consumption used in training.
Under the model category, companies had to take several factors under consideration, from describing capabilities and limitations for the foundation model, and describing any foreseeable risks and mitigations.
Further, they would have to benchmark the foundation model with public and industry standards, and report the results of internal and external testing on the model.
Under deployment, models would have to disclose when content from a generative model is machine generated and non human generated, disclose to EU member states where the foundation model is on the market, and provide sufficient technical compliance for downstream compliance with the EU AI Act.
Based on these requirements, researchers gave each model a score from zero to four for each requirement based on their level of compliance.
Results
The results provided a wide range of compliance, with some providers scoring less than 25% compliance, including AI21 Labs, Aleph Alpha, Anthropic) and only one provider scoring 75% compliance (BigScience).
BigScience’s Bloom was the only AI model to have only one zero score, in the matter of making clear to EU member states where in the market their AI was – companies across the board struggled with this section as well, with only GPT – 4 (OpenAI) and Claude (Anthropic) scoring two, and Luminous (Aleph Alpha) scoring one.
Four areas the research identified with persistent challenges and some of the poorest scores were copyrighted data, compute/energy, risk mitigation, and evaluation/testing.
In the case of copyrighted data, this has been an issue since the prolific rise of generative AI models like ChatGPT and many art generators. Few providers disclose information regarding the copyright status of training data, especially since most train on data sourced from the internet, which has a large proportion of copyrighted data.
The legal ramifications for training on copyrighted data has yet to be decided, but the EU would require companies to at least disclose the copyrighted date they use in training the models.
Uneven energy reporting was also a noted issue – this could be, as the paper noted, that reporting energy consumption for training foundation models is currently contentious. Therefore, reporting, even if done, can prove unreliable.
The research also revealed that companies are not adequately disclosing their risk mitigation protocols to the EU. The EU has classified certain applications of AI models as ‘high-risk’ such as those that are for the targeting of political ads, and even those used for hiring purposes.
A Time investigation purported that OpenAI lobbied the EU to exclude their general AI model from being classified as ‘high risk’ which would require it to follow more stringent rules regarding transparency, explainability, and responsibility.
The report found that relatively few models share how they mitigate risks associated with AI, and the efficacy of these mitigations.
Further, the research discovered an absence of evaluation standards or an auditing ecosystem.
Foundation model providers were found to rarely audit a models’ performance, even when it came to intentional harms. This becomes a great issue when investigating generative AI systems, which cannot simply be audited, and the answers they produce cannot be explained or traced.
Recommendations
Despite no AI model studied currently meets the EU’s requirements, the Stanford researchers believe compliance is feasible.
Recommended
- Cloud First ‘23 | The Whats, Whys, and Hows of Cloud Centres of Excellence
- NCSC Removes Thousands of Email Scams With Reporting Service
- Cabinet Office Engages AND Digital to Enhance its Find a Grant Service
Foundation model providers can look to those already with best practices to improve their compliance. Researchers identified “low-hanging fruit” in terms of improvements as being making documentation available to downstream developers.
Further, creating an industry standard by working with those outside the AI industry, like those in academia and civil society.
The researchers also targeted the EU AI Act to clear up some of its ambiguity – such as detailing what is necessary to disclose to comply with a ‘description of a model’s performance.”
Researchers also noted the technical expertise necessary to properly audit these AI models, and that the EU must make technical resources and talent available to member states as they try to enforce compliance.





