Site navigation

DIGIT EXPO West: Harnessing Data To Achieve AI Success

Elizabeth Greenberg

,

data ai
At DIGIT’s inaugural Expo West event in Glasgow, Dr Janet Bastiman, chief data scientist, Napier AI, explained how to battle through bad data to succeed in AI.  

“There’s a reason it’s called data science,” Dr Janet Bastiman, chief data scientist at Napier AI said to a packed keynote hall at DIGIT Expo West. “Because science is magic that works.”

While advances in AI may seem magical, they can cast a dangerous spell on businesses and projects alike if its use does not have a solid enough foundation. After all, science requires a sturdy basis and methodology to ensure it works.

Dr Bastiman explained some of the most common and prolific AI downfalls and how understanding data, and data science principles, is essential to successful AI implementation.

AI projects can be multi-million pound projects, and make critical decisions in sectors like healthcare, financial services, and public service. They can be used to detect anomalies for natural disasters, fight financial fraud, and screen for cancer.

With many companies jumping on the AI hype train, however, they could be on the road to nowhere, or worse – crashing and burning.

A lot of teams might have a simplistic look at how AI systems actually work: “Identify the problem, gather some data, feed it into AI, mess around a bit and you’ll have a solution.

“This is the magic you’ve been fed,” Bastiman mused—and she would know, considering her deep and impressive experience in and around data.

Bastiman served as the chair of the Royal Statistical Society’s data science and AI section, was part of the FCA’s synthetic data expert group, and is now the chief data scientist at Napier AI, using AI to fight financial crime.

She knows that “no company sets out to fail,” but without a solid foundation of data, these can be doomed to fail.

Building this requires robust data management, understanding, and a stringent scientific method. Bastiman presented some baseline essentials in building this data foundation.

Building Your Data Backbone

AI is just data science in action. Data is the foundation of AI, and ensuring it is understood and sturdy is an essential first step of using AI.

“You probably all heard that really twee saying that ‘data is the new oil,’” Dr Bastiman said. “Well, if we follow that analogy, and it is a fair one, we do not use crude oil straight from the ground.”

It does not matter how advanced a machine is, or how skilled the driver is – crude oil is not going to power a vehicle. The same is true for data. It needs to be properly understood, filtered, and harnessed for specific use cases.

Dr Bastiman provided integrated starting points for ensuring that data is robust enough for AI implementation.

“So the first question that I am always asked by customers whenever they want an AI project is: ‘How much data do I need?’ And it is a really difficult question to answer because the answer is always: ‘It depends.’”

Availability, reliability, representation and bias are all things to consider from the amount of data one inputs.

Data can be corrupted even before it finds its way to you – biases and assumptions can be made before it even ends up on a system. Understanding data providence and what assumptions were made prior to its implementation is therefore vital.

Understanding if a data set is representative of what the data is being used to convey – for instance, using a national basketball team to estimate average height across a country’s population, would not be representative. While these assumptions and knowledge bases are obvious to us, they need to be implemented in a system.

This is where understanding use cases is so vital – it is important to know what the AI is being used for so that the appropriate amount of data can be used.

Trying to represent an entire population may require more representative data, but sometimes, less is more.

Irrelevant data can cloud an AI system’s vision, having it consider things outside the realm of its scope.

“When image recognition first came out, there was the famous problem of the background of the image being important to the classification,” Dr Bastiman said.

“So Huskies, for example, were not identified for any feature of the dog just from the snowy background. So if you showed the AI an array of images asking it to recognize a dog with a snowy background, it would always say Husky.

“So that background was irrelevant, but the network learned something it shouldn’t have done.”

Distilling data is essential, but documenting this is also vital in ensuring that bias is mitigated and assumptions can be tracked.

Cutting 95% of data away to hone in on an AI system may be required, but keeping this data for other uses, and marking where and how this data was cut, can allow analysts to track where a data set may have gone wrong.

While some data may be irrelevant, data gaps may exist, and more data may need to be generated to fully encapsulate and solve an issue.

This is where synthetic data can be useful, but the tool also has associated risks and potential pitfalls.

Synthetic data cannot be used on its own, Dr. Bastiman says. It needs to not only represent data, but have the same shape as real data.

“So the same statistics, the same correlations and interdependencies,” she explained.

“You can’t just generate randomness or you’ll learn nothing or at best your models will learn the wrong things.

“If you don’t have any real data to start with, then you’re making it up again from your bias, which is a huge problem.”

Even with real data, mitigating bias is difficult. “We are all biased based on our experience,” Dr Bastiman said. “And all of that information affects how we deal with data, what we believe to be true.”

Dr Bastiman continually referenced documentation and testing as essential to developing a data set – it is data science after all, and taking note of assumptions, processes, and testing is vital to ensuring the AI model actually works, and mitigated biases.

“Testing is absolutely critical to data science.”

Without proper testing, models that are unrepresentative, use incorrect assumptions, or are biased can enter an algorithm with dire consequences.

But what happens when you just do not have the right data for a project?


Recommended reading


Managing Management

Dr Bastiman dedicated some of her keynote to the issue of dealing with managers, directors, and executives when using data science and AI.

“It’s really hard if you’re asked to do something that isn’t great, and until you’re at a level where you can say no, you may be forced to do these things as a data scientist,” she said.

These asks can range from using a smaller dataset than required, to ethical disagreements and moral quandaries on biases, implementations, to what is actually possible based on the available data.

Explaining why a project just won’t work can be a challenge, however.

“It’s really difficult as it depends on the statistical literacy of your management team,” she said. “In an ideal world, they all would be. In the real world, they tend not to be.”

“And in that ideal world, the conversation should go a little bit like this. You express concern about a project, you say it’s unethical, your boss who’s lovely is very accepting of this, asks what you need, and you say I’ll try something else. That would be lovely.”

But the real world is often far from ideal. Often, data scientists are met with passive aggressiveness, threats to their job, a “do it or I’ll find someone who will” attitude.

First, Bastiman issued a plea to stakeholders.

“If you are a stakeholder and your data person is expressing concerns, please listen to them. Particularly if it’s a business critical project, they are expressing concerns for a reason.

“It’s not because they’re incapable, it’s because they have bad data. And bad results are never good for a company.”

But if that still does not work, Bastiman issued advice for data specialists themselves to explain their data concerns.

Showing your work – explicitly why a project just won’t work – can be one of the most effective ways to prove to management that you are not incapable, but that rather a job is impossible.

“A lot of the things that we’re working on have real world impacts,” she asserted.

“You need to make sure it’s right. Show them whether it can be done or not, what you need to fix it, or whether it just can’t be done.”

Elizabeth Greenberg

Staff Writer

Latest News

Cybersecurity Editor's Picks Recruitment Security

Comment | Building Cyber Talent Takes More Than a Degree

Culture Featured Technology

Inside TecTonic’s Growing Innovation Market Square

Cybersecurity

Revolut Leaked Customer Data to Fake Government Email Account

Cybersecurity Editor's Picks Security

Welsh SMEs Urged to Strengthen Cyber Defences