“Data is the new oil,” was what British mathematician Clive Humby said in 2006 during a time when the potential of data was only just being realised.
Now, data punctuates every aspect of our lives, and is used to personalise and optimise what you buy, watch, and share, right down to how your community operates.
However, usable data is not so readily accessed; first it comes in the form of mass amounts of raw data points which can take a lot of time to decipher due to the data being held across individual systems, multiple organisations, and formats.
Further, public sector data – collected by public bodies like the NHS, the government, and law enforcement – is difficult to access, often hidden in safe havens where sensitive data can only be reached under strict controls.
This presents a crossroads: while data is vital to developing products and services that benefit the health, economic, social and environmental wellbeing of Scotland, public sector data needs to remain secure and private.
It’s this quandary that has heralded the proliferation of synthetic data – a data set that is artificially generated, and mimics the statistical properties of real-world data.
Using mathematical models to create datasets offers researchers insights without compromising anyone’s privacy. For the researcher, synthetic data is quicker and easier to access.
Speaking with DIGIT, Dr Lynne Adair, data curation manager at Research Data Scotland (RDS) talked about the role RDS plays in Scotland, and how synthetic data is becoming a growing part of what it does.
Synthetic Data
RDS says it is “working to improve the economic, social and environmental well-being in Scotland by enabling access to and linking to data about people, places and businesses for research in the public good.”
“We are facilitating access to public sector data for research purposes and synthetic data is really for us an important means to improve and speed up access to the data, which can be a long and complicated process,” said Dr Adair.
Public sector data for research is accessed via a safe haven, which Dr Adair says is a secure room with no internet access that only allows approved researchers from approved projects to use. Any output from the data observed has to be approved, cleared, and then released.
“Synthetic data can help researchers at the early stages of a project determine what a dataset looks like, whether it contains the variables they need and is suitable for their purposes,” said Dr Adair.
Synthetic data can also be used to develop code without the safe haven environment, before researchers get access to the real data, thus speeding up the whole project process.
Research Data Scotland
RDS is a not-for-profit charitable organisation established in 2021 by the Scottish government in a collaboration with public bodies – like Public Health Scotland, and National Records Scotland – and leading academic institutions in Scotland.
The COVID-19 pandemic showed how synthetic data could be used by organisations to provide data sets to researchers during a time where it was difficult to get access to real data.
On RDS, MSP John Swinney said: “Better connections between organisations and the data they hold will be instrumental to finding innovative solutions for issues such as sustainable employment, financial security for families and low-income households, and the wellbeing and mental health of children and young people.”
“We’re quite a new organisation,” said Dr Adair. “Our aim is to produce synthetic datasets for training, data discovery and code development purposes on an ongoing basis, for public sector datasets in Scotland.”
Now around two years into operation, RDS released its 2023-2024 business plan. According to the plan, 2023 has ambitions to improve the quality of researcher data. RDS then plans to shift gears to broadening the range of users to include the public and third sector, and eventually to researchers outside of Scotland to attract investment.
In March of this year, RDS invested £56,000 in nine projects across Scotland to support public engagement in data. The winners were selected from over 30 applicants and included projects from the University of Edinburgh, Dundee, Glasgow, as well as CodeClan, Grampian Regional Equity Council, and People Know How.
Creating Synthetic Data
Synthetic data is made in various ways. The first step is to decide what data points researchers want to include and get an idea of patterns and attributes it has to make sure the final synthetic data product is similar to the real data.
Next researchers would choose a generation method such as open source tools like Synthpop. Synthpop synthetic dataset variables are synthesised one-by-one using sequential regression modelling, which is statistical technique used to understand and predict the relationship between variables over time or in a specific sequence. Other programs might use randomisation to do the same thing.
Once the synthetic data is created, it is then compared with the real data, to make sure it is similar, and accurate. At this point, any sensitive information is removed. In synthetic data generation it is vitally important that no personal or private details are disclosed.
Synthetic data does have a broad spectrum on the fidelity of synthesis. High fidelity data contains more details and nuances, whereas low fidelity would contain omissions, giving a less detailed snapshot.
Therefore when the fidelity is higher, the data is closer to the ‘real thing’. However, if there is a higher analytical value, there is a greater disclosure risk. “We need to consider what level of fidelity is necessary and this may differ depending on the use of the data,” reads the RDS synthetic data strategy document.
“Low fidelity data sets might be totally fine for your purposes. For the data discovery or the code development side of things that could be fine,” said Dr Adair.
“For researchers, it might be better to have a more high fidelity data set for code development, so these are all issues that we’re looking at and thinking about because we’re still at quite an early stage.”
After a synthetic dataset is created, there also needs to be considerations on how it will be accessed and where it will be stored. The fidelity of the data, accreditations of users, and what it is being used for would all influence these protocols.
Recommended
- Report: UK Home to Europe’s Highest Number of Future Unicorns
- Centrica to Build Battery Storage Project in Perthshire
- 6th Edition of Scots Fintech Festival Coming in September
Scotland’s Role
A global forecast report predicts the synthetic data generation market will grow by a CAGR of 45.7% to £1.6 billion in 2028. Scotland is well positioned to be a key player in this market due to large investments being made into harnessing data technology and capabilities.
In 2018, as part of the Edinburgh and South–East City Region deal, Edinburgh set out to be the data capital of Europe. Now the city competes with a number of other data hubs like Dublin, Amsterdam, and Stockholm, but still maintains a strong presence in the field.
According to research from Accenture, the talent is here, but often gets pulled away from Scotland to other hubs due to more competitive salaries.
However, work from organisations like RDS whose goal is to make public sector data in Scotland easier to access, will likely result in more interesting data roles.
Synthetic data also has the ability to draw more researchers here who will want to use the data, turning Scotland into the research and data hub it is on the journey to becoming.





