Site navigation

Fighting Fire with Fire: Using LLMs to Take Down Cybercriminals

Michael Edgar

,

LLMs Cybercriminals
Half of all social engineering cyber attacks are pretexting incidents, where the attacker impersonates someone else to manipulate the victim. 

This method has been proliferated as large language models (LLMs) become more commonplace and intelligent. Attackers use AI to target what is often the most vulnerable component of any organisation, the human element, where a person is a critical point in the chain of attack. 

Of all recorded breaches, 74% of them involved a human element, according to Verizon’s 2023 Data Breach Investigations Report. The human element in this context is social engineering or misuse, which has spiked in use since 2021, when only 35% of successful breaches started this way. 

Security researchers have pointed to AI as the catalyst for this change, since it can be used to translate and refine attacks. 

“Cybercriminals are becoming more sophisticated in their attacks, using AI and existing ransomware code to drill deeper into victims’ systems and extract sensitive information.

“AI-created malware is adept at avoiding detection in traditional antivirus models and public ransomware cases have exploded relative to last year,” read a cyberthreats report from cybersecurity provider Acronis.

Behind this all is a burgeoning marketplace on the dark web, selling AI tools like WormGPT to help threat actors carry out attacks. In response, researchers in South Korea are training an LLM to help give security practitioners an upper hand in the source of cybercrime. 

Enter: DarkBERT

While AI is being used to enrich and automate cybercriminal activity, the same can be used to the advantage of cybersecurity researchers.

According to the South Korean researchers, DarkBERT is trained on over 6 million pages of dark web content. Compared to its “vanilla counterpart,” trained on surface internet data, DarkBERT outperforms current LLMs in the context of understanding the dark web’s complex internal language and structure. 

This LLM can be used in multiple cases, say the researchers. First, to detect ransomware leak sites. Ransomware is one of the top action types present in breaches, making up  almost a quarter (24%) of attacks according to the Data Breach Investigations report. DarkBERT was able to identify ransomware leak sites on the dark web more accurately than models trained on the surface-web. 

DarkBERT is also able to monitor developing attacks before they happen in “noteworthy threads.” Threat actors will often share information like admin access credentials, source code, or information on access attempts on the dark web. By monitoring these threads successfully, DarkBERT could be used to notify victims quickly, and locate the source of stolen data. 

The last use case identified by the South Korean researchers was DarkBERT’s fill function. Due to the underground nature of the dark web, many deals and advertisements happen in coded language and slang. When compared to a surface-web trained LLM, DarkBERT showed a higher aptitude of recognizing the meaning of slang words. 

Michael Edgar

Staff Writer, DIGIT

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data