Site navigation

Microsoft Research Shows Agentic AI Multitasking Like Humans

Graham Turner

,

CORPGEN framework
Researchers at Microsoft have introduced CORPGEN, a friendly-sounding framework designed to help AI agents manage multiple complex workplace tasks simultaneously.

Microsoft researchers have outlined a new framework designed to help AI agents operate more like knowledge workers, arguing that today’s evaluation methods fail to reflect the realities of modern corporate life.

In a research paper pithily titled CORPGEN: Simulating Corporate Environments with Autonomous Digital Employees in Multi-Horizon Task Environments, the team proposes an agent architecture intended to equip AI systems with the memory, planning, and learning capabilities required to handle multiple complex tasks simultaneously over extended periods.

The research begins from a familiar workplace scenario: by mid-morning, a typical knowledge worker may already be balancing a client report, a budget spreadsheet, a slide deck, and an email backlog, all interdependent and demanding attention at once. However, the researchers argue that most leading AI models are still assessed one task at a time.

To address that disconnect, they developed what they call Multi-Horizon Task Environments (MHTEs), designed to replicate the demands of sustained, overlapping corporate workloads.

Within these environments, an agent must manage multiple complex tasks at once, with each task involving between 10 and 30 dependent steps carried out during a single five-hour session.

When the researchers ran MHTEs at scale on several leading AI agents, they identified four recurring weaknesses. Memory constraints prevented agents from holding details for multiple active tasks simultaneously. Information from one task interfered with reasoning about another.

Task dependencies formed complex webs rather than simple sequences, requiring constant checks on upstream progress before downstream work could continue. Finally, every action cycle required reprioritisation across all active tasks, rather than simply resuming from the last point of activity.

The researchers also tested three independent agent systems under increasing workloads. As concurrent tasks rose from 12 to 46, completion rates dropped from 16.7% to 8.7% across all systems, highlighting a scalability problem in current approaches.

Testing AI in Realistic Corporate Workloads

Microsoft’s proposed solution is CORPGEN, which introduces the concept of “digital employees”: large language model-powered agents with persistent identities, role-specific expertise, and structured work schedules.

These agents operate Microsoft Office applications through graphical user interface automation and are designed to function consistently within MHTEs over hours of continuous activity.

Each working day begins with a structured plan generated from memory loaded from previous sessions. Agents then cycle repeatedly through overlapping tasks, retrieving context, reasoning through actions, and persisting results across more than 50 interleaved tasks. At the end of the day, they generate reflections and consolidate experience into long-term memory to inform future sessions.

CORPGEN’s architecture is designed to address the four weaknesses identified in earlier testing. Hierarchical planning breaks objectives into daily goals and then into immediate decisions, reducing the need to review all available tasks at every step. Subagents perform complex operations such as web research in isolated contexts to prevent cross-task contamination.

Beyond this, a tiered memory system enables selective recall of relevant information, rather than maintaining all data in active context. Adaptive summarisation compresses routine observations while preserving critical details to control memory growth.

Because these mechanisms are model-agnostic, the team tested CORPGEN across three different agents. In each case, they report consistent gains, arguing that improvements stemmed from the architecture itself rather than the strength of any individual base model.

From Digital Worker to Virtual Organisation

The framework also explores how multiple digital employees might collaborate.

When several agents operate within the same environment, coordination takes place through standard workplace communication channels, including email and Microsoft Teams, without predefined rules or shared internal state. One agent may send an email requesting data; another processes it in a later cycle and responds.

Over time, these exchanges form organisational patterns, with some agents assuming leadership roles and others providing support, while shared documents act as connective tissue.

If a communication path fails, such as through an email delivery error, agents reroute messages through alternative channels to maintain workflow. The researchers describe the outcome as a virtual organisation that behaves like a real one without explicit programming to enforce that structure.

Benchmark Results and Performance Gains

CORPGEN was evaluated using a multi-task benchmark combining up to 46 tasks in a single six-hour session. According to the researchers, baseline systems showed steady performance declines as workload increased, whereas CORPGEN maintained or improved completion rates. At 46 tasks, CORPGEN completed 15.2% of tasks compared with 4.3% for baseline systems, around 3.5 times higher.

The team introduced CORPGEN’s components sequentially to assess their impact. An initial orchestration layer and cognitive tools delivered moderate gains. However, the largest improvement came from experiential learning, whereby agents store records of completed tasks and reuse them when encountering structurally similar work. This mechanism increased completion rates from 8.7% to 15.2%.


Recommended reading


The researchers also found that evaluation methodology significantly influenced results. When assessing actual output files produced by agents, findings aligned with human judgements approximately 90% of the time. By contrast, evaluations based on screenshots and action logs aligned only around 40% of the time, suggesting that some commonly used assessment approaches may underestimate agent performance.

Looking ahead, the team argues that memory and retrieval systems, rather than raw model capability alone, may represent a key bottleneck in deploying AI agents effectively in real-world corporate settings. They suggest that agents capable of learning from prior successes and applying those patterns to similar tasks can build an advantage over systems that treat each assignment in isolation.

Future work will examine whether agents can maintain memory across multiple workdays and how they coordinate when operating in teams. The researchers are also exploring ways to increase speed and reliability by combining different methods of interacting with software.

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data