Abstract
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Community
MatrAIx: Simulating the World with 8.3 Billion AI Agents
A research team of over 200 scientists, organized by researchers from Harvard and MIT and including more than 40 scientists from OpenAI, Anthropic, Google DeepMind, and xAI, has published new research on using AI agents to simulate human behavior. Together they built ๐ ๐ฎ๐๐ฟ๐๐๐ , an evaluation infrastructure that tests AI systems and digital products with diverse populations of simulated users. To date, MatrAIx contains 8.3 billion persona profiles spanning 1,290 persona dimensions and 1,010 application tasks. The paper reports 18,189 simulated user trials, with GPT-5.5, Claude Opus 4.8, and Claude Haiku 4.5 serving as the underlying models that power the personas.
This research tackles a question of growing importance:
Can AI agents imitate the diversity of human behavior and preferences?
The study first examines whether agents can simulate human preferences in evaluation settings. A high-quality user study requires recruiting, screening, testing, and statistical analysis. Reaching a few dozen participants is difficult, and comparing thousands of users at once is even more challenging. Running large-scale tests with real users is too time-consuming and costly to be practical.
The core idea behind MatrAIx is to give LLM agents personas, so that they can imitate diverse user groups and simulate large-scale human evaluations. The 8.3 billion persona-driven virtual users in MatrAIx can complete a wide range of testing tasks quickly and accurately. This addresses several long-standing problems in product testing: high cost, long turnaround, and limited coverage.
The MatrAIx platform has three core components: 1. the Persona-8B persona database, 2. four types of simulation environments, and 3. evaluation tasks across 25 domains.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents (2026)
- Qwen-CUA: Native Computer Use for (almost) Everything (2026)
- SWE-Together: Evaluating Coding Agents in Interactive User Sessions (2026)
- MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents (2026)
- How Benchmarks Mis-Score Computer-Use Agents (2026)
- Fara-1.5: Scalable Learning Environments for Computer Use Agents (2026)
- OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.04205 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
