• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Monday, July 20, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Al, Analytics and Automation

Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep

Josh by Josh
July 19, 2026
in Al, Analytics and Automation
0
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep


Research agents already handle real knowledge work today. Teams delegate competitive mapping, due diligence, and literature review to them. However, most benchmarks test a single answer, not large evidence-backed collections. Perplexity targets that gap with a new open benchmark.

Perplexity released WANDR (Wide ANd Deep Research). It is an open benchmark and evaluation harness. It is built around 500 realistic, challenging data-collection tasks for knowledge work. WANDR is the wide sibling of Perplexity’s DRACO benchmark for deep research. DRACO asks whether an agent produces an accurate, complete, objective long-form report. WANDR instead asks whether it can build a large collection with evidence.

READ ALSO

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model

What is WANDR

At its core, WANDR tests two demands together. Wide means discovering a large, often open-ended set of qualifying entities. Deep means investigating every entity enough to support each claim with evidence. Combining both changes the problem for agents. A few compelling examples are not enough here. A polished narrative built on incomplete research also falls short.

To capture this, WANDR uses a composable qualification key hierarchy. One task might request company(n) -> employee(m) -> url(k). This means n qualifying companies, m employees each, and k supporting pages each. Every complete path through the tree gets validated independently. The same structure can represent a flat list, nested search, or matrix.

A Concrete Task Example

To ground that hierarchy, consider the released ceo_cfo_appointments task. It asks for at least 70 US-based companies. Each must have a CEO or CFO appointment first announced between March 1 and April 30, 2026. For each, the agent supplies one authoritative appointment page. A subtask adds a listing-authority page per company. Together, the task requires 140 source-backed records.

Concretely, the two hierarchies and one submitted record look like this:

# Task hierarchies
company(70) -> company_appointee(1) -> url(1)   # 70 appointment records
company(70) -> url(1)                           # 70 listing records

# One record the grader re-fetches and re-checks (values are illustrative)
{
  "item":     "Example Corp - new CFO",
  "url":      "https://issuer.example.com/press/cfo-appointment",
  "excerpts": ["Example Corp today named Jane Doe as Chief Financial Officer, effective April 2026."],
  "answer":   "Jane Doe appointed CFO; announced April 2026"
}

Realistic Tasks, Generated At Scale

Beyond single examples, WANDR builds its tasks from real usage. It starts from de-identified patterns seen in production, not synthetic prompts. A semi-automated pipeline then turns those patterns into tasks. The pipeline runs four stages: seeding, authoring, admission, and curation. It uses an interleaved author-critic loop with mechanical linting.

As a result, the median task asks for 50 members and 245 records overall. Across all 500 tasks, WANDR calls for 170,495 source-backed records. Tasks split into 167 lower, 166 middle, and 167 higher difficulty. Difficulty depends on per-record work, not scale alone.

How WANDR Grades Agents

Unlike fixed answer keys, WANDR grades each claim against cited evidence. Every record contains an item, URL, selected excerpts, and answer. The grader re-fetches the page during evaluation. It checks whether the page is usable and in scope. It then verifies the excerpts truly appear and support every requirement.

These binary record verdicts then roll up through the hierarchy. Precision measures the quality of what a system submitted. Recall measures quality-adjusted completion, filling any shortfall with zeros. Soft scores give partial credit to incomplete members. Hard scores count only members whose full subtree is correct.

The Benchmark Results

Using that method, Perplexity ran six production systems on all 500 tasks. Its own Search as Code (SaC) system leads. Still, no system comes close to solving the benchmark.

System Soft F1 Hard F1 Notes
Perplexity (Search as Code) 0.363 0.133 $5.20/task, 14.9-min median, 3.82M tokens/task
Anthropic 0.249 0.072 Closest on quality, but more time, money, tokens
Others (best) 0.121 0.035 OpenAI, Exa faster and cheaper, but lower scores

With more effort, Perplexity reaches 0.447 soft F1 at the xhigh setting. Cost across settings spans more than four orders of magnitude. It ranges from $0.03 per task up to $324.83 per task.

Beyond the leaderboard, four findings stand out. First, partial progress is common, but complete coverage is not. Every system shows soft recall below soft precision. Second, scale compounds the problem sharply. Deeper hierarchies hurt most, since each branch adds a failure point. Third, discovery is the first structural bottleneck. Top-level discovery completion ranges from 0.611 to 0.951 across systems. Under-delivery, not duplicate merging, explains most missing volume. Fourth, finding a usable page is usually easy. Turning it into complete evidence is the hard part. For Perplexity, 41.4% of pages miss a substantive requirement. Also, 57.5% of excerpts fail to support the full claim. Its soft F1 falls from 0.531 under a retrieval-only check to 0.363 under the full verdict.

Notably, Search as Code fits this task shape well. An agent can express retrieval, filtering, fan-out, joins, deduplication, and stopping logic as a program. Deterministic compute then handles repeated operations outside the model context.

Use Cases With Examples

Practically, WANDR maps to jobs teams already automate. A market analyst needs every qualifying competitor, with matching evidence for each. A due-diligence team needs dozens of companies, then ownership, executives, and financing. Talent sourcing needs many candidates, each with supporting profile pages. WANDR tests exactly these wide-and-deep collection patterns at professional scale.

Because grading is per-record, teams can localize failures precisely. The score tree isolates loss to discovery, enrichment, or evidence extraction. This diagnosis helps engineers improve one weak stage at a time.

Key Takeaways

  • WANDR is an open benchmark with 500 evidence-heavy, wide-and-deep tasks.
  • Tasks use a qualification key hierarchy validated path by path.
  • Grading is reference-free; the grader re-fetches and checks cited evidence.
  • Perplexity Search as Code leads at 0.363 soft F1 and 0.133 hard F1.
  • Discovery and complete evidence remain the biggest failure points.

Check out the Technical details and Repo here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.



Source_link

Related Posts

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared
Al, Analytics and Automation

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

July 20, 2026
Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model
Al, Analytics and Automation

Someone Fine-Tuned OpenBMB’s MiniCPM5-1B on Claude Fable 5 Traces to Ship a 657MB Local Thinking Model

July 20, 2026
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
Al, Analytics and Automation

Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost

July 19, 2026
Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite
Al, Analytics and Automation

Google Cloud’s Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite

July 18, 2026
Following the questions where they lead | MIT News
Al, Analytics and Automation

Following the questions where they lead | MIT News

July 18, 2026
Al, Analytics and Automation

Build an Agentic Event Venue Operator with MongoDB Atlas, Voyage, and LangGraph

July 17, 2026
Next Post
Is It Worth Your Budget in 2025?

Is It Worth Your Budget in 2025?

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025

EDITOR'S PICK

Turn Pinterest Predicts 2026 Trends Into Traffic-Ready Keywords (with free resource)

December 12, 2025
What Gaming VCs Actually Look For: Team, Metrics, and the Power of Resilience June 2025 (Updated)

What Gaming VCs Actually Look For: Team, Metrics, and the Power of Resilience June 2025 (Updated)

June 24, 2026
I Found the 8 Best Security Compliance Software on G2

I Found the 8 Best Security Compliance Software on G2

October 30, 2025

Kinetiq and NLogic Partner to Advance TV Ad Intelligence

May 27, 2025

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • Measuring What Matters: Q&A With AMEC Chair Raina Lazarova
  • ROAS Goal for Maximize Number of Conversions
  • The Galaxy Card Is Samsung’s Answer to the Apple Card
  • I Tested G2’s 10 Best Marketing Automation Tools: My Review
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions