• About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
Saturday, August 1, 2026
mGrowTech
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions
No Result
View All Result
mGrowTech
No Result
View All Result
Home Al, Analytics and Automation

Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Josh by Josh
August 1, 2026
in Al, Analytics and Automation
0
Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks


Supabase has open sourced Supabase Evals, its benchmark and framework for testing how well AI agents build using Supabase. It runs coding agents including Claude Code, Codex, and OpenCode against real tasks, such as building a schema, debugging a failed Edge Function, or fixing a broken RLS policy, then scores the result. It powers the public leaderboard at supabase.com/evals and an internal regression suite monitored daily.

Is it deployable?

Yes, today. supabase/evals is public under Apache-2.0 and runs locally via pnpm.

READ ALSO

The Current State of Agentic AI

OpenAI’s Widened Probe Turns Up More Agent Escapes – Unite.AI

  • Industries: Developer tooling, cloud infrastructure, data platforms, and regulated backends in fintech or healthcare, where an agent writing a wrong RLS policy is a security incident.
  • Applications: Regression-testing docs and skill edits, gating SDK releases, and comparing agent harnesses head to head.
  • Constraints: Local-stack runs need a Docker daemon, provider API keys, and ports 54321–54329 free.

How the harness works

Supabase defined three dimensions: products (database, auth, storage, edge-functions, realtime, cron, queues, vectors, data-api), topics (RLS, security, migrations, SQL, SDK, observability, self-hosting, tests, declarative-schema), and stages (build, deploy, investigate, resolve). It then picked the smallest scenario set touching each dimension once, grounded in support tickets, bug reports, and GitHub issues.

Scenarios split into two suites. Benchmark scenarios cover breadth and are published. Regression scenarios cover known failure modes, refresh daily, and do not move published scores.

Every scenario runs against a real environment. The framework boots a hosted-like stack and a local CLI project in containers, so agents call the actual MCP server and CLI. A platform-lite runtime exposes a Management API-compatible surface backed by @supabase/lite. Scoring combines deterministic checks with LLM-as-a-judge. Agents get one retry before grading.

Each eval directory holds PROMPT.md (task plus frontmatter), EVAL.ts (the scorer), and optional remote/ and local/ starting states. Shipping a local/ workspace, or declaring interface: cli, boots a Docker sandbox with the real CLI installed.


Findings

Agents pass most scenarios with no skill loaded. In the Build stage, Opus 5 and Kimi K3 both scored 100% unaided. Skills closed the rest of the gap: Sonnet 5 rose from 78% to 100%, GPT-5.6 Sol from 89% to 100%, and GPT-5.4 mini from 78% to 89%.

Three weaknesses surfaced. Agents hand-write migrations instead of using declarative schemas, prompting a skill guidance update. Agents verify auth by hand rather than reaching for @supabase/server, prompting a package selection guide. And docs usage varies sharply: Codex / GPT-5.6 reads roughly 8 docs pages per scenario versus about 2 for Claude Code, which checks docs in under 40% of scenarios even with skills loaded.

Key Takeaways

  • Supabase open sourced supabase/evals under Apache-2.0.
  • Scenarios run against real containerized Supabase stacks, not mocks.
  • Scoring mixes deterministic checks with LLM-as-a-judge; one retry allowed.
  • Skills mattered least for top models, most for smaller ones.

Check out the Technical details and GitHub Repo. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.



Source_link

Related Posts

Al, Analytics and Automation

The Current State of Agentic AI

August 1, 2026
OpenAI’s Widened Probe Turns Up More Agent Escapes – Unite.AI
Al, Analytics and Automation

OpenAI’s Widened Probe Turns Up More Agent Escapes – Unite.AI

July 31, 2026
Connecting research to policy on Capitol Hill | MIT News
Al, Analytics and Automation

Connecting research to policy on Capitol Hill | MIT News

July 31, 2026
JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI
Al, Analytics and Automation

JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI

July 31, 2026
An Introduction to Loop Engineering
Al, Analytics and Automation

An Introduction to Loop Engineering

July 31, 2026
OpenAI Cuts API Prices on Its Two Cheaper GPT-5.6 Tiers – Unite.AI
Al, Analytics and Automation

OpenAI Cuts API Prices on Its Two Cheaper GPT-5.6 Tiers – Unite.AI

July 31, 2026
Next Post
Cyberattacks Hit Water Facilities In Seven States Across The US

Cyberattacks Hit Water Facilities In Seven States Across The US

POPULAR NEWS

Trump ends trade talks with Canada over a digital services tax

Trump ends trade talks with Canada over a digital services tax

June 28, 2025
15 Trending Songs on TikTok in 2025 (+ How to Use Them)

15 Trending Songs on TikTok in 2025 (+ How to Use Them)

June 18, 2025
Communication Effectiveness Skills For Business Leaders

Communication Effectiveness Skills For Business Leaders

June 10, 2025
Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

Comparing the Top 7 Large Language Models LLMs/Systems for Coding in 2025

November 4, 2025
App Development Cost in Singapore: Pricing Breakdown & Insights

App Development Cost in Singapore: Pricing Breakdown & Insights

June 22, 2025

EDITOR'S PICK

A Coding Implementation to Build and Train Advanced Architectures with Residual Connections, Self-Attention, and Adaptive Optimization Using JAX, Flax, and Optax

A Coding Implementation to Build and Train Advanced Architectures with Residual Connections, Self-Attention, and Adaptive Optimization Using JAX, Flax, and Optax

November 12, 2025
Crisis Management in Health Tech: A Leadership Guide For AI-Driven Medicine

Crisis Management in Health Tech: A Leadership Guide For AI-Driven Medicine

August 27, 2025
Prime Video’s Blade Runner 2099 Now Has A Trailer And Release Date Of November 25

Prime Video’s Blade Runner 2099 Now Has A Trailer And Release Date Of November 25

July 24, 2026

Inside United Airlines’ ‘Mean Girls Day’ campaign and the pivot that made it work

April 16, 2026

About

We bring you the best Premium WordPress Themes that perfect for news, magazine, personal blog, etc. Check our landing page for details.

Follow us

Categories

  • Account Based Marketing
  • Ad Management
  • Al, Analytics and Automation
  • Brand Management
  • Channel Marketing
  • Digital Marketing
  • Direct Marketing
  • Event Management
  • Google Marketing
  • Marketing Attribution and Consulting
  • Marketing Automation
  • Mobile Marketing
  • PR Solutions
  • Social Media Management
  • Technology And Software
  • Uncategorized

Recent Posts

  • 4 questions every brand should answer before launching a podcast
  • Cyberattacks Hit Water Facilities In Seven States Across The US
  • Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks
  • Google might launch a ‘Pixel Tag’
  • About Us
  • Disclaimer
  • Contact Us
  • Privacy Policy
No Result
View All Result
  • Technology And Software
    • Account Based Marketing
    • Channel Marketing
    • Marketing Automation
      • Al, Analytics and Automation
      • Ad Management
  • Digital Marketing
    • Social Media Management
    • Google Marketing
  • Direct Marketing
    • Brand Management
    • Marketing Attribution and Consulting
  • Mobile Marketing
  • Event Management
  • PR Solutions