Weekly Paper Radar

2026-09-21 → 2026-09-27 · scanned 3855 unique arXiv papers · 3 candidate(s)

01

Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows

arXiv:2609.34017 · recall: citation · cites 1 seed paper(s)
agentmulti-agent

Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow.

Metadata

Authors: Uliana Elina

Affiliations: Unknown

Published: 2026-09-27

02

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

arXiv:2609.27490 · recall: citation · cites 1 seed paper(s)
agent

AI research agents must predict the effects of computational changes after budgeted experiments.

Metadata

Authors: Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li

Affiliations: Unknown

Published: 2026-09-23

03

BudgetVerify: Budget-Tiered Verification for Financial QA

arXiv:2609.33052 · recall: citation · cites 1 seed paper(s)

Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute.

Metadata

Authors: Janet Jenq, Hongda Shen

Affiliations: Unknown

Published: 2026-09-27