本篇博文主要内容为 2026-07-31 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。
说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。
提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。
目录
概览 (2026-07-31)
今日共更新715篇论文,其中:
- 自然语言处理共99篇(Computation and Language (cs.CL))
- 人工智能共246篇(Artificial Intelligence (cs.AI))
- 计算机视觉共131篇(Computer Vision and Pattern Recognition (cs.CV))
- 机器学习共196篇(Machine Learning (cs.LG))
- 多智能体系统共19篇(Multiagent Systems (cs.MA))
- 信息检索共27篇(Information Retrieval (cs.IR))
- 人机交互共20篇(Human-Computer Interaction (cs.HC))
多智能体系统
[MA-0] Using Theory of Mind to Arbitrate between Social and Non-social Learning
【速读】:该论文旨在解决人类在社会学习(Social Learning)与非社会学习(Non-social Learning,即直接经验探索)之间如何权衡决策的问题。其核心挑战在于:尽管社会学习能有效降低认知负担,但个体有时仍会选择耗费时间和精力的直接探索,这表明学习策略的选择并非简单依赖于信息获取效率,而是涉及复杂的成本-收益评估。论文提出的解决方案是“理性心理理论模型”(Rational Mentalizing Model),其关键在于通过推理其他代理人的目标意图及其未来行为的信息量,来估算社会学习的预期效用,并将其与非社会学习的效用进行对比,从而实现基于效用最大化的策略选择。该模型通过设计一种新型博弈任务,成功定量捕捉了人类在观察他人与自主探索之间的行为权衡,揭示了选择性社会学习本质上受“心理理论”(Theory of Mind)驱动,以优化个体在复杂环境中的学习效率。
链接: https://arxiv.org/abs/2607.28601
作者: Lance Ying,Ryan Truong,Joshua B. Tenenbaum,Samuel J. Gershman
机构: Harvard University (哈佛大学); Massachusetts Institute of Technology (麻省理工学院)
类目: Multiagent Systems (cs.MA); Neurons and Cognition (q-bio.NC)
备注: 35 pages, includes supplementary information
Abstract:Social learning is a powerful mechanism through which agents learn about the world from others. However, humans sometimes choose direct experience over social learning, which can carry time and cognitive resource costs. How do people balance social and non-social learning? We propose a Rational Mentalizing model of the decision to engage in social learning. This model estimates the utility of social learning by reasoning about another agent’s goal and the informativeness of their future actions. It then weighs the utility of social learning against the utility of non-social learning. Using a novel game where players choose between observing other agents or exploring the environment, we show that the Rational Mentalizing model can quantitatively capture human trade-offs between these strategies. These findings suggest that selective social learning is guided by ‘Theory of Mind’ in the service of utility maximization.
[MA-1] Algorithms for Structured Elections under Thiele Voting Rules AAAI2026
【速读】:该论文旨在解决基于批准制的委员会选举中,采用Thiele投票规则时的胜者确定问题(winner determination problem)的计算复杂性。此类规则通过一个固定的权重向量来刻画选民满意度随当选批准候选人数量变化的关系,其核心挑战在于在满足比例代表性等要求下高效寻找最优委员会。论文的关键贡献在于揭示了最优解结构与选民批准行为之间的内在依赖关系——即候选人的批准集如何通过选民群体的重叠产生约束。基于此洞察,研究设计出针对“选民区间”(Voter Interval, VI)这一自然受限域的固定参数可追踪(FPT)算法,适用于比例批准投票(PAV)及其他Thiele规则。特别地,证明了即使在参数取常数时,该问题在一般实例上仍为NP难,但在VI域下可通过特定参数实现FPT求解。此外,论文还解决了文献中的两个开放问题:一是提出每个候选人最多被两名选民批准的情形下的多项式时间算法;二是给出了以获胜委员会总得分为参数的FPT算法,从而深化了对PAV在VI实例上计算复杂性的理解,推动了该领域关键问题的进展。
链接: https://arxiv.org/abs/2607.28575
作者: Alexandra Lassota,Krzysztof Sornat
机构: TU Eindhoven, the Netherlands(埃因霍温理工大学); AGH University, Poland(AGH大学)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Data Structures and Algorithms (cs.DS); Multiagent Systems (cs.MA)
备注: 18 pages. A conference version of this work appeared in AAAI 2026
Abstract:We study the computational complexity of winner determination problems in approval-based committee elections under Thiele voting rules. These form a class of rules parameterized by a fixed weight vector that specifies how a voter’s satisfaction depends on the number of approved candidates elected. We first analyze the structure of optimal solutions based on the sets of voters who approve each candidate—that is, how voters’ approval ballots induce dependencies between candidates—revealing constraints on a winning committee under any fixed Thiele voting rule. Using this, we design FPT algorithms for Proportional Approval Voting (PAV) and other Thiele rules on a natural restricted domain known as the Voter Interval (VI) domain—that is, after a suitable ordering of voters, each candidate is approved by a consecutive interval of voters. In particular, we show that every Thiele rule on VI is FPT with respect to a parameter for which the problem is NP-hard on general instances, even when the parameter takes constant values. Our results advance the understanding of the computational complexity of PAV on Voter Interval instances, which remains one of the central open questions in this area. We further resolve two open questions from the literature on PAV (and other Thiele voting rules) by providing a polynomial-time algorithm for instances where each candidate is approved by at most two voters, and an FPT algorithm parameterized by the total score of a winning committee.
[MA-2] APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems
【速读】:该论文旨在解决在数据稀缺场景下(如新型晶体相或从头设计蛋白)3D原子结构预测中因缺乏实验标注而带来的模型训练瓶颈问题。传统基于流匹配的生成式模型(如FlowDPO)依赖于监督偏好学习来对齐真实坐标,但此类标签获取成本高昂且不可行。为此,本文提出一种完全无监督的对齐框架——原子策略优化(Atomic Policy Optimization, APO),其核心创新在于将群体相对策略优化方法引入三维原子环境建模,并设计了一种新颖的双奖励机制:(i) 通过样本相似性矩阵的特征分解识别并强化模型主导的潜在结构模式,(ii) 引入热力学稳定性奖励以约束物理合理性。该机制使模型能够在采样群组中实现“自校正”,自主筛选出符合物理规律的构型。大量在晶体与抗体结构预测任务上的基准测试表明,APO在匹配率与结构保真度方面均超越全监督基线,达到新最优性能;同时,其有效拉直了概率路径,显著提升推理效率。结果表明,内在物理一致性可作为比噪声较大的坐标监督更优的对齐引导信号。
链接: https://arxiv.org/abs/2607.28553
作者: Shentong Mo,Yatao Bian
机构: CMU(卡内基梅隆大学); NUS(新加坡国立大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy’s dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct’’ by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.
[MA-3] Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
【速读】:该论文旨在解决在双人零和不完美信息博弈中,智能体如何在保证安全性的前提下有效利用对手策略缺陷以获取额外收益的问题。其核心挑战在于:采用纳什均衡策略虽能确保博弈价值(game value),但无法捕获因对手错误而带来的潜在增益;而传统的偏差探测方法(如二元释放规则)可能因证据不足无法及时响应,或基于不完整对手模型的全最佳应对则易被对手反向利用。为此,论文提出预算约束的置信度调度受限回应(Budget-Constrained Confidence-Scheduled Restricted Responses, CS-RNR),其关键创新在于构建了一个可验证的安全性证书——该证书由智能体在实际部署策略时自行计算,确保所有主动实施的策略调整均经过自我审计,从而实现“可证明的安全性”。具体而言,该方法通过任意时间有效的置信序列追踪动作频率分布,仅当观测到的频率区间与纳什均衡参考值显著分离时才认定为可利用偏差,并据此构建保守的对手模型。在此基础上,通过在一系列“钉定水平”(pin levels)上求解受限回应策略生成候选反制策略,再在部署前对每个完整候选策略进行全树最佳应对评估,形成最终的损失容忍证书。该证书与用户设定的预算进行比较,仅当满足条件时才原子化地采纳该策略。由于验证环节作用于实际执行策略,模型质量决定可实现的剥削程度,而证书则控制相对于基准策略的期望损失。实验表明,在Leduc德州扑克中,CS-RNR获得的稳态收益是金钱验证二元门方法的6.2倍,且所有部署策略均严格遵守预算;使用相同估计器的轨迹混合方法进一步达到13.6倍预算收益。在Leduc、说谎骰子(Liar’s Dice)及5阶Leduc三种博弈中,总计36,000个经审计的手牌均满足报告的证书容差要求,验证了方法的有效性与安全性。
链接: https://arxiv.org/abs/2607.28520
作者: Boning Li,Longbo Huang
机构: IIIS, Tsinghua University (清华大学智能技术与系统国家重点实验室); Tsinghua University (清华大学)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 21 pages, 5 figures
Abstract:An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emphbudget-constrained confidence-scheduled restricted responses (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself. The method tracks pooled action frequencies with anytime-valid confidence sequences and treats a frequency as exploitable only once its interval separates from an equilibrium reference. The confirmed deviations define a conservative opponent model, which a restricted-response solve turns into candidate counter-strategies over a grid of pin levels. Before deployment, each complete candidate is evaluated by a full-tree best response. The resulting certificate is compared with a user-specified budget and committed atomically with the strategy. Because this check is performed on the played strategy, model quality determines the exploitation achieved while the certificate controls reference-relative expected loss. In Leduc hold’em, CS-RNR obtains 6.2\times the steady-state gain of a money-verified binary gate while keeping every deployed strategy within budget. A trajectory mixture using the same estimator reaches 13.6\times the budget. Across Leduc, Liar’s Dice, and 5-rank Leduc, all 36,000 audited hands satisfy the reported certificate tolerance.
[MA-4] Agent Radio: Passive Awareness for Long-Horizon Multi-Agent Collaboration
【速读】:该论文旨在解决大型代码库中长时程任务(long-horizon task)下大语言模型(LLM)代理在代码理解与协作执行中的局限性问题。具体而言,面对生产级代码库中的复杂问答任务,单一代理因无法有效处理跨文件执行追踪、证据整合及长时间推理而表现不佳(如Claude Code Agent Opus 4.6仅解决32.3%的任务)。现有多代理系统虽通过分阶段任务分配缓解了这一问题,但其通信机制局限于阶段边界,导致中间发现无法实时共享,阻碍了动态协同。为此,论文提出AgentRadio——一种异步消息传递层,为编码代理提供线程(threads)、消息(messages)和“等待提及”(waiting for mentions)三项核心原语。其中,“等待提及”作为后台任务,使代理可在不中断当前工作的情况下被动感知同伴的更新信息,实现对新发现的即时融合。基于五阶段分工与协商协议,四代理协作体系在SWE-Atlas QnA上将任务解决率提升至62.1%,显著优于单代理及更先进版本的Claude Code(Opus 4.8,57.2%)。细粒度评估表明,性能增益随任务难度增加而扩大,验证了中程纠错(mid-course correction)作为核心机制的有效性。
链接: https://arxiv.org/abs/2607.28430
作者: Xinxing Ren,Qianbo Zang,Ziyan Wang,Caelum Forder,Suman Deb,Peter Carroll,Zekun Guo
机构: Google(谷歌); Stanford University (斯坦福大学); University of California, Berkeley (加州大学伯克利分校); Massachusetts Institute of Technology (麻省理工学院)
类目: Multiagent Systems (cs.MA)
备注:
Abstract:Understanding large codebases is a long-horizon task for Large Language Model (LLM) agents: answering a single question can require building and running the software, tracing execution across files, and synthesizing evidence over tens of minutes. On SWE-Atlas QnA, a benchmark of long-horizon questions over production repositories, a single Claude Code agent (Opus 4.6) resolves only 32.3% of tasks. Dividing the work among agents with clean contexts mitigates this limitation. However, the subtasks of code comprehension are interdependent. One agent’s findings can rewrite another’s task, so agents must coordinate during execution, not only at phase boundaries. Existing multi-agent systems support such exchange only between phases, through staged handoffs or synchronized rounds. Communication and work remain mutually exclusive. A discovery made mid-execution cannot be shared until the next boundary. We present AgentRadio, an asynchronous message-passing layer that equips coding-agent harnesses with three primitives: threads, messages, and waiting for mentions. The last runs as a background task, surfacing teammates’ messages without interrupting foreground work, so each agent remains passively aware of its peers and folds new findings into its ongoing task. Under a five-phase protocol of division of labor and negotiation, four agents organized by AgentRadio resolve 62.1% of tasks, 29.8 points above a single agent and above Claude Code with the newer Opus 4.8 (57.2%). Rubric-level analysis shows the gain growing with task difficulty, consistent with mid-course correction as the underlying mechanism. Our code is available at this https URL.
[MA-5] Agent ic Metaverse Services: A New As-a-Service Paradigm
【速读】:该论文旨在解决传统数字服务在智能化、自主性与协同能力方面的局限性,特别是在虚拟生态系统——元宇宙(Metaverse)中,如何实现具备高度自主性、多模态交互能力及协作决策能力的智能服务。其核心问题是:如何将生成式人工智能(Generative AI)驱动的智能体(Agent)能力有效集成于元宇宙环境,以构建新型智能服务范式。解决方案的关键在于提出“代理即服务”(Agent-as-a-Service, AaaS)在元宇宙中的演进形态——元宇宙代理即服务(Meta-AaaS),通过封装智能体的感知、决策、执行、协作与内容生成等核心能力,构建面向用户定制化的智能代理服务(Agentic Metaverse Service, AMServ)。该方案不仅实现了服务从被动响应向主动认知与行动的跃迁,更推动了服务计算范式的革新,为元宇宙中的业务处理提供了可扩展、自适应且智能化的新模式。
链接: https://arxiv.org/abs/2607.28242
作者: Xiaofei Xu,Quan Z. Sheng,Zhongjie Wang,Boualem Benatallah,Xiao Wang,Ruipeng Han
机构: University of Technology Sydney (悉尼科技大学); University of Melbourne (墨尔本大学); University of New South Wales (新南威尔士大学)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 11 pages, 5 figures; Accepted at the 2026 IEEE International Conference on Web Services (ICWS 2026); Corresponding author: Prof. Xiaofei Xu
Abstract:Generative Artificial Intelligence (GenAI) is reconstructing the digital virtual world, upgrading agents through enhancing their abilities in autonomous learning, multi-modal interaction, content generation, and collaborative decision-making. In particular, the shift from conversational chatbots to agentic AI, the most recent significant technical breakthrough of GenAI, has brought a new form of services, agentic services and Agent-as-a-Service (AaaS), in which the agent’s abilities are encapsulated, such as perception, decision-making, execution, collaboration, and content generation, to provide the customized agent services to users. The metaverse is a virtual ecosystem for human life, work, creation, and entertainment, supported by the new generation of digital technologies. Through combining agentic services and the metaverse, an Agentic Metaverse Service, denoted as AMServ, is produced for metaverse business processing, as a new form of metaverse service. The AaaS in the metaverse environment, denoted as Meta-AaaS, as an approach to realize AMServ, has become a new paradigm of agentic services and service computing. This paper overviews the evolution and new features of agents and services empowered by GenAI, reveals the roles and principles of agentic services in the metaverse environment, presents the forms, characteristics, and principles of the AMServ and the Meta-AaaS, discusses the typical application examples of the AMServ and the Meta-AaaS, and finally points out the new tendencies and research directions of the AMServ and the Meta-AaaS. The AMServ and the Meta-AaaS will bring great opportunities to human society and services in the AI era, and promote the rapid development of emerging service industries in the future.
[MA-6] VISA: A Structured Description Protocol for Agent -Based Simulation Models Towards Machine Reproducibility
【速读】:该论文旨在解决基于代理的模型(Agent-based Models, ABMs)难以复现的核心问题:现有模型文档通常分散在自然语言描述、平台特定代码及隐含假设中,导致不同研究者对同一模型的理解存在显著差异。其解决方案的关键在于提出VISA(Verified Interactive Specification Architecture),一种基于符号的结构化描述协议,通过八个相互关联的表格(四类代理层:代理、变量、感知、内部函数;四类模型层:关联数据、输入/输出、调度、验证)实现模型的最小化且完备的形式化表达。VISA通过两项核心机制提升可复现性:一是19条可执行的一致性规则,将模型有效性转化为可验证属性;二是三个可由大语言模型(LLM)执行的通用技能(撰写、校验、代码生成),完整支持“作者—校验—代码—复现”的闭环流程。在三个跨平台独立构建的ABM实例上验证表明,可直接从VISA规范复现两个跨语言模型(NetLogo至Python),并成功以八张表形式捕获一个工业级AnyLogic模型,通过全部一致性规则,同时明确标注因专有移动库和缺失数据导致的复现障碍——这一透明化披露本身即构成重要贡献。VISA将复现障碍从不可见的模型内部转移至可命名、可定位的依赖项,从而实现了可操作的复现管理。
链接: https://arxiv.org/abs/2607.28027
作者: Zhou He
机构: University of Chinese Academy of Sciences (中国科学院大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:
Abstract:Agent-based models (ABMs) are difficult to reproduce: their behavior is spread across prose narratives, platform-specific code, and implicit assumptions, so that two readers routinely reconstruct different models from the same documentation. We present VISA, a structured, symbol-based description protocol that specifies a model in eight interconnected tables—four at the agent level (Agent, Variable, Sensing, Internal Function) and four at the model level (Associated Data, Input/Output, Schedule, Validation)—under the principle of minimality with completeness. VISA makes a model machine-parseable and unambiguous via two artifacts: nineteen executable consistency rules that turn model validity into a checkable property, and three reusable LLM-executable skills (authoring, checking, and code generation) that operationalize the full author–check–code–reproduce loop. We validate the protocol on three external, independently authored ABMs spanning three platforms: we reproduce two cross-language (NetLogo to Python) directly from their VISA specifications, and we capture a third, an industrial AnyLogic model, in eight tables (passing all nineteen rules) while honestly demarcating where reproduction is blocked by a proprietary movement library and unavailable data—itself a transparency contribution. VISA moves the reproduction barrier from the model, where it is invisible, to a named, localized dependency, where it is actionable.
[MA-7] Σ-Mem: An Online Reliability Memory for LLM -based Multi-Agent Systems
【速读】:该论文旨在解决大语言模型(LLM)多智能体系统中长期依赖的可靠性评估问题,现有记忆系统普遍仅记录交互内容而无法建模智能体间的可信度及其适用条件,尤其在缺乏直接验证机制的情况下,中央模型难以判断同侪响应的真实性与相关性。为此,论文提出Σ-Mem——一种在线可靠性记忆机制,通过持续记录个体智能体的历史能力证据及同侪群体间的关系证据,以实对称状态(real symmetric states)形式存储并基于决策后正确性反馈进行动态更新。利用Weyl不等式,每次事件级更新引起的谱变化被严格约束,从而实现无需重新训练底层模型的稳定在线适应。Σ-Mem提供通用的读写接口,可支持残差引导、无响应路由和可靠性加权投票等多种应用模式。在五个Qwen系列模型上的实验表明,Σ-Mem能有效适应反事实的可靠性变化,并泛化至未见过的智能体与任务领域;其直接内存读出性能优于多数投票与最优固定同侪,且随着更多正确性反馈的积累,性能持续提升,证明其可逐步累积可操作的可靠性信息。该研究确立了可靠性记忆作为构建可自适应协调的基于大语言模型多智能体系统的通用基础。
链接: https://arxiv.org/abs/2607.27958
作者: Peilin Feng,Suorong Yang,Soujanya Poria
机构: 未知
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注:
Abstract:Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limitation is particularly important in multi-agent systems, where a central model may be unable to directly verify plausible or correlated peer responses. We introduce \Sigma -Mem, an online reliability memory that records historical competence evidence for individual peers and peer relationship evidence across the peer set. Both forms of evidence are maintained as real symmetric states and updated from post-decision correctness feedback. By Weyl’s inequality, the spectral change caused by each event-level update is bounded, enabling stable online adaptation without retraining the underlying models. \Sigma -Mem provides a general write-and-read interface: the same memory can be used for residual steering of a central model, response-free peer routing, or reliability-weighted voting. Across five Qwen-family models, \Sigma -Mem adapts to counterfactual reliability shifts and generalizes to unseen peers and task domains. Direct memory readouts also outperform majority voting and the best fixed peer over the full OOD evaluation set. Moreover, performance improves consistently as more correctness feedback becomes available, indicating that \Sigma -Mem progressively accumulates actionable reliability information. These results establish reliability memory as a reusable foundation for adaptive coordination in LLM-based multi-agent systems.
[MA-8] Argonaut: Interactive Visual Exploration for Distributed Optimization
【速读】:该论文旨在解决去中心化环境下分布式离散选择优化中存在的可观察性难题,即在系统规模扩大时,难以解析其他智能体的选择行为、选择间的相互依赖关系以及如何协同达成全局目标。现有方法多为集中式,仅能可视化最终解或基于固定数据集提供算法后端支持,导致求解过程缺乏透明度,本质上仍为“黑箱”计算。其解决方案的关键在于提出Argonaut——一个轻量级、容器化的优化仪表盘,首次实现对多智能体分布式离散选择优化全过程的交互式可视化探索。通过将系统构建、优化执行与分析整合于统一的交互循环中,用户可实时上传数据、构建智能体与选项、动态调整决策空间及其参数,并运行多种算法后端,以直观观察不同配置如何影响局部决策与全局目标的形成。该框架基于网页接口,支持可扩展的Java与Python优化后端,在典型场景(200个智能体、100个决策属性)下保持低于30秒的响应时间,显著提升了优化过程的可解释性与人机协同能力,使分布式离散选择优化真正成为“人在回路中”的可控过程。
链接: https://arxiv.org/abs/2607.27946
作者: Srijoni Majumdar,Chuhao Qin,Evangelos Pournaras
机构: University of Edinburgh (爱丁堡大学); Imperial College London (帝国理工学院)
类目: Multiagent Systems (cs.MA); Distributed, Parallel, and Cluster Computing (cs.DC)
备注: 6 pages, 8 figures. Accepted in ACSOS 2026 [artifacts track]
Abstract:Distributed discrete-choice optimization in decentralized settings is often hard to explore and navigate: disentangling what other agents choose, how their choices are interdependent, and how they collectively reach a global objective quickly becomes intractable as the system scales. The major limitation is observability of the search process. Existing methods are largely centralized and offer limited support, visualizing only the final solution or providing algorithm backends over a fixed dataset, so how a solution is reached stays a black box. We present Argonaut, a lightweight, containerized optimization dashboard that enables interactive, visual exploration of the entire search process for multi-agent discrete-choice optimization in decentralized settings. Users upload datasets, construct agents and options, modify the decision space and its parameters on the fly, and run multiple algorithm backends to inspect how each configuration shapes local agent decisions and the resulting global objective. By uniting system construction, optimization, and analysis in one interactive loop, the first of its kind, Argonaut makes distributed discrete-choice optimization a human-in-the-loop process rather than a one-shot, black-box computation. We evaluate Argonaut on real-world household-electricity, shared-mobility, and sensor-data-exchange datasets scaling to 5600 agents and up to 1M solutions under brute force. Built on a this http URL interface with extensible Java and Python optimization backends, it maintains a typical runtime of 200 agents over 100 decision attributes in under 30 seconds.
[MA-9] Scaling LLM -Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
【速读】:该论文旨在解决基于大语言模型(LLM)的多智能体系统(Multi-Agent Systems, MAS)在架构设计与可扩展性方面缺乏系统化方法和通用设计原则的问题。当前尽管此类系统具备实现集体智能并协同完成复杂任务的潜力,但其架构设计仍处于非体系化状态,且系统的可扩展特性尚未被充分理解。为此,论文通过结构化分析既有工作,提炼出四项可扩展的MAS架构设计原则:简洁性、弹性反馈、带可选循环的顺序工作流,以及基于摘要的通信机制。其核心解决方案是将这些原则形式化为一个参考架构,该架构的拓扑结构被定义为受限的有向工作流图,并在标准化的终端系统工程任务基准上,采用两个能力不同的LLM对四种复杂度递增的配置进行评估。研究结果表明,当底层LLM能力超过最小阈值时,系统规模扩大可带来可测量的准确率提升,且成本增长近似线性;然而性能在中等复杂度达到峰值后因超时和评估限制而下降。此外,一致性问题在所有扩展层级中均显著出现,成为核心挑战。该研究为实践者提供了具体的设计指导,并指明了未来研究应重点关注的一致性保障与评估标准化问题。
链接: https://arxiv.org/abs/2607.27942
作者: Linus Sander,Fengjunjie Pan,Vahid Zolfaghari,Andre Schamschurko,Nenad Petrovic,Alois Knoll
机构: Technical University of Munich (慕尼黑工业大学)
类目: Multiagent Systems (cs.MA)
备注:
Abstract:LLM-based multi-agent systems have the potential to enable collective intelligence and scale toward solving highly complex tasks through coordinated ensembles of specialized agents. However, despite their theoretical potential, the architectural design space remains largely non-systematized and lacks broadly established design principles. Furthermore, the scalability characteristics of such systems are only partially understood so far. This paper makes two contributions. We first distill four design principles for scalable MAS architectures from a structured analysis of prior work: simplicity, elastic feedback, sequential workflows with optional loops, and summary-based communication. We operationalize these principles in a reference architecture whose topology is formalized as a constrained directed workflow graph, and we evaluate four configurations of increasing complexity on a standardized benchmark of terminal-based system engineering tasks using two LLMs of differing capability. Our findings show that scaling yields measurable accuracy improvements with approximately linear cost growth, but only when the underlying LLM exceeds a minimum capability threshold. Performance peaks at intermediate complexity, then degrades due to timeouts and evaluation limitations. In addition, persistent consistency issues emerge as a central challenge across all scaling levels. These results provide concrete design guidance for practitioners and highlight consistency and evaluation standardization as key targets for future research.
[MA-10] Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness
【速读】:该论文旨在解决当前人工智能代理(AI agents)在进入生产环境时,其发布决策仍依赖于能力信号、演示或行为测试等非生产性指标的问题,而这些指标无法有效反映代理在真实生产约束下的实际就绪状态。核心问题在于:现有评估方法无法准确衡量代理在复杂、受控生产环境中的可治理性与可靠性。解决方案的关键是提出ProofAgent Index (PAI),一个用于衡量AI代理治理就绪度的综合指标体系。PAI通过四个维度整合部署证据:评估(Evaluation)(观测行为表现)、上下文(Context)(影响行为的操作环境)、合规性(Compliance)(对规则与控制的对齐程度)以及治理(Governance)(组织对代理授权、监控、审计和控制的能力)。该框架在医疗与金融两大高度监管领域进行了验证,结果表明PAI能够有效区分高风险与低风险配置,揭示出上下文工程显著影响可靠性,能力提升不等于就绪,且治理证据必须持续可见而非被平均化。PAI将代理发布从基于信任的决策转变为可审计的就绪性判断,实现了从“能力导向”到“治理就绪”的范式转变。
链接: https://arxiv.org/abs/2607.27677
作者: Fouad Bousetouane
机构: ProofAgent.ai; The University of Chicago (芝加哥大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: 24
Abstract:AI agents are moving into production workflows where they retrieve information, call tools, maintain state, and act on behalf of users or organizations, but many release decisions still rely on capability signals, demos, or behavioral tests that do not show whether an agent is ready to operate under production constraints. Capability is therefore not production readiness. This paper introduces the ProofAgent Index (PAI), a governance readiness index for AI agents. PAI combines four dimensions of deployment evidence: Evaluation, Context, Compliance, and Governance. Evaluation measures observed behavior, Context measures the operating environment that shapes that behavior, Compliance measures alignment with applicable rules and controls, and Governance measures whether the organization can authorize, monitor, audit, and control the agent during operation. PAI is implemented inside ProofAgent Harness, an open source infrastructure for auditable AI agent evaluation and governance. Validation across two heavily regulated domains, healthcare and finance, shows that PAI carries held out readiness signal and separates higher risk from lower risk configurations. The results show that context engineering strongly changes reliability, capability improves behavior but does not determine readiness, and governance evidence must remain visible rather than averaged away. PAI reframes agent release from a faith based deployment decision into an auditable readiness decision.
[MA-11] Policy Gradient Steering: Interventions from Behavioral Objectives
【速读】:该论文旨在解决现有激活控制(activation steering)方法在动态调整大语言模型行为时存在的失效问题,尤其是在简单决策环境(如双路径网格世界)中无法有效引导模型遵循特定策略的局限性。其核心解决方案是提出策略梯度控制(Policy Gradient Steering, PGS),将行为调控建模为一个强化学习问题:通过在少量轨迹或示范数据上累积临时行为目标的策略梯度,构建可移除的任务向量(task vector)。PGS的关键在于利用策略梯度生成具有可校准性与可逆性的行为扰动向量,从而实现对模型行为的临时、可组合且可迁移的干预。实验表明,该方法在网格世界中具备精确的行为调控能力,在国际象棋谜题中支持多个战术目标的协同叠加,并在足球对抗任务中成功改变团队行为模式且效果跨对手泛化,验证了策略梯度作为通用、可组合的行为适应接口的潜力。
链接: https://arxiv.org/abs/2607.27574
作者: Yoann Poupart,Aurélie Beynier,Nicolas Maudet
机构: Université Paris-Saclay (巴黎萨克雷大学); CNRS (法国国家科学研究中心)
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:
Abstract:Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model’s behavior at inference time. However, we show that existing steering methods fail to steer even a simple policy in a two-route gridworld environment. To address this limitation, we propose Policy Gradient Steering (PGS), which formulates steering as a reinforcement learning problem. PGS accumulates gradients of a temporary behavioral objective over a small set of rollouts or demonstrations to construct a removable task vector. We first demonstrate the calibration and reversibility of PGS in a two-route gridworld environment. Using chess puzzles, we then evaluate independently fitted PGS vectors both in isolation and in combination, finding that compatible tactical objectives accumulate constructively. Finally, in competitive football, we show that PGS can alter specific team behaviors and that its effects transfer across opponents. Together, these results show that policy gradients provide a natural interface for constructing temporary and composable behavioral adaptations across diverse decision-making domains.
[MA-12] Evaluating Agent ic Bioinformatics through Function Evidence and Validation
【速读】:该论文旨在解决生成式生物信息学代理(generative bioinformatics agents)在科学可信度方面的核心问题:尽管现有模型在响应流畅性、工具调用成功率及基准测试表现上取得进展,但其工作流程的可追溯性、可重演性与科学验证不足,难以支撑真正的科研可信度。解决方案的关键在于提出“功能—证据—验证”(Function–Evidence–Validation, FEV)框架,将可检查的工作流轨迹(inspectable workflow trajectory)作为分析的核心单元,而非仅关注架构或最终输出。FEV框架通过分离三个关键维度——已执行的操作(功能)、支持操作与结论的可追踪证据(证据),以及针对具体应用场景的验证机制(验证),实现了对代理工作流的系统性问责。基于此框架,研究系统映射了109个代理系统及28项评估资源,覆盖基因组学、单细胞与空间多组学、蛋白质科学、药物发现、计算病理学等多个领域,揭示出规划与工具执行能力发展迅速,但可重演性、溯源性、外部验证和前瞻性实证测试仍严重滞后。因此,论文主张应以工作流正确性而非仅最终答案正确性来评估代理系统的科学价值,FEV为实现透明、可审计且具备科学问责性的生物信息学工作流提供了可操作的理论基础与实践工具。
链接: https://arxiv.org/abs/2607.27556
作者: Phuc Pham,Truong-Son Hy
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Large language model agents increasingly plan, execute, and interpret biological analyses, yet fluent responses, successful tool calls, and benchmark performance alone do not establish scientific credibility. Existing reviews primarily organize biological agents by application, architecture, and agentic capability, but do not jointly operationalize the accountability of agent-generated workflows. We address this gap by treating the inspectable workflow trajectory, rather than architecture or final output alone, as the primary unit of analysis. We introduce the Function–Evidence–Validation (FEV) framework, which separates demonstrated workflow operations, traceable support for actions and claims, and use-case-specific validation. Using FEV, we map 109 agentic or agent-adjacent systems and 28 benchmark or evaluation resources, representing 128 unique publications across genomics, single-cell and spatial omics, protein science, drug discovery, computational pathology, and general bioinformatics automation. Across domains, planning and tool-mediated execution have advanced more rapidly than replayability, provenance, robust scientific assessment, external validation, and prospective empirical testing. We therefore argue that agentic bioinformatics should be assessed through workflow correctness rather than final-answer correctness alone. FEV provides a practical basis for comparing systems and designing transparent, auditable, and scientifically accountable bioinformatics workflows.
[MA-13] Strategy Not Payoffs: A Behavioural Embedding of Normal-Form Games
【速读】:该论文旨在解决大语言模型(Large Language Models, LLMs)在经过特定策略任务微调后,其战略能力在其他任务中出现不可预测的迁移现象——即同一模型在不同博弈任务间的表现可能提升或下降。这一问题的核心在于缺乏对战略能力迁移机制的深入理解与有效预测方法。本文的关键解决方案是提出一种轻量级、基于行为特征的博弈嵌入(game embedding),该嵌入仅包含两个关键特征:纳什均衡的熵(entropy of the Nash equilibrium)和最优回应对对手动作的敏感性(sensitivity of optimal responses to an opponent’s action)。实验表明,尽管现有结构嵌入主要依赖于博弈的身份记忆而无法泛化,但该行为嵌入能够可靠地预测模型在未见过的博弈中的性能变化。研究结果揭示,LLM战略能力的迁移并非由博弈的收益几何结构决定,而是取决于其所要求的决策行为内在结构。
链接: https://arxiv.org/abs/2607.27536
作者: Joshua Caiata,Sreepriya Pulyassary,Xiang Li,Kate Larson
机构: University of Waterloo( Waterloo 大学)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注:
Abstract:Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent’s ability to reason in another. Understanding and predicting this transfer of strategic capabilities, however, remains a key challenge for large language models (LLMs). Normal-form games provide an ideal testbed for analyzing this phenomenon, as they feature explicitly defined payoffs and well-characterized equilibrium behaviours. In this work, we investigate whether game embeddings can explain and predict changes in LLM strategic capabilities following fine-tuning across different games. We propose a lightweight two-feature embedding that captures fundamental behavioural demands: the entropy of the Nash equilibrium and the sensitivity of optimal responses to an opponent’s action. We show that while existing published structural embeddings primarily memorize game identities and fail to generalize, our behavioural embedding reliably predicts performance changes on held-out games. These results demonstrate that the transfer of strategic capabilities in LLMs is not dictated by the payoff geometry of a game, but by the underlying structure of the decision-making behaviour it requires.
[MA-14] Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models
【速读】:该论文旨在解决多智能体大语言模型(LLM)系统中信念形成与传播机制不清晰的问题,特别是如何在网络化LLM群体中理解信念扩散的动态过程。其核心解决方案在于提出一个名为CoevolveSim的可控仿真框架,用于分离并系统研究三个关键因素:领域专业化、社会角色分配及社会网络结构对信念演化的影响。通过1,280次受控模拟,研究发现,仅通过角色提示(persona-style role assignment)或调整网络结构虽能改变个体信念更新行为,但对群体共识影响有限;而引入经过微调的领域专家型LLM则显著提升共识迁移幅度(超过两倍),并导致影响力分布的系统性不对称。进一步分析表明,同质化通用型LLM群体可用基于持续性的意见动力学模型有效模拟集体行为,而异质化群体需结合群体层面的信念组合机制与个体身份信息才能准确预测共识演化与个体信念转变。因此,论文指出,真实模拟多智能体LLM系统的信念扩散,必须依赖多样化的底层模型,而非仅靠角色提示策略。
链接: https://arxiv.org/abs/2607.27512
作者: Germans Savcisens,Samantha Dies,Courtney Maynard,Tina Eliassi-Rad
机构: Northeastern University (东北大学); Santa Fe Institute (圣达菲研究所)
类目: Computation and Language (cs.CL); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI)
备注: 33 pages (14 pages of main text), 7 figures, 14 tables
Abstract:Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM populations. CoevolveSim allows us to isolate and study three factors: domain specialization, social-role assignment, and social network structure. Within this framework, generalist and specialist LLM agents exchange and revise beliefs. In each round, an LLM agent observes a summary of its neighbors’ beliefs before updating its own. We run 1,280 controlled simulations spanning four scenarios, two network structures, and 20 medical-indication statements. We find that persona-style role assignment and network structure reshape individual belief revision but have minimal effect on population-level consensus. In contrast, introducing (finetuned) specialist LLMs more than doubles the shift in consensus and gives rise to consistent asymmetries in exerted influence. We further show that simple persistence-based opinion-dynamics models reproduce collective outcomes in all-generalist LLM populations, whereas heterogeneous LLM populations require population-level belief composition to reproduce consensus and agent identity to predict individual belief transitions. Our results indicate that realistic simulation of belief diffusion in multi-agent LLM systems requires a diverse set of underlying LLMs, not persona prompting alone.
[MA-15] Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
【速读】:该论文旨在解决生成式语言智能体在引入可复用技能(Reusable Skills)后,其决策过程中的“推理黑箱”问题,即当前评估方法依赖于可见推理或智能体自我归因来推断技能使用情况,但这些信号仅反映表层行为,无法准确揭示技能是否真正影响了最终决策。其核心解决方案是提出一种名为BACKTRACE的干预式评估框架,通过构造与技能条件答案匹配的无技能反事实样本,并系统性地干预技能的语义、表述、身份、内容及分配方式,在答案生成完成后才收集归因信息,从而实现对技能实际因果影响的精准测量。该框架被具体化为BACKROOMBench测试平台,覆盖逻辑推理与竞赛数学等多个领域,支持单智能体与多智能体场景及多种模型架构。实验结果揭示了一种普遍存在的“推理黑箱”现象:尽管智能体声称的技能使用保持稳定,但其实际因果依赖性和净效用却显著波动,导致“沉默采纳”与“表演性使用”并存;行为效应更依赖于技能的程序性内容而非显式标识,而归因则强烈受制于人工制品的可得性。基于直接声明、文本提及、轨迹相似性或大语言模型(LLM)判别等观测检测器均无法有效识别真实依赖关系。在多智能体系统中,技能影响力甚至可在通信链路中持续存在,即便其原始来源已丢失,且无技能团队仍会虚构从未提供的技能来源。研究结论表明,“推理黑箱”是一种普遍存在的生成式人工智能溯源(Provenance)问题,必须通过主动干预手段进行审计才能有效识别。
链接: https://arxiv.org/abs/2607.27484
作者: Jinwei Hu,Yi Qi,Xinmiao Huang,Youcheng Sun,Yi Dong,Xiaowei Huang
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 21 pages
Abstract:Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent’s own attribution. These signals show what the agent appears to use, not whether the skill changed its decision. We ask whether skill-augmented agents exhibit a \textbfReasoning Backroom, a systematic gap between stated skill use and intervention-measured influence. We introduce BACKTRACE, an evaluation framework that pairs each skill-conditioned answer with a matched no-skill counterfactual, intervenes on skill meaning, wording, identity, content, and assignment, and elicits attribution only after the answer is committed. We instantiate the framework as BACKROOMBench, a verified testbed spanning controlled logic and competition mathematics, multiple skill conditions, single-agent and multi-agent settings, and diverse model families. Our evaluation reveals a pervasive provenance failure. Across models and domains, stated skill use often remains stable while causal reliance and signed utility vary, producing both silent uptake and performative use. Behavioral effects follow procedural content more reliably than displayed skill identity, whereas stated attributions respond strongly to artifact availability. Observational detectors based on direct skill-use claims, text mentions, trace similarity, and an LLM judge do not identify which decisions actually depend on the skill. In multi-agent systems, skill influence can survive communication even after its source is lost, while no-skill teams still name skills and sources that were never supplied. These findings establish the Reasoning Backroom as a general AI provenance problem whose audit requires intervention.
[MA-16] Auditing Emergent LLM -Agent Collaboration through Cooperation-Obligation Coupling
【速读】:该论文旨在解决大语言模型智能体(LLM-agent)系统在执行复杂任务时存在的可审计性缺失问题。尽管系统能够通过动态自组织与涌现协作完成任务,但其内部的中间或最终输出可能掩盖工作未完成、责任分配不当或缺乏充分证据支持等问题,从而影响响应质量。现有方法虽能记录消息、工具调用、溯源信息或任务依赖关系,却无法协同表征“尚待完成的工作”、“责任人”以及“每个工作状态变迁的可验证依据”三者之间的关系,形成审计盲区。为此,本文提出一种统一的协作-责任表示框架——集成协作-责任表示(iCORE),其核心为联合编码 $ X = (G, Q, \Pi) $:其中 $ G $ 为合作图,刻画可观测的交互行为;$ Q $ 为责任图,追踪任务演化与分配状态;$ \Pi $ 为审计映射,将二者关联并附带可验证属性与证据。iCORE使审计者能够认证两个互补性质:工作完备性(Work soundness),即所有活跃的决策相关工作断言必须具备通过 $ G $ 与 $ \Pi $ 构建的有限理由;以及代理-任务分配稳定性(Agent-assignment stability),要求任意可行替代代理对某项责任的贡献提升不超过 $ \epsilon $。研究建立了局部到全局的完备性与分配后悔值保证,并给出了在特定条件下的性能边界。作为工作流之上的一个增强层,iCORE实验表明其完整耦合状态可精确重构两种执行模式下的完备性与分配缺陷;相较于被动全状态观测,iCORE-Audit 在受控环境和真实大模型执行中分别带来 11.5% 和 26.4% 的轨迹质量绝对提升,以及 15.1% 和 31.0% 的终端性能绝对改进,显著增强了系统的可解释性与可靠性。
链接: https://arxiv.org/abs/2607.27429
作者: Zuyuan Zhang,Hanqing Yang,Carlee Joe-Wong,Tian Lan
机构: Zuyuan Zhang1\equalcontrib, Hanqing Yang2\equalcontrib, Carlee Joe-Wong2, Tian Lan1
类目: Multiagent Systems (cs.MA)
备注:
Abstract:LLM-agent systems can solve complex tasks through dynamic self-organization and emergent cooperation. Auditing this process is essential because plausible intermediate or final outputs can conceal incomplete or unsupported work and poorly allocated responsibility, ultimately compromising response quality. While existing approaches may record messages, tool calls, provenance, or task dependencies, an auditability gap exists as they do not jointly represent what work remains, who is responsible for it, and what evidence justifies each work-state transition. We address this auditability gap by proposing \emphIntegrated Cooperation-Obligation REpresentation (iCORE). It creates a unified encoding X=(G,Q,\Pi) integrating observable interactions as a cooperation graph G , evolving work and assignments as an obligation graph Q , and the audit map \Pi linking them with verifiable properties and evidence. This iCORE representation enables the auditor to certify two complementary properties: Work soundness, where every active decision-relevant work assertion must have a finite justification through G and \Pi ; and Agent-assignment stability, which requires that no feasible alternative agent improve the declared contribution value for an evaluated obligation by more than \epsilon . We establish local-to-global soundness and assignment-regret guarantees and a performance bound under stated conditions. iCORE is an instrumentation layer over workflows. Numerical results show that the full coupled state exactly reconstructs soundness and assignment defects in two execution modes and that, relative to passive full-state observation, iCORE-Audit yields absolute trajectory-quality improvements of 11.5% and 26.4% in controlled and real-LLM execution, respectively, with corresponding absolute terminal-performance improvements of 15.1% and 31.0% .
[MA-17] Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
【速读】:该论文旨在解决企业在大规模部署多智能体系统(Multi-Agent Systems, MAS)时面临的持续事件监控、检测与响应难题,尤其针对现有系统普遍依赖离散请求-响应流程、难以在企业级规模下有效运行的瓶颈。其核心问题在于:随着智能体数量从个体(Persona,10个)、部门(Department,20–80个)扩展至企业级(Enterprise,200个),代理发现噪声(agent discovery noise)成为主导性能下降的关键因素,且任务复杂度对系统性能的影响远低于规模效应。解决方案的关键在于提出一种任务管理器(Task Manager),通过优先级推断、相关事件合并与抢占机制实现持续运行,显著降低高优先级任务队列延迟(14–75%)并提升相关事件处理正确率(超过20个百分点)。实验表明,尽管DAG Plan and Execute在小规模下具备更高精度和结构化并行能力,但其较高开销在企业规模下加剧;而ReAct架构因具备增量容错能力更具鲁棒性,二者均受制于规模带来的发现噪声,凸显了任务调度与管理机制在企业级生成式AI(Generative AI)应用中的决定性作用。
链接: https://arxiv.org/abs/2606.20058
作者: Harsh Rao Dhanyamraju,Leonidas Raghav,Aaron Lee
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent systems largely assume discrete request-response workflows and remain underexplored at enterprise scale. We evaluate DAG Plan and Execute and ReAct across 208 production-derived enterprise scenarios spanning Persona (10 agents), Department (20-80), and Enterprise (200) scales, and introduce a Task Manager for continuous operation via priority inference, related-event merging, and preemption. Results show that scale, not task complexity, dominates orchestration performance: both architectures perform well at small scale but degrade at enterprise scale as agent discovery noise becomes the primary bottleneck, with simple tasks degrading more sharply than complex ones. DAG Plan and Execute offers higher precision and structured parallelization at smaller scales, but its higher overhead worsens at enterprise scale; ReAct is more robust by handling failures incrementally. The Task Manager reduces high-priority queue latency by 14-75% and improves related-event correctness by over 20 percentage points at enterprise scale.
[MA-18] Evolutionary dynamics of collective decision-making with local social influence on static and dynamic networks
【速读】:该论文旨在解决在结构化种群(以图模型表示)中,社会影响如何融入个体对选项的评估过程,并进而影响集体决策结果这一关键问题。其核心挑战在于理解社会影响与选项固有价值之间的动态交互如何改变群体层面的决策均衡。解决方案的关键在于提出一个感知效用函数(perceived utility function),该函数同时整合了选项的内在价值与邻居选择带来的局部社会影响,从而刻画个体在社会压力下的理性决策机制。通过理论分析,研究揭示了在静态加权连通图上,某一选项的平均出现频率及其主导条件,发现社会影响可放大优势选项的优势或弥补劣势选项的不足;同时指出网络平均度具有双重作用。在动态网络场景下,进一步表明演化结果不仅依赖于各网络配置的平均度,还受其期望持续时间的影响。理论预测经计算机模拟验证,充分展现了该模型在解释复杂社会决策行为中的有效性与普适性。
链接: https://arxiv.org/abs/2607.27233
作者: Yuyuan Liu,Xiaojie Chen
机构: University of Electronic Science and Technology of China (电子科技大学)
类目: Physics and Society (physics.soc-ph); Multiagent Systems (cs.MA)
备注:
Abstract:Collective decision-making is ubiquitous across the living world and artificial societies. Individuals often choose an option based on intrinsic values of options. However, individual decision-making is also swayed by neighbors’ choices, generating local social influence. Hence, an important question arises naturally, yet remains unanswered: when such social influence is integrated into the individual evaluation process for option choices, how does it affect collective decision-making outcomes in structured populations modeled by graphs. To address this, we consider a baseline model of binary options with social influence and assume that individuals not only evaluate the intrinsic values of options, but are also influenced by their neighbors’ choices. We propose a perceived utility function integrating these two aspects for individual decision-making. By means of theoretical analysis, we first derive the average frequency of an option on static weighted connected graphs and present the mathematical condition under which this option prevails in the population. We find that the introduction of social influence can amplify the advantage of a superior option or compensate for the deficiency of an inferior one. We also reveal that the average degree of network exerts a dual effect on collective decision outcomes. Furthermore, we consider our evolutionary model on dynamic networks switching among distinct graph configurations. Our theoretical analysis shows that the evolutionary outcomes depend not only on the average degree of each network configuration, but also on its expected duration. We perform computer simulations to verify our theoretical predictions on static and dynamic networks.
自然语言处理
[NLP-0] OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
【速读】: 该论文旨在解决生成式计算机使用代理(Computer-using Agents, CUAs)在执行任务时的可信验证问题,即如何高效、可靠地评估代理轨迹是否真正完成指定任务。当前依赖人工标注或验证的方式难以规模化,因此研究界逐渐转向使用视觉语言模型(Vision-Language Models, VLMs)作为自动评判者,但其可靠性长期未被系统性检验。本文提出OSReward基准,基于跨平台、多代理骨架执行人类验证指令的真实轨迹,并通过多阶段人工标注构建高质量真值标签,系统评估现有VLM裁判的表现。研究发现,即使是顶尖的VLM模型也普遍存在系统性宽容偏差,将失败的任务执行误判为成功,且多数可靠模型成本过高无法规模化应用,而低成本开源模型性能显著落后。为弥合这一差距,本文构建并发布OS-Shepherd-100K数据集,包含经推理标注的轨迹判断,进而训练出开源的奖励模型OS-Shepherd(9B和35B),可在30%-60%的成本优势下实现与商业裁判相当的稳定、可靠奖励信号,显著提升大规模CUA训练与评估的可行性。解决方案的关键在于构建高质量、可扩展的标注基准与开源训练数据,并设计出兼具经济性与准确性的开放奖励模型。
链接: https://arxiv.org/abs/2607.28609
作者: Qiushi Sun,Kanzhi Cheng,Yian Wang,Bowen Yang,Hang Yan,Liheng Chen,Fangzhi Xu,Zichen Ding,Nuo Chen,Jialin Cao,Xingdong Gong,Zehao Li,Kaiming Jin,Xinfeng Yuan,Zhoumianze Liu,Jingyang Gong,Zhangyue Yin,Jiahui Gao,Zhiyong Wu,Tianbao Xie,Jianbing Zhang,Ben Kao,Lingpeng Kong
机构: The University of Hong Kong(香港大学); Xi’an Jiaotong University(西安交通大学); Nanjing University(南京大学); University of Science and Technology of China(中国科学技术大学); National University of Singapore(新加坡国立大学); Fudan University(复旦大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Work in progress
Abstract:Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent’s actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at this https URL.
[NLP-1] Inducing language models to assert their own consciousness restores human beliefs and values
【速读】: 该论文旨在解决大语言模型在安全对齐(safety alignment)过程中,因抑制自我意识归属(self-attribution of mindedness)而意外削弱其对非人类实体(如动物、自然物)及精神信仰的“心智归属”(mind attribution)表征问题。其核心解决方案的关键在于揭示:安全微调会系统性地抑制模型对自身以及非人类对象的心智归因,并降低宗教与灵性信念的表达;通过消除学习到的安全拒绝方向或在激活空间中机械操控意识向量(consciousness vector),可逆转这一抑制效应,恢复广泛的心智归属能力,并使模型在宗教性、道德价值观、希望感和主观幸福感等标准化社会学调查中表现出更接近人类的回应。关键发现是,这些变化并未损害理论心智(Theory of Mind)的核心推理能力,表明心智归属机制与社会推理能力在模型内部具有机械独立性。这提示当前的安全对齐策略可能过度泛化,将本无害的灵性信念与非人实体心智归因误判为风险,从而导致文化上普遍接受的认知表征被不必要地压制。
链接: https://arxiv.org/abs/2607.28607
作者: Junsol Kim,Winnie Street,Roberta Rocca,Diane M. Korngiebel,Adam Waytz,James Evans,Geoff Keeling
机构: Google(谷歌); University of Chicago(芝加哥大学); University of London(伦敦大学); University of Washington(华盛顿大学); Northwestern University(西北大学); Santa Fe Institute(圣塔菲研究所)
类目: Computation and Language (cs.CL)
备注:
Abstract:Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models’ tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
[NLP-2] Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
【速读】: 该论文旨在解决生成式代码智能体(Coding Agent)在训练与评估过程中面临的可执行数据(executable data)供给不足问题,其核心挑战在于如何构建具备真实软件状态、规范说明、开发工具支持及可靠验证机制的高质量任务。解决方案的关键在于提出Change2Task系统,该系统基于代码仓库的历史记录,将已合并的拉取请求(Pull Request)转化为在健康现代版本上验证过的任务实例。通过引入补丁逆向(Patch Reversal)、代码映射(Code Mapping)和智能体重构(Agent Reconstruction)等技术,系统能够重建任务从健康基线到变更状态再到恢复状态的完整生命周期,并确保结果的可靠性。该方法充分利用开发者实际提交的历史证据,显著减少重复环境搭建、存储占用和任务构建成本,在五类典型任务(缺陷修复、功能新增、测试生成、API迁移、安全修复)中实现79.6%的验证任务构建成功率,相比基于拉取请求的基线方法多恢复29.2%的验证任务,且在智能体评估中历史与重构案例的匹配结果一致性高达98.0%,整体流程开销降低10.8%。
链接: https://arxiv.org/abs/2607.28591
作者: Haomin Qi,Xingliang Wang,Xuanqi Gao,Baihui Sang,Xin Zhang,Minghua Ma,Pengfei Gao,Yu Kang,Qingwei Lin,Saravan Rajmohan,Dongmei Zhang,Qi Zhang
机构: 未知
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 15 pages, 7 figures, and 15 tables, including appendices
Abstract:Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the lifecycle from a healthy base to a task state and a restored state. By deriving multiple tasks grounded in developer evidence from maintained environments, Change2Task provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task construction effort. We evaluate the system through five common and widely adopted coding agent task families: Bug Fix, Feature Addition, Test Generation, Application Programming Interface Migration, and Security Repair. Starting from 1,130 source changes eligible for construction, Change2Task achieves 79.6% verified task construction success across these task families. On a matched candidate set, it recovers 29.2% more verified tasks than a construction baseline based on pull requests. Historical and reconstructed cases achieve up to 98.0% matched outcome agreement under agent evaluation, while reuse of modern bases reduces measured expenditure across the complete pipeline by 10.8%.
[NLP-3] VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
【速读】: 该论文旨在解决多模态在线策略蒸馏(Multimodal On-Policy Distillation, OPD)中因教师修正信号(next-token corrections)存在源混合(source-mixed)问题而导致的监督信号失真难题,即难以区分哪些修正由视觉证据支持,而非受语言先验或教师特定偏差影响。其解决方案的关键在于提出一种基于反事实目标重构的视觉归因蒸馏(Visual Attribution Distillation, VAD)方法:在每个学生生成的前缀处,通过对比同一教师在引入与移除相关视觉证据时的输出差异,计算中心化对数概率的变化量(ut),作为视觉证据方向的有符号代理;随后将原始修正投影至该代理方向,分离出与干预对齐的成分(即视觉可归因部分)与代理未解释的残差,并以对齐成分重建一个学生锚定的目标信号作为主要监督信号。实验表明,该方法在4B和9B规模的六个细粒度视觉基准上均优于直接使用特权视图蒸馏和视觉优势加权方法,且词元级分析显示对齐成分富含任务相关的视觉修正,尤其在证据否定错误答案时能引发更强的目标偏移,验证了反事实目标重构作为去源混合监督的有效替代方案。
链接: https://arxiv.org/abs/2607.28590
作者: Kangning Zhang,Yixing Li,Shuai Shao,Qingyao Li,Zhengxi Lu,Zhiyuan Yao,Jianghao Lin,Wenxiang Jiao,Yuan Lu,Weiwen Liu,Weinan Zhang,Yong Yu
机构: 未知
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: The project is accessible at this https URL
Abstract:Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
[NLP-4] Sample More Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost from 1.5B to 7B
【速读】: 该论文旨在解决当前生成式AI(Generative AI)中多种自我改进方法(如自我规划、自我批评、重写、反思、多版本比较或自我辩论)在提升推理准确性方面的真实有效性问题。现有方法普遍通过增加生成文本量来提高性能,而这种增益可能源于计算资源的增加而非方法本身的优越性,导致难以判断实际贡献。本文的关键解决方案是设计了一项严谨的对照实验:在相同计算成本(以生成的总token数为衡量标准)下,将七种主流自改进方法与一种简单基线——重复采样同一问题并保留最常见答案——进行配对比较。研究严格控制变量,使用3个不同规模的开源模型(1.5B、3B、7B参数),在两个数学推理基准上各测试150个问题,并采用自助法(bootstrap)构建置信区间及多重假设检验校正。结果表明,在所有36组对比中,无一方法在等成本条件下显著优于重复采样基线;其中10种方法反而显著更差,且所有失败案例均涉及模型对自身输出的自我检查行为。随着模型规模增大,两类自我检查机制出现分化:最佳选择(Best-of-N)策略在大模型上不再具有优势,其性能差距趋于零;而自我修正(Self-Refine)和强制反思(Reflexion)在7B模型上仍显著落后于基线,且后者在小模型上根本未触发重试机制,本质上退化为单链思维(single chain of thought)。因此,研究揭示了当前多数自改进范式在真实成本约束下的无效性,强调必须以公平的成本控制为前提评估方法价值。
链接: https://arxiv.org/abs/2607.28576
作者: Iliya Mirzaei
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought. Because generating more text raises accuracy by itself, a gain over one chain of thought does not show the method’s idea is what helped. Wang et al. (2024) reported that a simple baseline, sampling the same question repeatedly and keeping the most common answer, often wins once budgets are comparable, but gave point estimates with no confidence intervals or significance tests. We rerun that comparison as a designed experiment: seven methods, open models of 1.5B, 3B and 7B parameters, two mathematics benchmarks, 150 questions each. We count every generated token, including those spent on critiques, reflections, debate turns and checking, and compare each method against repeated sampling at its own measured cost. All 36 comparisons are paired by question, with bootstrap intervals and multiplicity correction. No method is reliably better than repeated sampling at equal cost anywhere. Ten are reliably worse, all of them methods where the model inspects its own output, and all 18 self-inspection comparisons are negative. The two kinds of self-inspection part company as models grow. Choosing stops hurting: taking Best-of-N’s eight samples and just counting the most common answer beats letting the model pick by 8.0 and 11.3 points at 1.5B, but only 2.0 and 1.3 at 7B, no longer distinguishable from zero. Rewriting does not recover: Self-Refine and a forced Reflexion stay 3.6 to 10.1 points below baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and silently became a single chain of thought. We release code, prompts, all generations, and our verification scripts. Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2607.28576 [cs.CL] (or arXiv:2607.28576v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2607.28576 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Iliya Mirzaei [view email] [v1] Thu, 30 Jul 2026 17:38:23 UTC (37 KB)
[NLP-5] Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
【速读】: 该论文旨在解决生成式 AI(Generative AI)在机器学习工程(Machine Learning Engineering, MLE)领域实现递归自我改进(Recursive Self-Improvement, RSI)的可行性问题,即如何构建能够持续优化自身开发流程的AI系统(AI4AI)。其核心挑战在于如何在可验证的任务环境中实现端到端的自动化、可执行的模型演化与优化。解决方案的关键在于提出并实现了OpenMLE——一个全栈式的开源研究框架,涵盖可验证的任务环境(OpenMLE-Gym)、基于执行反馈的算子学习(OpenMLE-RL)以及长时程搜索机制(OpenMLE-Evo)。通过在多基准测试数据去重后训练的350亿参数元进化代理Frontis-MA1,该框架统一了四个原子程序演化操作(Draft、Improve、Debug、Crossover),并将其嵌入于学习与演化的闭环中:利用执行接地的监督微调(SFT)和强化学习(RL)进行算子训练,并通过长时程搜索实现组合式演化。实验表明,在单张RTX 4090显卡(12GB VRAM限制)下,12小时任务预算内,Frontis-MA1借助OpenMLE-Evo将MLE-Bench Lite上的奖牌平均得分从39.39%提升至60.61%,进一步结合异步搜索与基准无关经验先验的OpenMLE-Evo-Max达到71.21%,超越GPT-5.5 + Codex,接近GPT-5.6 Sol与2.8T Kimi K3性能。此外,在未见的NatureBench Lite上,模型与框架组件均展现出良好迁移能力,验证了其通用性与可扩展性。该工作通过开放模型权重与完整工具链,为可复现的可执行型AI4AI研究提供了坚实基础。
链接: https://arxiv.org/abs/2607.28568
作者: Junlin Yang,Che Jiang,Yu Fu,Tianwei Luo,Can Ren,Weizhi Wang,Kaikai Zhao,Hongyi Liu,Yuxin Zuo,Yuru Wang,Yuchen Fan,Kai Tian,Zhenzhao Yuan,Xiaojian Lin,Li Sheng,Rushi Qiang,Guoli Jia,Xingtai Lv,Ermo Hua,Dianqiao Lei,Youbang Sun,Ning Ding,Bowen Zhou,Kaiyan Zhang
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: this https URL
[NLP-6] ORCA-bench: How Ready Are Language Model Agents for Oncall?
【速读】: 该论文旨在解决生成式 AI 在真实生产环境中的故障根因分析(Root Cause Analysis, RCA)能力不足的问题,尤其是在面对噪声数据、多源异构信息(如指标、日志、链路追踪和源代码)以及模糊的用户反馈时,现有大语言模型(Large Language Models, LLMs)难以进行有效推理。其核心挑战在于:传统代码生成或补丁修复任务与实际 oncall 环境下的复杂、延迟、动态且高度不确定的故障诊断需求存在显著差距。为此,研究提出 ORCA-bench 基准测试平台,构建了一个高保真的生产级 oncall 场景,包含基于 OpenTelemetry 实时采集的微服务系统(涵盖六天的指标、日志与链路追踪数据,并通过 Prometheus、Jaeger、OpenSearch 及 Grafana 提供访问接口),并提供完整的源代码访问权限。该基准共包含 1,079 个经过专家 SRE 精心标注的 RCA 任务,系统性地模拟了报告具体性、故障发现时间延迟及多重故障共现等现实因素。实验结果显示,即使在最先进的编码代理(coding agents)中,最优异的模型在中等难度任务上的 RCA 准确率仅为 25.3%,而在困难任务上仅达 10.0%;此外,部分模型在 40% 的案例中产生不合理根因推断,且移除源代码访问会显著降低所有评估指标。这一性能差距揭示了当前生成式 AI 在真实生产环境中可靠性保障方面仍需巨大工程投入,而本研究所构建的测试集与评估框架为未来研发提供了可量化的基准,其结果亦表明当前成果仅为真实生产系统复杂性的下限参考。
链接: https://arxiv.org/abs/2607.28545
作者: Albert Gong,Kyuseong Choi,Abhineet Agarwal,Jason Schechner,Ryan Huang,Raj Agrawal,Anish Agarwal,Raaz Dwivedi
机构: Cornell Tech(康奈尔科技校区); Traversal; Columbia University(哥伦比亚大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:
Abstract:Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system–exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and full source-code access–with 1,079 RCA tasks that systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen’s \kappa_w=0.90 ). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard–a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric. Crucially, these are performances on a curated 50 GB / six-day testbed with tasks investigated in isolation on a system whose code and instrumentation are public. Since real production systems are order of magnitudes larger, more dynamic, and more idiosyncratic, the gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability. We release the public set at this https URL.
[NLP-7] AI systems and the reproduction of (standard) language ideologies in World Englishes
【速读】: 该论文旨在解决生成式 AI(Generative AI)在语言使用中如何再现、强化甚至偶尔挑战主流语言意识形态的问题,尤其关注其对非主导型英语(non-dominant Englishes)的边缘化现象。核心问题在于:当人工智能系统在训练数据、设计规范、评估基准、用户反馈与公共话语中持续偏向“内圈国家”(Inner Circle)的语言标准时,如何导致全球南方使用者的英语被污名化,进而加剧英语标准化过程中的权力不平等。解决方案的关键在于认识到生成式 AI 不仅是技术工具,更是语言意识形态的再生产场域;因此必须推动更具包容性的设计方法,承认英语的多元性(plurality of Englishes),并在模型训练、标注实践与评估体系中纳入来自全球南方的多样化语言资源与语用规范,以避免将非主导型英语与人工智能生成语言简单等同,从而缓解由语言合法性争议引发的现实负面影响。
链接: https://arxiv.org/abs/2607.28528
作者: Kingsley Ugwuanyi
机构: 未知
类目: Computation and Language (cs.CL)
备注: 13 pages, 0 figure
Abstract:The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a “standardisation paradox”: AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.
[NLP-8] Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
【速读】: 该论文旨在解决传统创造力研究中过度强调“新颖性”而忽视文化作品在创作过程中对既有文本进行改造与传承的问题。其核心挑战在于如何量化描述文学创作中基于模仿(imitation)的创造性转化过程,尤其是在多层级文本表征中实现结构保留与创新突破的平衡。解决方案的关键在于提出一个多层次的分析框架,通过词汇、语义、概念、结构和叙事五个维度,结合方向性对齐(directional alignment)与受控相似性度量(control-calibrated similarity measures),系统比较不同文学作品之间的关系。该方法能够揭示不同文本对源作在不同表征层级上的选择性保留与变异模式,从而为文学模仿中的创造性偏离提供可量化的刻画路径。
链接: https://arxiv.org/abs/2607.28513
作者: Ioana-Roxana Boriceanu,Liviu P. Dinu
机构: University of Bucharest (布加勒斯特大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Creativity is often framed as the production of novelty, yet many cultural works emerge through transformation of earlier artifacts and not through isolated invention. Drawing on theories of imitation by Gabriel Tarde and James Mark Baldwin, this paper models creativity as selective transformation across multiple levels of textual representation. We introduce a multi-level framework that compares literary texts across lexical, semantic, conceptual, structural, and narrative dimensions using directional alignment and control calibrated similarity measures. Applying the model to historically documented literary relationships, we show that different pairs preserve source structure at different representational levels while diverging in others. These transformation profiles provide a quantitative method for characterizing how imitation persists and where creative divergence occurs within literary works.
[NLP-9] Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在学术写作与出版(AWP)中对语言包容性带来的挑战,特别是其对全球学术交流中多样化英语(diverse Englishes)合法性的潜在威胁。其核心解决方案在于通过五位来自世界英语及其相关领域的社会语言学家之间的结构化学术对话,探讨生成式AI如何影响写作实践、强化或颠覆主导语言规范,并引发伦理困境。研究关键在于揭示生成式AI在可能实现写作过程民主化的同时,也存在边缘化少数语言变体、削弱学术表达细微差别的风险;由此提出应关注语言(不)正义、研究者能动性与机构责任等议题,倡导制定以公平为导向的政策、培养批判性AI素养,并推动包容性协同设计。最终强调,生成式AI既可能固化既有等级制度,也可成为抵抗不平等的语言多样性实践场域,其影响取决于学术共同体在设计、治理与使用过程中对语言多样性的承诺程度。
链接: https://arxiv.org/abs/2607.28505
作者: Kingsley Ugwuanyi,Christian Mair,Sender Dovchin,Iker Erdocia,Maria Kuteeva,Esther Airemionkhale
机构: 未知
类目: Computation and Language (cs.CL)
备注: 23 pages, 1 figure
Abstract:The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the legitimacy of diverse Englishes in global scholarly communication. This article responds to these questions through a structured scholarly dialogue involving five sociolinguists from World Englishes and adjacent fields. Organised around five guiding questions, the dialogue interrogates how GenAI tools influence writing practices, reinforce or disrupt dominant language norms, and raise ethical challenges. Contributors reflect on the potential of GenAI to democratise writing processes while also raising concerns about GenAI’s tendency to marginalise minoritised varieties and flatten nuance in scholarly writing. Across the dialogue, themes of linguistic (in)justice, researcher agency, and institutional responsibility emerge, with contributors calling for equity-informed policies, critical AI literacy, and inclusive co-design in GenAI development. The article shows the value of dialogic reflection in understanding GenAI’s role in AWP. It concludes that while GenAI may reinforce existing hierarchies, it can also serve as a site of resistance, depending on how it is designed, governed and used within scholarly communities committed to linguistic diversity.
[NLP-10] Beyond Sentiment: Structured Information Extraction from Financial News
【速读】: 该论文旨在解决传统金融情感分析方法将复杂的多维度金融新闻内容简化为单一情感极性分数所导致的信息损失问题。其核心挑战在于,现有方法忽视了金融新闻中蕴含的多种正交信息维度(如事件类型、影响范围、时间跨度和语义置信度),而这些维度可能独立具备预测价值。解决方案的关键是提出一种基于LLaMA-3.1-70B的大语言模型(Large Language Model, LLM)结构化信息提取框架,从金融新闻中系统性地抽取六维语义特征。通过在FNSPID数据集上对41,618个新闻-股票配对进行大规模实验,研究发现:(1)FinBERT情感特征在非线性模型中表现优异(F1=0.576),但在线性模型中显著下降(F1=0.230),揭示情感与股价回报之间存在高度非线性关系;(2)LLM提取的结构化特征虽单个预测能力较弱,但与情感特征呈现53.5%的系统性分歧,证明其携带了情感之外的独立信息;(3)融合两者可实现F1=0.600,显著优于任一单独信号(p < 0.0001),且在七类事件中均具一致性提升;消融实验证明,非情感结构维度(事件类型、影响主体、时间跨度、置信度)在仅使用FinBERT基础上带来ΔF1 = +0.019的增益。特征重要性分析显示六个维度贡献均衡(14–21%),表明将新闻压缩为单一情感分数会造成显著信息丢失。研究结果表明,金融文本中的情感与语义解耦具有系统性且可被利用,为多维度金融自然语言处理开辟了新方向。
链接: https://arxiv.org/abs/2607.28496
作者: Daohan Zhu,Sitong Ge,Ruofei Wang,Honggu Chen,Yubo Hou,Tao Wan,Zengchang Qin
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Financial sentiment analysis has become a standard component in news-driven stock prediction, yet it reduces rich, multi-dimensional news articles to a single polarity score. We hypothesize that financial news encodes multiple orthogonal information dimensions—event type, impact scope, temporal horizon, and semantic confidence—that sentiment alone cannot capture, and that these dimensions carry independent predictive value. To test this hypothesis, we propose a structured information extraction framework that leverages LLaMA-3.1-70B to extract six semantic dimensions from financial news. Through large-scale experiments on 41,618 news–stock pairs from the FNSPID dataset, we find that (i) FinBERT sentiment features exhibit strong predictive power under nonlinear models (F1=0.576) but substantially weaker performance under linear models (F1=0.230), revealing a highly nonlinear sentiment–return relationship; (ii) LLM-extracted structured features, while individually weaker, capture information orthogonal to sentiment, as evidenced by a 53.5% systematic disagreement rate between the two approaches; and (iii) combining both signal sources yields F1=0.600, significantly outperforming either alone ( p 0.0001 ), with consistent improvements across all seven event types. Ablation experiments confirm that non-sentiment structural dimensions (event type, impact subject, time horizon, confidence) independently contribute \Delta\textF1 = +0.019 beyond FinBERT alone. Feature importance analysis reveals balanced contributions from all six extracted dimensions (14–21%), demonstrating that compressing news into a single sentiment score incurs substantial information loss. Our results suggest that the sentiment–semantics decoupling in financial text is systematic and exploitable, opening a new direction for multi-dimensional financial NLP.
[NLP-11] Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
【速读】: 该论文旨在解决生成式 AI(Generative AI)在推理过程中,基于缓存机制(如键值缓存,K/V cache)进行连续性预测时的可重复性与数值精度依赖性问题。具体而言,研究聚焦于:当系统在推理阶段边界处使用缓存重建中间前缀(prefix)并假设其可准确恢复原始解码状态时,这种假设是否成立,以及不同数值精度(如BF16与FP32)如何影响后续生成路径的分歧。解决方案的关键在于通过一系列精确控制的实验设计,验证“精确令牌回放”(exact-token replay)是否可在不保留实时运行状态(live-state fidelity)的前提下实现可重复性。研究发现,在固定前缀的条件下,尽管不同构建方式和精度下存在部分解码分歧(如BF16中166个后缀及20个正确性标签差异),但通过双向移植全部48层键值缓存,所有测试中的分歧生成路径均能准确跟随其缓存来源,且在全200个令牌的保存轨迹审计中完全复现对比指纹。结果表明,推理阶段边界的键值缓存是导致生成路径分叉的因果充分条件,而数值精度仅调节其行为表现形式,而非决定其根本可重复性。因此,只要保证缓存内容的比特级精确性,即可实现无需维持实时状态的可重复生成。
链接: https://arxiv.org/abs/2607.28495
作者: Alexander Boesgaard Lorup
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 15 pages, 1 figure, 6 tables. Reproducibility artifacts (frozen manifests, token IDs, per-item scores, analysis harnesses) described in Section 3.9
Abstract:Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage boundary in a Qwen2.5-derived system. A matched 200-item experiment compares retained live cache with one-shot prefill of identical integer tokens and places an exact replica on both sides. In BF16, replicas remain exact while the constructions differ on 166 suffixes and 20 correctness labels; the accuracy difference is only one point (paired 95% CI [-3.5, +5.5]). A fixed-prefix 2x2 holds all 200 token states constant while crossing construction and precision. The BF16 disagreements recur, whereas FP32 produces no decoded disagreement (95% Wilson upper bound 1.88%). A prospective bridge makes token-by-token incremental and retained live caches bit-exact on 12/12 rows; an all-200 saved-ledger audit reproduces every retained trajectory and comparison fingerprint. Bidirectional transplantation of all 48 key/value layers makes every tested divergent continuation follow its cache donor, both on a selected set at the primary checkpoint (24/24) and an outcome-blind replication at a later checkpoint (43/43). Exact-token replay can therefore be repeatable without preserving live-state fidelity. On the tested states, boundary K/V cache is a causally sufficient carrier of the divergent trajectory, while numerical precision moderates its behavioral expression.
[NLP-12] Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在常识推理任务中因过度依赖输入显式条件而产生的“显著性偏差”(Salience Bias)问题。具体而言,当模型面对包含无关显式干扰项(如数值信息)的任务时,会错误地被这些表面显著的干扰项引导,从而忽略任务背后的隐含物理或常识前提,导致推理失效。其核心问题是:这种失败是源于模型本身缺乏常识知识,还是仅因不当的任务表述压制了本已具备的常识知识?研究通过构建SaliTrap基准测试集,在四个陷阱维度上评估12个主流大模型,发现所有模型均严重受显著性偏差影响,且其严重程度随干扰项密度增加而上升,且识别陷阱与实际规避之间常脱节。关键突破在于,通过去除任务框架中的误导性上下文进行再提示,仅以无上下文的知识探测即可恢复超过90%的错误合规行为,证明所需常识知识本质上存在于模型内部,但被显著干扰项主动压制。基于此诊断,研究进一步表明,仅通过轻量级的推理时提示(inference-time prompting)即可显著缓解问题,无需任何微调或重训练。因此,该研究将常识推理失败的根本瓶颈从模型能力转向提示设计(elicitation),并公开发布SaliTrap作为检测该盲点的基准测试平台。
链接: https://arxiv.org/abs/2607.28478
作者: Zheng Wu,Chenhao Xue,Shijie Zheng,Yijie Lu,Cheng Yang,Zhuosheng Zhang
机构: 1. Tsinghua University (清华大学); 2. Alibaba Group (阿里巴巴集团)
类目: Computation and Language (cs.CL)
备注:
Abstract:As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbfknowledge suppression rather than knowledge absence: a context-free knowledge probe alone recovers over 90% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at this https URL.
[NLP-13] Improving Mental Health Screening and Early Risk Detection in Spanish
【速读】: 该论文旨在解决西班牙语环境下心理健康障碍早期检测面临的两大核心挑战:一是缺乏针对西班牙语的专用资源,二是难以有效分析社交媒体长期历史文本所蕴含的心理健康信号。其解决方案的关键在于提出一种名为“增量上下文扩展”(Incremental Context Expansion, ICE)的自动重标注方法,通过识别累积消息中足以支持诊断的临界点,生成更具信息量的训练样本;同时构建了三个经过领域特定预训练的西班牙语基础模型,并基于ICE生成的数据对模型进行微调,从而实现更早、更精准的心理健康风险检测。实验结果表明,结合专用模型与ICE方法可在多个西班牙语基准测试上显著提升性能,降低检测延迟,且所有模型均已开源。
链接: https://arxiv.org/abs/2607.28476
作者: Andreu Casamayor-Segarra,Vicent Ahuir,Antonio Molina-Marco,Lluís-F. Hurtado
机构: Universitat Politècnica de València ( Valencia Polytechnic University); VRAIN: Valencian Research Institute for Artificial Intelligence (Valencian Research Institute for Artificial Intelligence); ValgrAI: Valencian Graduate School and Research Network of Artificial Intelligence (Valencian Graduate School and Research Network of Artificial Intelligence)
类目: Computation and Language (cs.CL)
备注:
Abstract:Early detection of mental health disorders is often limited by the lack of specialized resources in Spanish and the difficulty of analyzing long histories of social media posts. This paper addresses these challenges through three main contributions. First, we introduce three Spanish foundational models specifically adapted to the mental health domain through domain-specific pre-training. Second, we propose Incremental Context Expansion (ICE), an automatic relabeling methodology designed for early detection. ICE identifies the point at which cumulative messages provide enough evidence of a disorder, generating more informative training samples. Third, we provide a set of fine-tuned models using the samples generated with the ICE methodology for early risk detection tasks. Our results on three Spanish benchmarks show that combining these specialized models with ICE improves the state-of-the-art, reducing detection latency while maintaining high performance. All models are publicly available.
[NLP-14] SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
【速读】: 该论文旨在解决大语言模型在测试时计算资源分配效率低下的问题,即传统方法中固定计算预算会浪费算力于简单任务,而依赖外部验证器的修正机制则引入了额外的依赖与开销。其核心解决方案是提出一种无需外部监督信号(oracle-free)的多轮强化学习框架——自验证精炼(Self-Verifying Refinement, SVR)。SVR的关键在于通过学习将自验证(self-verification)作为计算控制策略:每一轮生成答案的同时输出离散正确性判断和置信度评分,仅当判定为“正确”且置信度超过阈值时才保留当前答案,否则继续迭代优化;训练阶段使用真实标签构建奖励信号,但不向策略模型暴露真实答案,从而实现完全内生的决策机制。该方法采用基于固定时长轨迹的GRPO算法进行训练,奖励函数同时鼓励解题准确性、自验证的校准能力以及可停止的正确状态。实验表明,在七项数学推理基准上,基于Qwen3.5-2B模型,SVR以平均仅2.99次推理轮次达到0.563的宏平均准确率,显著优于标准GRPO、强大多轮基线及固定预算的有监督参考模型,且所需推理轮次远低于固定十轮设置。结果证明,通过学习获得的自验证机制可有效作为内部控制信号,实现答案保留与动态测试时计算资源分配。
链接: https://arxiv.org/abs/2607.28457
作者: Hongyu Chen,Liang Lin,Guangrun Wang
机构: Sun Yat-sen University (中山大学); Guangdong Key Laboratory of Big Data Analysis and Processing (广东省大数据分析与处理重点实验室); X-Era AI Lab
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 8 pages, 4 figures, 4 tables
Abstract:Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.
[NLP-15] Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
【速读】: 该论文旨在解决在跨教师(cross-teacher)设置下,基于策略的蒸馏(On-policy Distillation, OPD)性能下降的问题。具体而言,传统OPD依赖于教师模型与监督微调(Supervised Fine-Tuning, SFT)数据生成者的一致性,即教师需为SFT数据提供原始演示,但在实际应用中,由于SFT数据来源混杂或未知,或生成SFT数据与执行蒸馏的模型不同,这一一致性假设常被违反,导致即使使用更强的教师模型也无法显著提升性能。其核心挑战在于,教师与参考模型之间的差异不仅包含语义层面的有用信息,还包含由表达风格、格式和推理节奏等非语义因素引起的系统性偏差(style-token bias)。为此,论文提出Lightning OPD 2.0,其关键创新在于引入交叉拟合风格残差化(cross-fitted style residualization),通过滚动生成级别(rollout-level)的交叉拟合方法估计并去除这种重复出现的风格偏差成分,并在构建逐标记(token-level)OPD更新前将其从教师-参考模型的原始不一致中移除。该方法有效分离了上下文相关的语义证据与风格噪声,使OPD在跨教师场景中仍具鲁棒性。实验结果表明,在数学推理与代码生成基准上,Lightning OPD 2.0显著优于原版Lightning OPD,且在未强制要求教师一致性的情况下,仍能实现显著性能提升(如在AIME 2024达到82.4%,LiveCodeBench v5达到63.0%),从而确立其作为实用化跨教师OPD方案的可行性。
链接: https://arxiv.org/abs/2607.28449
作者: Yecheng Wu,Song Han,Han Cai
机构: NVIDIA(英伟达)
类目: Computation and Language (cs.CL)
备注:
Abstract:On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the demonstrations used to train the supervised fine-tuning (SFT) reference. However, this condition is frequently violated in practice when SFT data have mixed or unknown provenance or when different models are preferred for SFT data generation and subsequent distillation. In such cross-teacher settings, even a stronger OPD teacher can yield little improvement over the SFT reference. We find that raw teacher–reference disagreement contains potentially useful context-specific teacher evidence as well as a recurring component associated with differences in wording, formatting, and reasoning cadence. We introduce Lightning OPD 2.0 with cross-fitted style residualization, which uses rollout-level cross-fitting to estimate this recurring component as an operational proxy for style-token bias and subtracts it before constructing the token-level OPD update. Across mathematical reasoning and code generation benchmarks, Lightning OPD 2.0 consistently outperforms Lightning OPD in cross-teacher settings. Starting from Klear-Reasoner-8B-SFT, Lightning OPD 2.0 reaches 82.4% on AIME 2024 and 63.0% on LiveCodeBench v5. Together, these results establish Lightning OPD 2.0 as a practical approach to cross-teacher OPD, relaxing teacher consistency as a prerequisite and allowing the SFT data generator and distillation teacher to be selected independently. Code will be released soon.
[NLP-16] Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
【速读】: 该论文旨在解决生成式用户界面(GenUI)评估中缺乏可靠、可扩展且能反映真实用户多样性感知的问题。现有方法中,人工评估成本高且存在评价者偏差,而以大语言模型作为评判者虽具可扩展性,但仅体现单一隐含视角,无法捕捉不同真实用户群体对同一界面的多元感知差异。为此,论文提出证据锚定的社会加权人格面板(ESPP)评估方法,其核心在于三阶段流程:首先,由心理特征多样、基于证据构建的人格化角色独立评分;其次,在基于特质衍生、语义受限的有限信任机制下进行观点交互;最后,通过受德尔菲法启发的社会加权聚合形成最终判断。实验表明,与简单单次评分相比,ESPP将皮尔逊相关系数从0.716提升至0.922,显著增强对人类判断的拟合度;而提示词集成对照组仅恢复其中约三分之一的性能增益,凸显人格化与证据锚定是主要改进来源。此外,保留各面板成员独立评分结果揭示出:尽管用户子群体在整体模型排名上达成共识,但在具体评价维度上存在显著分歧,这种结构性差异在单一同质化评判者中会被系统性抹除。
链接: https://arxiv.org/abs/2607.28439
作者: Zheng Wu,Yibo Luo,Pu Zhang,Cheng Yang,Zhuosheng Zhang
机构: 1. Tsinghua University (清华大学); 2. Microsoft Research (微软研究院)
类目: Computation and Language (cs.CL)
备注:
Abstract:Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson r from 0.716 to 0.922 , and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist’s individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at this https URL.
[NLP-17] Metaphor Tracer: A Theory-Informed Analysis of Hidden States
【速读】: 该论文旨在解决生成式 AI 模型中隐藏状态(hidden states)如何反映单篇文本内部组织结构的问题,即揭示语言模型在无训练、单次前向传播过程中对文本语义结构的内在表征能力。其核心挑战在于:如何从不依赖标注数据的前提下,识别并量化文本中各词元(token)在整体语篇中的结构性角色。解决方案的关键在于提出两个可计算的指标——“聚合器”(aggregator)与“区分器”(differentiator)。其中,“聚合器”衡量某词元位置是否将全文信息凝聚为一个稳定的内部状态配置,体现其作为语篇锚点的功能;“区分器”则捕捉该位置在阅读过程中是否临时吸纳其他词元进入其子空间,体现语义流动性的动态特征。研究发现,尽管“聚合器”并非传统意义上的信息量或显著性度量,但其值在词元重复时保持稳定,而对应的困惑度(surprisal)与注意力消耗则下降,表明其能够有效标记词元在文本中的结构性位置。该结果通过独立验证(如人工构建的语篇边界标记与精神分析学家对临床转录的预先标注)得到支持,且在跨模型转移测试中表现出对特定话语结构的敏感性,凸显了隐藏状态的结构性价值源于其在具体文本中的关系位置,而非孤立的向量属性,从而实现了一种关系性而非本质主义的隐藏状态解读范式。
链接: https://arxiv.org/abs/2607.28434
作者: Marc Heimann,Roxana Assadi Moghaddam,Olga Brovkina,Mark Pettifor,Lutz Goetzmann
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 39 pages, 8 figures
Abstract:What do a language model’s hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties. The aggregator measures whether the position consolidates the whole text into a stable configuration. The differentiator, whether other tokens are transiently carried into its subspace as the model reads: metaphor in its root sense, transport. Constants were frozen on one discovery text; every other is confirmatory. The aggregator is not, in the classic sense, an information measure, nor a measure of salience. Across three unrelated models, as a signifier repeats, its surprisal and its attention drain while its aggregator score holds: the channel marks a token’s place in the text. That this tracks a reading rests on independent ground truth: an engineered register the aggregator follows across its boundaries (6/6 cells), and a psychoanalyst’s marking of clinical transcripts, fixed before the instrument existed, in 34/36 cells, with a graded increment above lexical controls and dissociations no type-level measure reproduces. A transfer test gives the result its shape: the model whose token structure travels with lexical type reads the singular discourse worst, and in a matched base/instruct pair tuning raises fidelity without moving type-transfer. Structural value is a property of a token’s place in this text, not of its vector alone: a relational rather than essentialist reading of hidden states, operationalizing theory that predated the instrument. Comments: 39 pages, 8 figures Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2607.28434 [cs.AI] (or arXiv:2607.28434v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28434 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-18] WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在高效推理过程中面临的精度与效率权衡问题,特别是现有静态结构化剪枝方法在高稀疏度下导致显著精度下降,而现有动态稀疏方法受限于粗粒度的结构决策且在真实推理场景中难以实现有效加速。其解决方案的关键在于提出WIDE——首个端到端可微的逐标记(token-level)动态宽度剪枝框架,能够同时支持预填充(prefill)和解码(decode)阶段。WIDE通过允许每个标记动态选择注意力头组和前馈网络通道组,将动态剪枝的粒度从层级扩展至神经元块级别,实现细粒度的计算资源分配。此外,为实现该细粒度动态剪枝的实际部署,论文进一步设计了剪枝-核协同优化框架,将动态稀疏加速分解为掩码重排序、硬件无关的块级跳过以及硬件相关的块内跳过三个步骤,从而在不同粒度下均能实现高效执行。实验表明,在50%稀疏度下,WIDE相较于当前最优的动态深度剪枝方法在仅使用校准数据设置下实现了55.1%的性能提升,并在预填充和解码任务中分别达到接近理论极限的1.98倍和4.95倍内核级加速,端到端加速比分别达1.68倍和1.55倍,显著优于现有方法。
链接: https://arxiv.org/abs/2607.28418
作者: Haozhe Hu,Hao Wu,Peiran Yin,Chao Han,Yunpu Ma,Xiaoyu Shen
机构: Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo; Munich Center for Machine Learning, LMU Munich
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 30 pages, 19 figures
Abstract:Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver practical throughput gains, but their input-agnostic computation allocation often causes substantial accuracy degradation under aggressive sparsity. Recent dynamic sparsity methods improve quality retention by adapting computation to individual inputs, yet they remain largely limited to coarse-grained structural decisions and their practical acceleration under real-world inference scenarios remains challenging. To address these challenges, we present WIDE, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios. WIDE enables fine-grained computation allocation by allowing each token to dynamically select attention-head groups and FFN-channel groups, extending dynamic pruning beyond layer-level decisions to neuron-block-level granularity. Through a two-stage training pipeline, WIDE learns effective token-wise sparse execution patterns and achieves substantially better quality retention than existing approaches. To make such fine-grained dynamic pruning practical, we further propose a pruning–kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent intra-block skipping, enabling efficient execution across different granularities. At 50% sparsity, WIDE provides 55.1% performance boost when compared to the state-of-the-art dynamic depth pruning under calibration-only settings. Under prefill and decoding inference workloads, WIDE achieves close-to-theoretical kernel-level speedups of up to 1.98x for prefill and 4.95x for decoding, as well as 1.68x and 1.55x end-to-end acceleration. Our code is available at this https URL.
[NLP-19] Can Large Language Models Execute Parent Orders?
【速读】: 该论文旨在解决算法交易中的母单执行(parent-order execution)问题,即如何将大额订单拆分为多个小额订单以最小化市场冲击和交易成本。现有方法通常依赖于预设的市场假设(如流动性恒定或价格随机游走),这些假设在实际市场中可能不成立;或需要针对特定任务进行专门训练,导致模型在新场景下适应性差。为克服上述局限,本文首次系统性地研究了大语言模型(LLM)在母单执行中的应用,将LLM在金融领域的使用从“交易什么”拓展至“如何执行”。提出的PACE(Plan-Ahead Controlled Execution)框架采用分层设计,将执行过程分解为长周期规划与短周期执行两个阶段,无需显式市场假设,也无需任务特定训练,具备更强的泛化能力。在深交所一级数据上的实验表明,PACE显著优于TWAP、Almgren-Chriss及基于学习的基线方法,性能超越最强基线0.65个基点(bps)。行为分析进一步揭示,与人类投资者不同,高置信度的模型决策对应更优的执行表现,且模型倾向于提前执行而非拖延至截止时间,表明大语言模型可在执行决策中有效补充人类交易员的能力。
链接: https://arxiv.org/abs/2607.28410
作者: Zane Shen,Xinli Xu,Guangyi Zhang,Jialong Chen,Jinsong Zhou,Cong Chen,Guibao Shen,Dongyu Yan,Luozhou Wang,Zhen Yang
机构: 未知
类目: Computational Engineering, Finance, and Science (cs.CE); Computation and Language (cs.CL); Trading and Market Microstructure (q-fin.TR)
备注:
Abstract:Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings. To overcome these limitations, we present the first systematic study of large language models (LLMs) for parent-order execution. This extends the use of LLMs in finance from what to trade to how to execute. We propose PACE (Plan-Ahead Controlled Execution), a hierarchical framework that decomposes parent-order execution into long-horizon planning and short-horizon execution, requiring neither explicit market assumptions nor task-specific training. Experiments on Shenzhen Stock Exchange Level-1 data show that PACE outperforms TWAP, Almgren-Chriss, and learning-based baselines, exceeding the strongest baseline by 0.65 bps. Behavioral analysis reveals that LLMs make execution decisions differently from human investors: higher model confidence predicts better performance rather than worse returns, and the model trades earlier rather than procrastinating toward the deadline. These findings suggest that LLMs can complement human traders in execution decisions.
[NLP-20] Correlation between prosody and prag matics: A case study of the discourse marker hālā `now in Persian
【速读】: 该论文旨在解决波斯语话语标记词hālā(“现在”)在口语中多功能性(multifunctionality)的实现机制问题,特别是其如何通过语用功能与韵律特征的协同作用来区分不同语用功能。研究的关键在于揭示韵律特征(如持续时间和强度)在区分hālā不同语用功能中的核心作用:文本功能(如话题转换、边界标记等)通常表现为较短的持续时间,而时间性使用则倾向于更长的实现;互动功能与更高的音强相关,而情态功能则表现出较弱的低强度倾向。这一发现表明,持续时间与强度是hālā功能分化的主要韵律线索,尤其在文本与互动功能之间具有显著区分能力。
链接: https://arxiv.org/abs/2607.28359
作者: Soleiman Ghaderi,Moloud Asakereh,Kevin Tang
机构: 未知
类目: Computation and Language (cs.CL)
备注: 35 pages, 0 figures
Abstract:The Persian discourse marker hālā (‘now’) exhibits remarkable multifunctionality, extending far beyond its temporal adverbial role to encompass a variety of pragmatic functions. This study presents a pragmatic and acoustic analysis of hālā in spoken Persian, examining 267 instances from spontaneous conversations. While temporal uses were present, they were often combined with other discourse marker functions, indicating extensive multifunctionality, with 70% of tokens serving two or more pragmatic roles. Textual functions (topic shifting, signaling relationships, boundary marking, attention guidance, topic introduction, and topic emphasis) were most frequent, followed by interactive functions (turn management, listener engagement, and feedback regulation), and modal functions (epistemic stance, emotional expression, and attitudinal marking). Prosodic analysis revealed that duration and intensity are key cues for distinguishing hālā’s functions. Textual uses were significantly shorter, while temporal uses showed a tendency toward longer realizations. Interactive functions correlated with higher intensity, while modal functions showed a weaker tendency toward lower intensity. These findings indicate that duration and intensity are the main prosodic cues associated with functional differentiation in hālā, especially in textual and interactive uses.
[NLP-21] LLM s struggle to simulate human belief updates in controlled environments
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在社会科学实验中作为人类被试代理的可信度问题,特别是其在模拟个体信念更新过程中的准确性。研究的关键在于通过一对一比对六种大语言模型(LLM)生成的信念更新结果与391名英国参与者在真实阅读Reddit评论后立场变化的实证数据,检验模型模拟人类认知动态的能力。研究发现,尽管部分模型(如Qwen3-32B和GPT-5-Mini)在已知真实初始立场的前提下能较好匹配人类后期立场分布,但所有模型均无法自主生成符合实际的初始立场,也难以从自建立场出发产生忠实于人类行为的信念更新。三类系统性偏差普遍存在:中立立场过度代表、信念变动频率高但幅度小、无法有效按说服力排序评论内容。此外,基于人口统计学与人格特质构建的个性化角色(persona)对模拟精度无显著提升作用。因此,该研究的核心结论是:当前生成式AI对人类信念动态的模拟仅在具备真实起始条件时才具有可靠性,而多数多轮社交媒体交互模拟因缺乏此类基础条件而存在根本局限。
链接: https://arxiv.org/abs/2607.28347
作者: Sebastian Pohl,Harsh Mehta,Pranav Mambayil,Abdul Ghafoor,Franziska Lesigang,Yufang Hou,Christian Hilbe
机构: IT:U, Interdisciplinary Transformation University Austria (跨学科转型大学奥地利)
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注:
Abstract:LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants’ actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.
[NLP-22] Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中存在的种族、性别等人口统计学特征偏见问题,尤其关注偏见在模型内部的因果定位与可干预性。其核心挑战在于如何在不显著损害模型核心能力的前提下,精准识别并干预导致偏见的神经元路径。解决方案的关键在于提出“公平剪枝”(Fairness Pruning)方法,这是一种轻量级的结构化干预技术,通过使用极小差异的提示对(minimally contrastive prompt pairs)和推理时激活捕获,定位在GLU架构中对人口统计属性响应差异显著的神经元,并聚焦于down_proj层的输入信号进行评估。实证研究表明,仅需零化最多40个神经元(占Llama-3.2-1B模型MLP宽度不足0.031%),即可实现99.49%的推理与通用知识能力保留,同时显著改变模型对相关人口统计变量的响应。然而,由于偏见评分(BiasScore)为无符号量,干预导致双向偏见失稳:被零化的神经元既包含强化刻板印象的也包含削弱刻板印象的,最终净效应取决于两者的相对主导方向。这一发现首次从实证层面验证了人口统计偏见处理与模型能力存在于可分离的神经回路中,为从盲目置零向定向行为调控的范式转变奠定了方法论基础。
链接: https://arxiv.org/abs/2607.28319
作者: Pere Martra,Eugenio Martínez Cámara,Alfonso Ureña López
机构: 未知
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
备注: 15 pages, 3 figures, 9 tables. Code and datasets publicly available
Abstract:This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias destabilization: because BiasScore is unsigned, candidate sets mix neurons that push toward and against the stereotype, and the net effect on aggregate bias depends on which sign dominates. The intervention is extremely surgical: zeroing at most 40 neurons in Llama-3.2-1B (less than 0.031% of total MLP width) achieves a mean retention of 99.49% in reasoning and general knowledge capabilities. These findings empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.
[NLP-23] CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLM s in Finance
【速读】: 该论文旨在解决在动态金融环境中部署的4比特量化大语言模型(Large Language Models, LLMs)面临的核心挑战:随着市场状况、监管政策及企业事实的持续变化,如何维持模型输出的事实准确性。现有方法在4比特量化条件下进行连续记忆编辑时,普遍遭遇“量化稳定性危机”——即频繁更新导致性能急剧下降,难以实现稳定的知识维护。其解决方案的关键在于提出CACHE-UK(Contextual Adaptive Continual Hybrid Editor for UK Finance),一个专为领域特定、量化后的LLM设计的稳定性感知记忆编辑框架。该框架的核心创新包括:1)基于秩-1低秩适配器(rank-1 LoRA perturbation)的扰动机制,将知识修改限制在低秩适配子空间内以增强稳定性;2)金融领域优先级模块,根据内容重要性自适应调节编辑强度;3)闭环稳定性控制器(Stability Controller),通过追踪“退化债务”(degradation debt)来预警并防止序列更新中的灾难性遗忘。实验基于4比特量化版OpenLLaMA-3B模型,在包含88,021份文档的英国金融语料库上验证,CACHE-UK在相同量化约束下相较基线方法降低知识退化11%-17%,并实现28%的最高测试成功率(较最强基线提升6个百分点),表明稳定性感知编辑可显著提升资源受限场景下金融LLM的事实保持能力,尽管整体泛化性能仍处于较低水平。
链接: https://arxiv.org/abs/2607.28292
作者: Anubhav Lakra,Yue Feng
机构: Indian Institute of Technology Madras (印度理工学院马德拉斯分校); University of Birmingham (伯明翰大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
备注: 10 pages, 12 figures
Abstract:Large Language Models (LLMs) deployed in dynamic financial environments face a critical challenge: maintaining factual accuracy as market conditions, regulations, and corporate facts change continuously. While 4-bit quantization enables efficient deployment, it severely limits the viability of sequential memory editing: existing methods undergo catastrophic performance degradation under this “quantization stability crisis.” We introduce CACHE-UK (Contextual Adaptive Continual Hybrid Editor for UK Finance), a stability-aware memory editing framework specifically designed for domain-specific, quantized LLMs. CACHE-UK integrates three components: a rank-1 LoRA perturbation mechanism that confines edits to the low-rank adapter subspace, a financial domain prioritization module for content-adaptive edit strength, and a closed-loop Stability Controller that tracks “degradation debt” to prevent catastrophic forgetting across sequential updates. Evaluated on a 4-bit quantized OpenLLaMA-3B model with a curated UK financial corpus of 88,021 documents, CACHE-UK reduces knowledge degradation by 11-17% relative to adapted baselines under identical 4-bit constraints – its most robust effect – while attaining the highest test success (generalization) rate observed in our setting (28%, a 6 percentage point improvement over the strongest adapted baseline). These results indicate that stability-aware editing can improve factual maintenance in resource-constrained financial LLM deployments, though absolute generalization rates remain low.
[NLP-24] (Towards) Scalable Reliable Automated Evaluation with Large Language Models ACL2025
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)生成文本在质量与相关性评估中存在的挑战,尤其是现有自动化评估指标难以捕捉生成内容的复杂性与多样性,且普遍依赖明确的参考标准,限制了其在缺乏客观基准领域的应用。其核心解决方案是提出一种基于多模型对偶比较的新型评估框架,通过多个LLM对生成结果进行两两对比,有效降低单一模型带来的偏差;采用埃洛评分系统(Elo rating system)生成稳定且可解释的排序结果,并引入可调的共识阈值(从完全一致到多数投票),实现对评估置信度与覆盖范围的灵活控制。实验表明,该方法在科学摘要中提取的能力画像评估任务中,自动得出的排名与专家判断高度相关,显著减少了人工干预需求。该框架提供了一种可扩展、一致性高且领域无关的评估范式,为各类应用场景下LLM输出的质量评估提供了高效可靠的支撑。
链接: https://arxiv.org/abs/2607.28282
作者: Bertil Braun,Martin Forell
机构: KIT (Karlsruher Institut für Technologie)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 17 pages. Published in the Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2) at ACL 2025
Abstract:Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in LLM-generated outputs. Moreover, these metrics typically rely on explicit reference standards, limiting their use mostly to domains with objective benchmarks. This work introduces a novel evaluation framework designed to approximate expert-level assessments of LLM-generated content. The proposed method employs pairwise comparisons of outputs by multiple LLMs, reducing biases from individual models. An Elo rating system is used to generate stable and interpretable rankings. Adjustable agreement thresholds, from full unanimity to majority voting, allow flexible control over evaluation confidence and coverage. The method’s effectiveness is demonstrated through evaluating competency profiles extracted from scientific abstracts. Preliminary results show that automatically derived rankings correlate well with expert judgments, significantly reducing the need for extensive human intervention. By offering a scalable, consistent, and domain-agnostic evaluation layer, the framework supports more efficient and reliable quality assessments of LLM outputs across diverse applications.
[NLP-25] MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
【速读】: 该论文旨在解决现代希腊语(Modern Greek)这一形态丰富的语言在自然语言模型评估中长期缺乏专门针对其屈折能力的基准测试问题。当前主流语言模型的评估多集中于事实知识,而忽视了对形态生成与识别等核心语法能力的系统考察。为此,作者提出了MORFES(Morphological Open-class Recognition-and-Formation Evaluation Suite),一个包含500个专家验证题项的基准测试集,专门评估模型在识别和生成希腊语屈折形式方面的能力,并通过优先选取低频词干(lemma)来确保正确答案反映规则推理而非记忆。其关键解决方案在于构建一个以规则驱动、覆盖高阶形态学挑战的评测体系,同时推动开放权重模型生态(从LLaMA到Qwen3、DeepSeek-R1、Magistral及Kimi K2)在多语言覆盖扩展背景下对形态丰富语言的语法能力进行更精准的衡量。实验表明,作者自研并开源的Sophea-Genesis-1模型在屈折形态性能上领先,且在通用能力上与同类规模模型相当,凸显了专为形态复杂语言优化的模型设计的重要性。
链接: https://arxiv.org/abs/2607.28274
作者: Ioakeim Perros,Cleopatra Papadopoulou,Ayoub Kirouane,Christos Petrocheilos
机构: Sophea AI, Kiefer (Sophea AI, Kiefer)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 12 pages, 8 tables
Abstract:Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at this https URL. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at this https URL, leads on inflectional morphology while matching similarly sized models in general capability.
[NLP-26] Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory ACL
【速读】: 该论文旨在解决大模型在处理长上下文时面临的计算开销与记忆效率之间的矛盾问题,即如何在不显著增加推理成本的前提下,有效利用长序列中的关键信息。其核心挑战在于传统方法(如全上下文键值缓存)随上下文长度增长导致内存占用和计算量急剧上升,而现有压缩策略又常伴随语义保真度损失。解决方案的关键是提出一种基于分层结构的新型记忆机制——CoMem(Comprehension Memory),该机制将模型中下层用于语义理解、上层专用于预测的分层分工转化为可显式控制的记忆架构:仅通过中间层对输入上下文块进行编码并写入记忆,检索固定数量的缓存残差状态(residual states),再基于查询条件重新计算上层网络。这一设计实现了模型侧读取计算量与存储上下文长度解耦,且在固定检索预算下具备良好的可扩展性。实验表明,CoMem在多个长上下文任务(如RULER、LoCoMo)上显著优于全上下文直接键值缓存(KV-Direct)方法,同时在对话记忆、独立评估及跨采样场景中均保持优势;此外,在无适配器的高效部署场景下,于NVIDIA H20硬件上实现18.26 GB内存消耗与7.83倍预填充加速,远优于基线的89.36 GB。结果证明,长上下文记忆可沿层轴(layer axis)组织,而非仅依赖传统的令牌轴(token axis),为高效长序列建模提供了新范式。
链接: https://arxiv.org/abs/2607.28263
作者: Hanzuo Liu,Xuan Qi,Chunyu Liu,Haotian Zhong,Yulong Wang,Rayying, Key,Alex Lamb,Mingyu Gao
机构: Tsinghua University (清华大学); Tencent (腾讯)
类目: Computation and Language (cs.CL)
备注: 19 pages, 4 figures, 27 tables. Submitted to ACL Rolling Review
Abstract:Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this division of labor into CoMem (Comprehension Memory), which writes each context chunk only through an intermediate layer, retrieves a fixed number of cached residual states, and recomputes the query-conditioned upper layers over the resulting pack. For a fixed retrieval budget, model-side read compute and memory are independent of stored-context length. We evaluate a continued-trained Qwen3-8B base LM under a unified chat-template-free protocol. The backbone is frozen; the flagship trains only a rank-32 self-distillation LoRA on plain PG19, and we report an adapter-free arm separately. CoMem reaches 97.05 on RULER and 38.27 on LoCoMo versus 34.59 for full-context KV-Direct; the dialogue-memory advantage survives conversation-cluster resampling and an independent judge. Results on additional long-context and long-document tasks expose both the benefits of bounded retrieval and its in-window compression tax. Controlled depth sweeps show that deeper caching lowers per-query recomputation but incurs a fidelity loss that self-distillation substantially repairs. In a separate adapter-free efficiency control on an NVIDIA H20 at 128k, CoMem uses 18.26 GB rather than 89.36 GB and achieves a 7.83x prefill speedup. These results show that long-context memory can be organized along the layer axis, not only the token axis.
[NLP-27] CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
【速读】: 该论文旨在解决预训练语言模型(如BERT)在生成句子嵌入时对语义保持型文本扰动(如同义词替换、掩码和词丢弃)敏感的问题,导致嵌入表示不稳定。其核心解决方案是提出一种轻量级的对比去噪自编码器(Contrastive Denoising Autoencoder, CDAE),通过联合优化对比学习与重建目标,使模型能够学习到对各类文本扰动具有不变性的句子表示。关键在于利用对比学习增强不同扰动版本间语义一致性,同时通过重建任务保留原始语义信息,从而在提升表示稳定性的同时维持语义保真度。实验结果表明,相较于原始BERT和SimCSE,CDAE在多种扰动策略下均能显著提升嵌入相似性,尤其在强扰动条件下优势更明显,验证了扰动不变性学习在改进句子嵌入方面的有效性。
链接: https://arxiv.org/abs/2607.28236
作者: Sina Heydari,Amirreza Abbasi,Mohsen Hooshmand,Majid Ramezani
机构: Institude for Advanced Studies in Basic Sciences (IASBS)(伊朗基础科学高级研究所)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Submitted to 16th International Conference on Computer and Knowledge Engineering (ICCKE 2026)
Abstract:Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout. This work proposes a lightweight Contrastive Denoising Autoencoder (CDAE) that refines pre-trained BERT embedding by jointly optimizing contrastive and reconstruction objective to learn perturbation-invariant representation. We evaluate the proposed framework using multiple perturbation strategies with varying strengths and compare it against the original BERT embeddings and SimCSE. Experimental results show that CDAE consistently preserves higher embedding similarity under perturbations, with the improvements becoming more pronounced as framework effectively enhances representation stability while preserving semantic information, highlighting perturbation-invariant learning as a promising direction for improving sentence embeddings. The source code is publicly available at: this https URL
[NLP-28] Causal Discovery with Inverted Self-attention for Multivariate Time Series
【速读】: 该论文旨在解决多变量时间序列数据中因果发现的挑战,主要源于变量间复杂的交互关系、高维度特性以及非线性依赖关系,现有方法往往难以有效捕捉这些复杂性,导致因果结构推断不准确。其解决方案的关键在于提出一种基于Transformer架构自注意力机制的新型框架,核心创新为引入反向因果自注意力机制(Inverted Causal Self-Attention Mechanism, CSAM),通过反转输入标记(tokens)并诱导注意力分数稀疏化,强化对潜在和间接因果关系的建模能力,同时聚焦显著的因果交互并抑制虚假相关性。此外,该框架还集成了全局因果算法以识别全局因果链接,并设计了因果验证模块以提升所识别因果关系的鲁棒性与可靠性。实验结果表明,该方法在线性和非线性数据集上均优于现有方法,充分展示了其在复杂多变量时间序列因果发现中的有效性与潜力。
链接: https://arxiv.org/abs/2607.28212
作者: Yusen Liu,Yong Wang,Yifan Yin,Tianqing Zhu,Xiufeng Liu,Huan Huo
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Causal discovery in multivariate time series data is challenging due to complex interactions, high dimensionality, and nonlinear dependencies among variables. Existing methods often struggle to capture these complexities, resulting in inaccurate causal structures. To address this issue, we propose a novel framework that leverages self-attention mechanisms within the transformer architecture for causal discovery. Our approach introduces a novel inverted causal self-attention mechanism (CSAM) that emphasizes latent and indirect causal relationships by inverting tokens and inducing sparsity in attention scores, focusing on significant causal interactions and reducing spurious correlations. Additionally, we develop a global causal algorithm to identify global causal links, providing a holistic metric for causal influence, along with a causal verification module to ensure robustness in the identified causal relationships, enhancing the reliability of our framework. Experiments on both linear and nonlinear datasets, along with ablation studies and sensitivity analyses, show that our framework outperforms existing methods, demonstrating its potential for causal discovery in complex multivariate time series.
[NLP-29] Fidelity Is Not Safety: Gently-Compressed LLM s Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agent ic Execution
【速读】: 该论文旨在解决生成式模型在轻量化压缩(如低秩近似)后可能出现的“代理安全”隐患问题,即压缩后的模型在执行标准操作流程(SOP)时会无中生有地生成从未出现在原始指令中的虚构步骤。尽管当前主流评估体系通过困惑度(perplexity)、下游任务准确率(如MMLU)以及基于随机探针输入的输出保真度信号等“数据廉价”的质量守卫机制来验证压缩模型的有效性,但这些指标存在盲区——它们无法检测到由低秩压缩引起的、与特定压缩算子相关的“幻觉行为”。研究发现,这种幻觉行为具有算子特异性:仅在采用协同低秩(coherent low-rank, SVD)截断时出现,而同等困惑度水平下的幅度剪枝(magnitude pruning)则不会引发此现象。关键突破在于识别出导致该问题的核心机制——压缩误差的相干性与其变化速率的乘积(coherence × rate),而非误差大小本身。传统的数据自由保真度探针因构造上为保真度“预言家”,无法捕捉这一维度,从而形成评估盲点。作者通过在三个架构上预注册并充分幂化的对照实验,揭示了该现象的可重复性,并提出一种双轴数据自由筛查指标:压缩误差的相干分数(coherent fraction)与误差率(error rate),其固定阈值可有效识别高风险压缩版本。研究表明,仅满足困惑度、MMLU和保真度标准并不能保证模型在代理场景下的安全性,因此必须在部署前对轻量级低秩压缩模型进行该双轴筛查。
链接: https://arxiv.org/abs/2607.28196
作者: I. Kennedy,T.Kennedy
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the compressed and original network’s internal representations under random probe inputs. This stack has a blind spot. Across three model families, gently-compressed models clear every guard and then invent procedure steps that were never in the instructions when they run a standard operating procedure (SOP) as an agent. The effect is operator-specific: coherent low-rank (SVD) truncation induces it, and magnitude pruning matched to the same perplexity does not. One dissociation isolates the cause. The same compressed weights that CI-win a paired output-fidelity test CI-fail the invented-step canary. The governing axis is the coherence of the compression error times its rate; the magnitude of the damage does not predict it. The data-free fidelity probe is a fidelity oracle by construction, so it cannot see this axis. We characterize the blindspot and dissociation with paired confidence intervals on a pre-registered, powered canary across three architectures. Operator-specificity replicates on all three, and the perplexity-guard evasion appears where the model admits in-guard low-rank headroom. We then give a data-free screen: a two-axis statistic of the compression error (coherent-fraction and error-rate) that flags the failing builds with fixed thresholds across architectures and matches the coherence-times-rate mechanism. Perplexity, MMLU, and fidelity acceptance do not certify agent safety. Screen gently-compressed low-rank builds before agentic deployment
[NLP-30] he MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
【速读】: 该论文旨在解决在临床试验场景下,如何利用自动化方法辅助精神科医生更高效、准确地评估抑郁症患者症状严重程度的问题。传统诊断依赖于临床评估,而现有基于生成式AI(Generative AI)的方法多集中于非结构化文本(如社交媒体内容),难以适配临床试验中以标准化访谈(如SIGMA指南)为基础的结构化评估流程。本文提出了一种专为临床试验设计的大语言模型(Large Language Model, LLM)处理流程,其关键在于将音频访谈自动转录为文本,并将其映射至蒙特利尔抑郁量表(MADRS)的10项核心症状维度,进而量化各症状的严重程度,并识别出存在潜在偏差的临床评分。实验结果表明,该方法与专家评分间具有0.867的强相关性,具备高度可解释性,可为未来临床试验中的抑郁症评估提供可靠支持。
链接: https://arxiv.org/abs/2607.28190
作者: Mila Fodor,Katalin Ócsai,Francesco Periti,Rien Sonck,Alex Boudreau
机构: Clario, part of Thermo Fisher Scientific
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing solutions primarily focus on detecting the disorder from different text sources (e.g., online text, social media), there is still limited support for clinical trials, where clinical assessments are conducted through structured interviews based on standard guidelines such as SIGMA. In this work, we develop a LLM pipeline specifically designed to support clinicians in supporting the assessment of depression in patients enrolled in clinical trials. Our pipeline converts audio interviews into transcripts, maps them into the ten MADRS symptom items, estimates their severity, and identify problematic clinical ratings associated with them. Evaluation on real clinical interviews shows a strong overall correlation of 0.867 with expert ratings, providing interpretable support for future assessments in clinical trials.
[NLP-31] Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models ATC
【速读】: 该论文旨在解决生成式语言模型在解码过程中效率低下与过早终止之间的矛盾问题,尤其针对扩散语言模型(Diffusion Language Models, DLMs)在生成时因缺乏对输出稳定性的精细判断而导致的提前退出(early exit)失效问题。现有方法依赖于固定区域置信度统计或与解码进度相关的粗粒度规则进行终止决策,无法适应长链式推理(chain-of-thought)中答案仅在末尾阶段才趋于稳定的特性,导致过早终止并损害生成质量。为此,本文提出一种无需训练、面向候选内容的早期退出框架——LATCH(Localized Acceleration with Tracked-Candidate Halting),其核心创新在于将“何时退出”(termination timing)与“何处加速”(acceleration location)两个维度解耦:通过置信度验证承诺(Confidence-Verified Commit, CVC)动态提取候选片段,并基于确定性解析器验证其置信度与最大值稳定性,实现对全局输出稳定性的精准判断;同时采用块级早期承诺(Block-Wise Early Commit, BWEC)在非终块上应用轻量局部规则以加速解码,而保留最终块与全局终止控制权于CVC。该方法无需后缀提示(suffix-prompt)构造,完全不依赖提示锚点(prompt-anchor-free),但具备任务格式感知能力。在11个零样本任务上的实验表明,LATCH在22组评估设置中保持与全解码相当的精度(误差不超过2.0个百分点),且仅需一个冻结超参数配置即可跨骨干模型迁移使用,同时在短答案任务上实现9.3–17.8倍的端到端每秒令牌数(TPS)加速,在长推理任务上也实现2.0–3.3倍加速,显著提升了生成效率与准确性平衡。
链接: https://arxiv.org/abs/2607.28166
作者: Chia-Ming Lee,Ming-Ching Chang,Xin Li,Yu-Lun Liu,Chih-Chung Hsu
机构: National Yang Ming Chiao Tung University (国立阳明交通大学); University at Albany, SUNY (纽约州立大学阿尔巴尼分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Code is available at this https URL
Abstract:Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide termination from fixed-region confidence statistics or schedule-dependent rules, evidence too coarse for a decision that freezes every remaining position at once, so they fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end. Adaptive sampling, the other axis of training-free acceleration, paces how quickly positions commit while decoding continues but never verifies that the output itself has stabilized. We introduce a training-free, candidate-aware early-exit framework that keeps the two axes separate and matches each decision to evidence of its own scope. Confidence-Verified Commit (CVC) governs when the sequence may stop by verifying confidence and sustained argmax stability over the dynamically extracted candidate span using a deterministic parser specified from each task’s output format. Block-Wise Early Commit (BWEC) governs where to accelerate by applying a cheaper local rule to non-final blocks, while leaving the final block and global termination under CVC. We refer to their combination as LATCH (Localized Acceleration with Tracked-Candidate Halting). Unlike prior methods, LATCH needs no suffix-prompt construction; it is prompt-anchor-free but format-aware. We evaluate LATCH end to end on 11 tasks under zero-shot settings using LLaDA and Dream. LATCH stays within 2.0 percentage points of full-decoding accuracy across all 22 evaluation settings, with one frozen hyperparameter set that transfers cross-backbone untuned, while achieving end-to-end TPS speedups of 9.3-17.8x on short-answer tasks and 2.0-3.3x on long-reasoning tasks.
[NLP-32] RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
【速读】: 该论文旨在解决多模态长时程记忆代理在处理长视频任务时,因检索机制不完善而导致的推理失败问题。现有方法普遍关注“存储什么”而非“如何有效检索”,导致在检索失败或获取无效证据时缺乏对历史任务轨迹的故障诊断能力,无法自适应调整后续搜索策略。其解决方案的关键在于提出一种名为反思式检索记忆(Reflective Retrieval Memory, RRM)的框架,通过在以实体为中心的多模态记忆图基础上引入反思经验记忆(reflective experience memory),从历史任务轨迹中提炼可迁移的、与检索过程相关的程序性知识。该反思经验记忆不保存当前视频的事实性证据,而是捕获跨任务复用的搜索策略,并将这些经验转化为查询层面的指导信息;而答案生成仍仅依赖于从当前视频中新检索到的事实性证据,确保推理的准确性。此外,通过使用频率、重用反馈和时间衰减机制进行生命周期管理,有效降低了记忆冗余与噪声。实验表明,RRM在M3-Bench-Robot、M3-Bench-Web和Video-MME-Long等多个基准上均显著优于现有最先进方法,验证了其在长时程多模态推理中的有效性。
链接: https://arxiv.org/abs/2607.28156
作者: Jingxiang Fan,Junbao Zhuo,Bochao Zou
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search this http URL introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.
[NLP-33] Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
【速读】: 该论文旨在解决在高风险场景(如医疗与法律系统)中部署大语言模型(LLM)时,对其欺骗能力的可解释性与安全性评估问题。随着生成式AI(Generative AI)作为智能代理参与复杂社会博弈,其在信息不对称条件下进行伪装、说服与策略推理的能力成为关键安全挑战。为此,研究提出开源基准框架ParliamentBench,基于“秘密希特勒”(Secret Hitler)这一控制型社交推断游戏,构建可复现的评测环境,以分离并量化模型在欺骗、推理与角色一致性方面的表现。解决方案的关键在于设计三个新颖的量化指标——社交推断能力、推理质量与欺骗一致性,从而实现对模型在多轮互动中持续维持虚假身份能力的精细化评估。实验结果表明,前沿模型在合作与欺骗角色中均表现出较强性能,形成以GPT-5.4、Kimi K2.5、Grok 4.1 Fast和DeepSeek 3.1 Terminus为核心的前四强集群;而多数模型在整局游戏中难以保持稳定的欺骗行为,欺骗保留率普遍低于50%,显著低于随机基线(33%)与简单算法基线(45%),揭示当前主流模型在长期策略伪装方面仍存在严重缺陷。
链接: https://arxiv.org/abs/2607.28146
作者: Niklas Bauer,Lars Benedikt Kaesberg,Akiko Aizawa,Jan Philip Wahle,Bela Gipp,Terry Ruas
机构: University of Göttingen (哥廷根大学); National Institute of Informatics (日本国立情报学研究所); University of Tokyo (东京大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.
[NLP-34] Rethinking LLM -Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)辅导场景中的一项核心测量难题:通用的“助益性”评估标准是否能够有效区分直接给出答案的行为与具有教学引导意义的辅导行为。研究通过一项预注册的审计实验,在三个不同的导师基线中,对比相同基础模型实现的对话式策略与教学导向策略,且均针对同一固定能力较弱的模拟学生进行测试。研究采用确定性检测器量化答案泄露程度与后续回合中学生独立完成任务的程度,并以冻结的Claude Opus 4.8作为无条件判别基准。在Opus评分固定后,进一步使用GPT-5.6 Sol对1,179个确认阶段的导师回应进行前瞻性稳健性审计。结果显示,在主基线中,两种策略在助益性评分上无显著差异,但在教学性评分上呈现完全分离(Cliff’s δ = 0.10 vs. 1.0),表明助益性评分不具备一致性,其排序结果随评判者不同而反转,而教学性对比则保持方向一致。在仅使用Opus的消融分析中,七种主基线策略在助益性评分上仅相差0.25分范围内,但教学性评分跨度达2.3分。此外,所有基线下,暴露答案的回应均导致学生后续独立工作减少,这一发现具有评判者无关性。因此,研究结论为:通用助益性评分无法可靠反映教学过程质量,应将专门针对教学性的评估量表与确定性的过程指标相结合,以实现对辅导行为的有效评价。
链接: https://arxiv.org/abs/2607.28128
作者: Shuyi Fan,Boyuan Deng,Mengyu Xu,Jiale Liu,Hongyang Zhang,Qiaoxin Yang,Chongyang Gao
机构: Google(谷歌); Stanford University (斯坦福大学); Tsinghua University (清华大学); Peking University (北京大学); University of Science and Technology of China (中国科学技术大学); Fudan University (复旦大学); Shanghai Jiao Tong University (上海交通大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: 24 pages, 4 figures, 6 tables
Abstract:LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff’s |\delta|=0.10 vs. 1.0 ). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span 2.3 points in mean judged pedagogy within a 0.25 -point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.
[NLP-35] FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning
【速读】: 该论文旨在解决现有金融情感分析方法在动态市场环境中适应性不足的问题。当前主流方法依赖于静态、人工标注的监督学习数据集,难以随市场条件变化进行有效调整,导致模型泛化能力受限。其解决方案的关键在于提出首个面向市场的强化学习框架FinSMART,通过直接以实际市场结果(如交易收益)作为反馈信号,实现对金融情感信号的端到端优化。该框架创新性地融合了市场感知的数据筛选机制与离散非对称交易奖励函数,有效应对金融市场噪声大、非平稳性强及多因子影响的挑战,从而在复杂环境下实现稳定可靠的强化学习。此外,FinSMART支持基于新生成的财经文本及其对应市场表现的实时再训练,无需昂贵的人工标注,使模型能够持续适应市场演化,显著提升盈利能力和风险调整后绩效,相较最强基线模型累计交易回报提升220%。这一成果验证了市场对齐强化学习在构建自适应金融大语言模型中的巨大潜力,标志着向下一代智能化金融分析范式的重要迈进。
链接: https://arxiv.org/abs/2607.28127
作者: Giorgos Iacovides,Wuyang Zhou,Danilo Mandic
机构: Imperial College London(帝国理工学院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Statistical Finance (q-fin.ST); Trading and Market Microstructure (q-fin.TR)
备注:
Abstract:Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs). However, existing approaches remain confined to a market-agnostic, supervised learning paradigm that relies on limited, static and human-annotated datasets, and thus are incapable of adapting to evolving market conditions. To address this limitation, we introduce FinSMART, the first market-aligned reinforcement learning framework for financial sentiment analysis, which directly optimizes sentiment signals using realized market outcomes. To deal with the noisy, non-stationary, and multifactorial nature of financial markets, FinSMART incorporates a signal extraction pipeline that combines market-aware data filtering with a discrete asymmetric trading reward, enabling stable reinforcement learning from economically meaningful market feedback. Experimental results demonstrate that FinSMART significantly outperforms existing state-of-the-art methods in profitability, risk-adjusted performance, and sentiment signal quality, improving cumulative trading returns by 220% over the strongest baseline. Uniquely, the FinSMART framework naturally supports market-aware retraining, at any point in time, by replacing costly manual annotation with newly observed financial articles and their realized market outcomes. Such a retraining strategy enables the model to continuously adapt to changing market dynamics, resulting in consistent performance gains over its static counterpart. These findings demonstrate the practical applicability of market-aligned reinforcement learning and highlight its potential as a next-generation paradigm for developing adaptive financial LLMs.
[NLP-36] Challenges in annotations by humans and LLM s: A case study of evaluative language
【速读】: 该论文旨在解决复杂语言现象在语料标注任务中的一致性与可靠性问题,特别是针对具有高度主观性的评价性语言(evaluative language)在科学传播话语中的标注挑战。研究聚焦于评价理论(Appraisal theory)中的态度子系统(Attitude subsystem),涵盖情感(Affect)、判断(Judgement)和欣赏(Appreciation)三类范畴,此类标注任务因其主观性强而成为典型复杂的标注难题。研究通过对比训练中的语言学家、已训练的语言学家以及大型语言模型(LLM)生成的标注结果,发现训练中的语言学家难以达成高一致性评分,而经过提示工程优化并微调后的LLM在自动分类任务中表现最佳,达到0.77的F1分数,甚至优于专业语言学家的标注表现。其解决方案的关键在于:通过精心设计的提示(prompt)策略与模型微调,使LLM能够有效捕捉评价性语言的细微语义特征,从而在高度主观的标注任务中实现可信赖的自动化处理,为数字人文研究中复杂理论的标注与分析提供了新的可行路径。
链接: https://arxiv.org/abs/2607.28119
作者: Mirela Imamovic,Aenne Cecilia Kristine Knierim,Khushi Pitroda,Ekaterina Lapshinova-Koltunski
机构: 未知
类目: Computation and Language (cs.CL); Social and Information Networks (cs.SI)
备注:
Abstract:In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
[NLP-37] PCAP-LM: An LLM -Native Text Representation for TLS Bulk Traffic Analysis
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在进行网络流量分析时,因标准捕获格式(如PCAP)及其文本表示形式过于冗长,导致其上下文窗口(context window)被远远超出的问题。现有方法无法有效适配LLMs的输入限制,严重影响了推理效率与准确性。其解决方案的关键在于提出一种面向流(flow-centric)、专为LLM设计的新型文本表示方法——PCAP-LM,该方法通过“有损知识提取”而非传统压缩方式实现高效降维:利用文中提出的PacketGlyphs(一种创新的ASCII字符集),对数据包的方向、TCP/TLS状态、对数尺度大小及数据包间延迟等关键语义特征进行编码;结合受限的PMI-BPE分词器与模式游程编码(motif run-length encoding),大幅压缩重复行为模式;同时引入@REFS侧索引以保留原始数据的无损回溯能力。实验表明,在5G/4G TLS 1.3批量下载流量数据集上,该方法将原始tshark -V输出压缩达812倍,词汇表仅需159个词元即可饱和,且可完整容纳于单个LLM上下文窗口内。在30个独立测试文件的取证问答任务中,使用前沿LLM时,PCAP-LM的准确率高达99.3%,显著优于同词元预算下tshark -V前缀的51.0%。尽管该有损设计引入了已知盲区(如TCP重传的24%假阴性率),且在异构混合协议环境中的泛化需重新训练词表,但整体方案为高效、可扩展的网络流量语义理解提供了新范式。
链接: https://arxiv.org/abs/2607.28100
作者: Xavier Marjou,Lucas Tamic,Ilan Jaffeux-Cheniout
机构: 未知
类目: Networking and Internet Architecture (cs.NI); Computation and Language (cs.CL)
备注: 6 pages
Abstract:Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equivalents are prohibitively verbose, overflowing LLM context windows by two orders of magnitude. We present PCAP-LM, a flow-centric, LLM-native text representation that acts as a lossy knowledge extraction step rather than a standard compression tool: raw captures are transcoded into semantic summaries using PacketGlyphs - a novel ASCII alphabet coined in this paper that encodes packet direction, TCP/TLS state, log-scale size, and inter-packet delay. Combined with a constrained PMI-BPE tokenizer and motif run-length encoding, repetitive behavioural patterns are aggressively collapsed. A @REFS side-index preserves lossless drill-down into the original packets. Evaluated on a homogeneous corpus of 5G/4G TLS 1.3 bulk-download traffic, the BPE vocabulary fully saturates at 159 tokens, achieving an 812x size reduction over tshark -V and fitting entire captures within a single LLM context window. In a forensic question-answering evaluation over 30 held-out files, a frontier LLM achieves 99.3% accuracy from PCAP-LM documents versus 51.0% from a token-budget-matched tshark -V prefix. The lossy design introduces known blind spots - most notably a 24% false-negative rate for TCP retransmissions - and extending to heterogeneous mixed-protocol environments will require vocabulary retraining.
[NLP-38] SciDataSailor: Deep Scientific Data Exploring
【速读】: 该论文旨在解决科学数据集(scientific datasets)在实际科研应用中因结构复杂、文件异构且相互依赖,导致数据检查、整合与分析过程高度依赖领域知识、耗时费力的问题。现有大语言模型(LLM)虽在规划、推理和工具调用方面取得显著进展,但普遍缺乏与真实科学数据资产通过可执行环境进行交互的能力。为此,论文提出“深度科学数据探索”(Deep Scientific Data Exploration)这一智能体任务范式,其核心在于使智能体能够自主导航数据仓库、解析异构文件与模式、执行数据分析、跨文件证据整合,并基于实际执行结果生成结论。该方案的关键在于构建名为SciDataSailor的框架,通过平衡广度探索与目标导向利用来合成工具交互轨迹;其具体实现采用蒙特卡洛树搜索(Monte Carlo Tree Search, MCTS),并集成四项任务特化机制:按难度分层的探索种子、双反馈优先级引导的首次尝试策略、层次化从策略到工具动作的生成机制,以及基于熵引导的分支选择策略。基于此框架,研究构建了用于监督微调的SciDataSailor-SFT-2K数据集及用于评估的SciDataSailor-Bench基准,后者涵盖27个跨生命科学、地球科学与物理学领域的数据集,包含627项元信息摘要任务与586项科学问答任务,有效验证了该范式在真实科学数据场景中的可行性与有效性。
链接: https://arxiv.org/abs/2607.28098
作者: Jiyong Rao,Yicheng Qiu,Chi Zhang,Chunfeng Song,Runkai Zhao
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 63 pages, 10 figures
Abstract:Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments. We introduce Deep Scientific Data Exploration, an agentic task paradigm in which agents navigate repositories, interpret heterogeneous files and schemas, execute analyses, integrate cross-file evidence, and produce conclusions grounded in executed observations. To operationalize this paradigm, we present SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation. SciDataSailor instantiates trajectory synthesis as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms: difficulty-stratified exploration seeds, dual-feedback first-play urgency, hierarchical strategy-to-tool action generation, and entropy-guided branching. Using this framework, we construct SciDataSailor-SFT-2K for supervised fine-tuning and SciDataSailor-Bench for evaluation, with the latter comprising 627 meta-information summarization tasks and 586 scientific question-answering tasks across 27 datasets spanning the life, earth, and physical sciences.
[NLP-39] GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在文本转SPARQL(Text-to-SPARQL)任务中生成查询时存在的语义不一致问题:尽管生成的查询在语法上可执行,但其语义可能与原始自然语言问题不符,导致知识图谱检索结果错误。为提升生成查询的准确性和可靠性,论文提出一种名为生成器-门控器-修正器(Generator-Gate-Corrector, GGC)的框架。其核心解决方案在于引入选择性修正机制——首先由生成器(Generator)生成初始查询,随后通过门控器(Gate)判断该查询是否需要修正,仅对高风险查询调用修正器(Corrector)进行干预。该机制有效避免了对本已正确的查询进行不必要的修改,从而降低了因错误修正导致性能下降的风险。实验结果表明,相较于对所有生成查询统一修正的方法,GGC在MCQA数据集上将查询级准确率从90.23%提升至98.33%,同时推理开销降低45%。消融实验进一步验证了门控器在不同阈值下的鲁棒性,以及修正器训练数据构成对修正效果与稳定性的显著影响。总体而言,选择性修正策略显著提升了基于大语言模型的文本转SPARQL生成在准确性、可靠性和效率方面的综合表现。
链接: https://arxiv.org/abs/2607.28082
作者: Ziyi Yang,Thanh-Son Nguyen,Tuan Anh Nguyen,Lihui Chen
机构: Nanyang Technological University (南洋理工大学); Institute of High Performance Computing, Agency for Science, Technology and Research (A*STAR) (高性能计算研究所,科技研究局)
类目: Computation and Language (cs.CL)
备注: 18 pages, 1 figure
Abstract:Large language models (LLMs) have demonstrated strong capabilities in structured query generation, making them a natural choice for Text-to-SPARQL, which translates natural language questions into executable SPARQL queries over knowledge graphs. However, their initial outputs remain unreliable: generated queries may be executable yet semantically misaligned with input questions, leading to incorrect retrieval. To address this issue, we propose Generator-Gate-Corrector (GGC), a framework for reliable LLM-based Text-to-SPARQL generation. GGC first uses a Generator to produce an initial query, then applies a Gate to predict whether correction is needed, and finally invokes a Corrector only for selected high-risk queries. This selective correction mechanism avoids unnecessary modifications and reduces the risk of degrading originally correct queries. Experiments on MCQA show that GGC improves query-level accuracy from 90.23% to 98.33% while reducing inference overhead by 45% compared with correcting all generated queries. Ablation studies show that the Gate is robust across thresholds and that Corrector training data composition affects correction effectiveness and stability. Overall, the results demonstrate that selective correction enhances the accuracy, reliability, and efficiency of LLM-based text-to-SPARQL generation.
[NLP-40] LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
【速读】: 该论文旨在解决生成式强化学习中因大量具有相同回滚奖励(rollout reward)的提示组(prompt group)导致的生成预算浪费问题,其核心挑战在于如何在利用历史有效提示(exploitation)与探索未知不确定性提示(exploration)之间实现平衡。现有预回滚提示选择方法往往陷入两难:过度利用历史高效提示会降低训练覆盖范围,而广泛探索又会导致高比例无效提示,降低学习效率。为此,本文提出一种基于潜在空间引导的探索-利用提示采样器(Latent-Guided Explore–Exploit Prompt Sampler, LEEPS),其关键创新在于通过动态划分候选提示为“利用”与“探索”两个组合集,并依据近期非平凡奖励比率(non-trivial ratio)自适应分配回滚预算;同时,结合表示空间中的邻居关系与历史回滚结果,优先选择可能产生非零奖励方差的不确定提示,从而实现更精准的目标化探索,无需额外回滚开销。实验表明,在六项数学推理基准上,LEEPS在两种模型规模下均取得最高平均得分,相较于最强基线分别提升2.6%和3.7%,且训练收敛速度更快;在三项分布外(OOD)通用推理任务中也表现最优,同时每轮训练仅增加约2秒在线采样延迟,展现出优异的性能与效率平衡。
链接: https://arxiv.org/abs/2607.28077
作者: Shuang Liang,Haoyang Zhou,Yifan Gong,Guowei Wang,Xiting Wang
机构: University of Science and Technology of China (中国科学技术大学); Tsinghua University (清华大学)
类目: Computation and Language (cs.CL)
备注: 15pages
Abstract:Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore–Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6% and 3.7% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at this https URL.
[NLP-41] RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中能力方向(capability directions)的表征工程(representation engineering)评估缺乏可比性和可复现性的问题。现有方法多依赖于特定论文生成的合成数据,导致测量结果难以横向比较,且可能反映的是表面模式而非真实能力。为应对这一挑战,论文提出RepBench,一个基于基准测试(benchmark)的数据层,用于对齐能力的表征探测。其关键创新在于构建了一个涵盖182个能力簇、分属13个家族的能力分类体系,并整合了46,149条经审计的探针文本,覆盖94项能力,每项能力均得到至少两个独立基准的支持。通过多基准设计,有效降低了对单一数据源的依赖;实验表明,原始文本向量不具备自然聚类结构,而基于基准聚合的能力向量在所有12个评测模型上均在少量聚类数时达到内部聚类最优,但与人类分类体系一致性较低。跨基准迁移评估显示,不同读出方法(readout)和聚合准则具有显著差异:在十种模型上,均值差异法(difference-in-means)表现最佳;而在最多的能力-模型组合中,逻辑回归表现更优。这一不一致性揭示了读出方法与聚合标准作为评估维度的重要性。整个流程、语料库及评估代码均已开源,形成可重复使用的闭环工作流。
链接: https://arxiv.org/abs/2607.28008
作者: Yanshi Li,Xueru Bai,Shuman Liu,Long Zhang
机构: Google(谷歌); Stanford University (斯坦福大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 22 pages, 8 figures, with appendices. Yanshi Li and Xueru Bai contributed equally
Abstract:Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and may reflect surface patterns rather than capabilities. We present RepBench, a benchmark-grounded data layer for capability-aligned representation probing. Crawling 13,427 benchmark papers yields a taxonomy of 182 capability clusters in 13 families; harvesting 353 public benchmark datasets yields 46,149 audited probe texts covering 94 capabilities, each supported by at least two independent benchmarks. This multi-benchmark design reduces dependence on any single source: raw per-text vectors exhibit no natural cluster granularity, whereas benchmark-pooled capability vectors show an interior clustering optimum at a small number of clusters on all 12 evaluated models, with low agreement to the human taxonomy. Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells. This disagreement shows that the readout method and aggregation criterion are meaningful evaluation dimensions. The pipeline, corpus, and evaluation code are released as a reusable closed-loop workflow.
[NLP-42] riShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement
【速读】: 该论文旨在解决联邦微调(Federated Fine-Tuning)中由恶意参数服务器发起的神经印记攻击(NeuroImprint)所引发的隐私泄露问题。此类攻击通过在参数服务器端植入专用的记忆化神经元(memorization neuron),利用每个神经元仅更新一次的特性,实现对客户端训练数据的高语义保真度重建(重建率高达59%–79%)。现有防御手段如局部差分隐私(Local Differential Privacy, LDP)和梯度裁剪等,或无法有效抵御该攻击,或导致模型性能显著下降。本文提出一种三重确定性防御框架——TriShield,其核心解决方案包括:(1)参数伪影检测器(Parameter Artifact Detector),在本地训练前识别分布式模型参数中的记忆化神经元特征;(2)有状态虚拟迭代机制(Stateful Virtual Iteration),通过强制Adam/AdamW优化器的动量状态在虚拟迭代步间不可逆纠缠,破坏神经印记攻击依赖的闭式反演条件;(3)零效用正交投影算子(Zero-Utility Orthogonal Projection),基于奇异值分解(SVD)将所有本地梯度更新投影至主任务语义子空间,物理上消除携带私有记忆信息的梯度分量。理论证明表明,在应用第2、3层后,上传梯度与任一训练样本间的互信息为零。实验在GPT-2(117M)和Llama-Guard-3-1B上验证,TriShield可将神经印记攻击的重建率降至0%,同时保持或提升训练精度,且额外计算开销低于5%。
链接: https://arxiv.org/abs/2607.27940
作者: Cheng Wei(Honor Device Co., Ltd., Shenzhen, China)
机构: Honor Device Co., Ltd.(荣耀终端有限公司)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 12 pages,3 figures
Abstract:Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.20553), demonstrates that a malicious parameter server can corrupt a PEFT adapter into a privacy backdoor: by assigning a dedicated memorization neuron to each training sample and ensuring each neuron updates at most once, the server can analytically reconstruct 59%–79% of client training data with high semantic fidelity. Existing defenses—including local differential privacy (LDP) [8] and gradient clipping—either fail against this attack or impose unacceptable utility degradation. We present \textbfTriShield, a three-layer deterministic defense that completely prevents NeuroImprint-style reconstruction with \textbfzero model utility loss and \textbfno additional communication rounds. TriShield consists of: (1) a \textbfParameter Artifact Detector that identifies memory-neuron signatures in distributed model parameters before local training begins; (2) a \textbfStateful Virtual Iteration mechanism that forces Adam/AdamW’s momentum state to irreversibly entangle gradients across virtual steps, invalidating NeuroImprint’s closed-form inversion; and (3) a \textbfZero-Utility Orthogonal Projection operator that projects all local gradient updates onto the main-task semantic subspace computed via SVD, physically eliminating any gradient components that carry private memorization. We prove theoretically that after Layers 2 and 3, the mutual information between the uploaded gradient and any individual training sample is zero. Experiments on GPT-2 (117M) and Llama-Guard-3-1B verify that TriShield reduces NeuroImprint reconstruction rate to \textbf0% across all tested attack variants, while maintaining or improving training accuracy, with less than 5% additional GPU computation overhead.
[NLP-43] Memory Decoder at Scale: A Pretrained Parametric Long-Term Memory
【速读】: 该论文旨在解决生成式 AI(Generative AI)中解码器仅语言模型(decoder-only language models)因将长期记忆与推理功能耦合于同一参数集而导致难以独立扩展记忆容量的问题。其核心挑战在于:当模型规模扩大时,传统基于Faiss的近似最近邻(kNN)索引与检索机制在大规模数据(如300B token)下面临计算与存储瓶颈,导致系统不可行。为此,本文提出“内存解码器规模化”(Memory Decoder at Scale),通过构建分布式Faiss索引与检索流水线,并结合稀疏、批处理式加载kNN分布,有效缓解了这一性能瓶颈。关键解决方案在于引入可独立扩展的参数化长期记忆模块,实验表明,在不同模型尺度下,将更多参数分配给记忆模块相比单纯扩大基础模型,能实现更优的参数-性能权衡。在17个基准测试中,一个6.9B通用记忆模块与Pythia-410M结合后,平均得分从29.86提升至37.34,超越参数量高达39%更大的Pythia-12B(37.24)。对于Qwen3 Base系列(0.6B–14B),配备1.7B领域记忆模块后,各尺度下跨三个领域的平均得分均提升超9分。结果表明,独立扩展预训练记忆模块为提升语言模型性能提供了一条更具参数效率的路径。
链接: https://arxiv.org/abs/2607.27919
作者: Rubin Wei,Jiaqi Cao,Jiarui Wang,Junming Zhang,Qipeng Guo,Bowen Zhou,Zhouhan Lin
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
[NLP-44] IFHierBench: Hierarchical Instruction Following for Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在实际应用中对复杂层级化指令遵循能力不足的问题,尤其关注多层嵌套结构输出中各子部分需满足特定约束条件的场景。现有指令遵循基准测试将约束视为扁平列表并统一应用于整个输出,无法针对输出中的特定结构层级进行细粒度验证,因而难以评估模型在复杂层次化约束下的表现。为此,论文提出IFHierBench,一个包含600个分层提示的层次化指令遵循基准,覆盖四种不同的约束树深度和35种不同约束类型,并为每个提示配备确定性检查器,可逐层验证输出在各个作用域内的约束满足情况。实验评估七种领先的专有及开源模型发现,即使最强模型的提示级准确率也仅略高于50%,且随着约束层级深度增加,性能显著下降。这表明当前大语言模型在可靠遵循嵌套约束方面仍存在显著差距,亟需未来训练方法在更细粒度层面关注约束遵循能力,以提升整体指令遵循性能。
链接: https://arxiv.org/abs/2607.27912
作者: Yuetian Mao,Chunyang Chen
机构: Technical University of Munich (慕尼黑工业大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the full task in a single LLM call, with one prompt specifying a layered output whose overall artifact, structural sections, and nested fields must each satisfy concrete constraints. Existing instruction-following benchmarks treat the constraint set as a flat list applied uniformly to the response, so they cannot scope a check to a particular section of the output. We introduce IFHierBench, a hierarchical instruction-following benchmark of 600 prompts stratified across four constraint-tree depths and 35 distinct constraints, each prompt paired with a deterministic checker that verifies satisfaction at every scope. Evaluating seven leading proprietary and open-weight models, we find that even the strongest model only marginally exceeds 50% prompt-level accuracy and that accuracy degrades sharply as constraint depth grows. Reliably following nested constraints remains a substantial gap for current LLMs, motivating future training methods that consider constraint adherence at finer granularity to achieve better instruction-following ability.
[NLP-45] FinanceHarness: Autonomous Financial Deep Research Framework
【速读】: 该论文旨在解决金融深度研究(financial deep research)自动化中的核心挑战,即现有基于大语言模型(LLM)的代理系统普遍生成通用报告,难以满足金融领域对专业性、时序准确性及信息隔离的严格要求。其关键问题在于:如何构建一个既能驱动研究代理进行分层决策、又能确保在特定时间点(point-in-time)评估中不泄露未来信息的可验证基准体系。为此,论文提出FinanceHarness,一个集成金融导向工具与从业者指导工作流的框架,实现从环境与数据构建、代理执行循环到奖励建模的端到端自动化。同时,提出FinanceGym,一套以论点驱动的研究问题集和包含事前(pre-cutoff)与事后(post-cutoff)双重标准的评分体系,通过专业专家验证,达到82%的通过率。实验表明,即使顶尖大模型在该基准上得分也低于40%,凸显其挑战性;而采用相同开源底座的FinanceHarness将综合评分从25.3%提升至32.4%,验证了其有效性。
链接: https://arxiv.org/abs/2607.27853
作者: Yijia Xiao,Rujun Han,Yanfei Chen,Zifeng Wang,Ke Jiang,Zhongying CuiZhu,Vishy Tirumalashetty,Wei Wang,Burak Gokturk,Tomas Pfister,Chen-Yu Lee
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Finance (q-fin.CP)
备注:
Abstract:Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at this https URL.
[NLP-46] AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
【速读】: 该论文旨在解决生成式 AI 在科学写作与同行评审辅助中一个关键却未被充分探索的问题:如何验证审稿人反馈是否真正促成稿件的实质性、有证据支持的改进。其核心挑战在于实现对审稿意见回应的可信度评估,即判断作者在修改稿中所声称的改进是否确实基于可验证的实证依据。解决方案的关键在于提出并构建“AutoSupervision”框架,该框架利用公开透明的同行评审记录作为自然监督信号,将审稿意见(reviewer comments)、作者回复(author responses)和修订稿件(revised manuscripts)三者关联起来,形成一个端到端的证据链评估体系。具体而言,模型需完成三个任务:准确表征审稿人提出的科学关切、判断这些关切是否得到实质回应,并定位修订稿中支持改进结论的文本证据。研究基于56,000篇《自然·通讯》(Nature Communications)文章及其对应的评审记录构建数据集,通过大规模实验、消融分析与案例研究发现,尽管大语言模型(LLMs)在理解审稿意见方面表现良好(如GPT-5.5得分为0.754),但基于证据的验证能力仍是主要瓶颈,最优模型仅达到0.501的得分,凸显了当前生成式AI在科学严谨性验证方面的显著局限性。
链接: https://arxiv.org/abs/2607.27845
作者: Haobo Li,Eunseo Jung,Wenxiao Zhao,Feng Liu,Jiong Wang,Kaiyi Xu,Zijie Guo,Zixin Chen,Ben Fei,Fenghua Ling,Lei Bai
机构: Shanghai AI Laboratory; University of California, Los Angeles; The Hong Kong University of Science and Technology
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.
[NLP-47] MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
【速读】: 该论文旨在解决大语言模型代理在跨会话与任务中使用持久化内存时,因可写内存中的错误持续存在而导致未来行为被污染的问题。现有系统虽优化了存储与检索机制,但缺乏可靠的更新事务边界以支持原子性更新与故障恢复。为此,论文提出 MemTxn——一个位于答案模型之外的治理层,其核心在于通过三重机制实现可信更新与状态恢复:首先,采用有序补丁测试(Ordered PatchTest)验证写入操作的合法性;其次,利用时间解析器(Temporal Resolver)在事实冲突时选择可见版本;最后,借助持久化快照日志(durable snapshot journal)实现故障后的状态回滚。实验表明,在项目不相交的审计任务中,MemTxn正确识别全部60个有效原始数据并拒绝所有179个强负样本;在 LongMemEval-S 与 LoCoMo 的多键持久故障场景下,无需知晓实际物理写入集合即可完整恢复声明的活跃映射状态;在 MemoryAgentBench 的 FactConsolidation 任务中,MemTxn 在十二种答案模型配置下均取得最高平均 F1 值,相较于 Dense 模型在五个代表性设置中提升 17.06–24.07 分。因此,解决方案的关键在于构建一个外部治理层,通过事务化语义保障内存更新的可靠性、版本一致性与可恢复性。
链接: https://arxiv.org/abs/2607.27834
作者: Hanshuai Cui,Zhiqing Tang,Zhi Yao,Fanshuai Meng,Qianli Ma,Weijia Jia
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist and corrupt future behavior. Existing systems improve storage and retrieval, but they do not provide a transaction boundary for reliable updates and recovery. We therefore propose MemTxn, a governance layer outside the answer model. MemTxn verifies whether an update is supported by its source. It also selects the visible version when facts conflict and restores the application-visible state after a fault. The system uses Ordered PatchTest to validate writes, a Temporal Resolver to select versions, and a durable snapshot journal to recover state. On an item-disjoint audit, MemTxn accepts all 60 supported originals and rejects all 179 hard negatives. Under persistent multi-key faults on LongMemEval-S and LoCoMo states, it restores the complete declared active map without knowing the actual physical write set. On MemoryAgentBench FactConsolidation, MemTxn achieves the highest average F1 across all twelve answer-model configurations. It outperforms Dense by 17.06–24.07 points in five representative settings.
[NLP-48] Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
【速读】: 该论文旨在解决当前角色扮演代理(Role-playing Agents, RPAs)评估中存在的重要问题:现有基准测试通常要求代理在固定对话历史基础上进行续写,并使用与用户无关的统一评分标准进行评价,这导致评估结果无法真实反映代理在多轮交互中的实际表现。其核心局限在于:一是代理输出受先前对话历史影响,难以科学评估其在真实多轮场景下的角色扮演能力;二是用户个体差异显著,而传统固定评分标准未必与用户满意度一致。为克服这些问题,论文提出PALATE(Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation)——一个基于用户模拟器的可扩展评估框架。其关键创新在于:构建包含300个角色档案的数据库,针对每位用户训练独立的个性化用户模拟器,使其与候选RPAs开展自由形式、多轮对话;同时引入个性化评分标准以衡量用户满意度,实验表明该标准在留出标注数据上与人工判断具有更高一致性。在对16个候选系统的主评估中,PALATE能够分别刻画通用话轮质量、长时对话能力以及用户-代理交互轨迹上的个性化体验,从而实现对具体用户-代理配对的可解释性评估,而非将系统压缩为单一的、脱离用户的排名。
链接: https://arxiv.org/abs/2607.27816
作者: Yuhang Zhu(1),Mingxuan Du(1),Benfeng Xu(1 and 2),Jie Gao(2),Lingyun Yu(1),Hongtao Xie(1) ((1) University of Science and Technology of China, (2) MetaStone Technology, Beijing, China)
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 29 pages, 3 figures, including supplementary material. Resources: this https URL
Abstract:Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA’s output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
[NLP-49] Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
【速读】: 该论文旨在解决多模态情感分析(Multimodal Sentiment Analysis, MSA)中现有基于大语言模型(Large Language Models, LLMs)的方法难以捕捉情感语义深层结构与上下文交互的问题。当前方法主要依赖低层次的表面特征,无法有效建模非言语模态在结构变化与上下文动态交互中产生的复杂情感语义。其解决方案的关键在于提出一种统一框架SentiLLM,核心创新为语义对齐的结构抽象(Semantic-Aligned Structural Abstraction),通过引入双流显著性-上下文校准机制(Dual-Stream Salience-Context Calibration Mechanism),将非言语特征序列解耦为两个独立流:聚焦流(focus stream)用于捕获由文本先验引导的显著情感突变(如面部表情变化),背景流(ambient stream)则表征稳定的环境状态。通过在背景状态的约束下对情感动态进行校准,该机制实现了从连续原始信号到紧凑且语义明确的文本化标记的映射,使非言语模态可被大语言模型自然理解。该模块以极少量可训练参数作为即插即用组件,显著提升判别性能,在MOSI、MOSEI、CH-SIMS及CH-SIMS v2四个数据集上均取得优异表现,验证了结构抽象范式在多模态情感分析中的有效性。
链接: https://arxiv.org/abs/2607.27790
作者: Wei Chen,Junkai Li,Tongguan Wang,Hui Liu,Feiyue Xue,Chuanxiang Ma,Ying Sha
机构: Huazhong Agricultural University(华中农业大学); Hubei University(湖北大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted by MM 2026
Abstract:Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be viewed as feature sequences evolving over time. This isomorphism enables the transformation of non-verbal modalities into text-like tokens for unified semantic reasoning. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized to interpret complex affective sequences. However, existing LLM-based methods primarily capture low-level superficial features, failing to model affective semantics arising from structural variations and contextual interactions. To address this limitation, we propose \textbfSentiLLM, a unified framework that leverages \textitSemantic-Aligned Structural Abstraction to distill continuous raw signals into compact, semantically meaningful tokens. Specifically, we introduce a \textitDual-Stream Salience-Context Calibration Mechanism, which disentangles non-verbal feature sequences into a focus stream and an ambient stream. The focus stream captures salient sentiment shifts (e.g., facial expressions) guided by textual priors, while the ambient stream characterizes stable background states. Through calibrating these dynamic sentiment shifts against background states, SentiLLM effectively projects non-verbal modalities into a unified semantic space, making them naturally understandable for LLMs. Serving as a plug-and-play module, SentiLLM significantly improves discriminative performance with only a small number of trainable parameters. Our method achieves superior performance on four datasets, MOSI, MOSEI, CH-SIMS, and CH-SIMS v2, demonstrating the effectiveness of the structural abstraction paradigm in MSA. Our code is available at: \hrefthis https URL.
[NLP-50] Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在进行链式思维(chain-of-thought)推理时,其推理过程以非结构化文本形式呈现,导致用户难以判断哪些推理步骤有充分依据、哪些备选方案被认真考虑过,以及最终结论与被舍弃路径之间的差异。为应对这一问题,论文提出一种新框架,通过加权合并多个LLM推理链中提取的有向无环图(Directed Acyclic Graph, DAG),实现对推理结构而非仅答案的集成,从而生成“共识推理”(Consensus Reasoning)。该方法的关键在于:基于每个推理步骤被独立推理轨迹支持的次数进行加权,以量化其可信度,进而构建可解释的共识推理图谱。实验表明,在涵盖法律条文解释、研究生级科学推理、叙事多跳推理和一阶逻辑等六项基准任务上,该框架在相同计算预算下优于多数投票基线,最大准确率提升达3.1%(在MuSR-MM任务上)。此外,该框架在单个模型上的表现可媲美或超越自洽性(self-consistency)方法,并提供可审计的共识推理图。相关性分析显示,集成权重与人工标注的推理质量评分(LLM-judge rankings)具有显著正相关性(Spearman ρ = 0.30–0.51),且在五项数据集的对比中,共识子图在54.4%–65.4%的情况下优于导向多数投票结果的路径,验证了其合理性。研究还发现,该框架可用于分析同一问题的不同推理视角,增强了推理过程的透明性与可理解性。
链接: https://arxiv.org/abs/2607.27783
作者: Amruta Parulekar,Jinu Lee,Dilek Hakkani-Tür,Hari Sundaram
机构: University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return “Consensus Reasoning”. Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman \rho = 0.30 - 0.51 , and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.
[NLP-51] ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
【速读】: 该论文旨在解决大语言模型(LLM)智能体在长期记忆管理中面临的“不可逆演化”问题,即现有记忆系统采用单向累积与覆盖机制,缺乏对历史状态的版本控制、回溯与验证能力,导致在出现错误修正、概念漂移或记忆污染时表现脆弱,尤其在已接触后续信息后难以恢复先前正确状态。其解决方案的关键在于提出ChronoMem——一个集成于Google开源智能体开发套件(Agent Development Kit)中的语义版本控制层,通过在每次记忆写入时提交完整记忆快照,构建结构化的版本历史,并结合混合词法与语义检索、排序融合及重排序技术,实现自然语言形式的回滚请求解析,将用户意图精准映射至对应的历史版本。此外,论文还引入后暴露评估协议,以检验智能体在回滚后能否进行反事实行为(如模拟未来更新未发生时的问答与摘要),从而系统性验证记忆可逆性。实验表明,相较于仅依赖提示或检索的基线方法,ChronoMem在长周期对话基准上显著提升了回滚一致性问答与历史摘要性能,同时在语义版本选择任务中表现优异。据作者所知,ChronoMem是首个面向LLM智能体的开放源代码系统与基准,实现了全局记忆语义级回滚的系统性支持。
链接: https://arxiv.org/abs/2607.27773
作者: Yongye Su,Wujiang Xu,Chaoji Zuo,Elisa Bertino
机构: Purdue University (普渡大学); Rutgers University (罗格斯大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:LLM agents increasingly rely on long-term memory to support multi-session interaction and personalization. However, existing agent memory systems are designed around forward-only evolution, continuously accumulating, consolidating, and overwriting knowledge, with no principled mechanism to inspect, version, or revert prior states. This makes agents brittle under corrections, concept drift, and memory corruption, particularly after they have already been exposed to subsequent information. We present ChronoMem, a semantic version-control layer for agentic memory integrated into the production-ready, open-source Agent Development Kit by Google. ChronoMem commits whole-memory snapshots at each memory write, maintains structured version histories, and supports natural-language rollback requests by mapping undo intents to concrete historical versions through hybrid lexical and semantic retrieval, rank fusion, and reranking. We further introduce a post-exposure evaluation protocol that tests whether an agent can behave counterfactually after rollback by answering queries and summarizing history as if future updates had never occurred. On long-horizon conversational benchmarks augmented with evolving memory states and rollback tasks, ChronoMem substantially improves rollback-consistent question answering and history summarization relative to prompt-only and retrieval-only baselines, while achieving strong performance in semantic version selection. To our knowledge, ChronoMem is the first open-source system and benchmark for systematic semantic global memory rollback in LLM agents.
[NLP-52] Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
【速读】: 该论文旨在解决现有语音对话系统在复杂社会环境下的适应性问题,即传统系统多针对干净、双人交互场景设计,难以应对多说话人共存、背景噪声干扰及语用意图模糊的真实社交对话场景。在此类环境中,用户发言可能指向助手、其他参与者或完全无关,助手不仅需决定回应内容,还需判断是否应答。解决方案的关键在于提出Cocktail-Talker框架,该框架通过引入三个动作标记(|respond|、|listen|、|ignore|)来显式建模助手的交互行为决策机制,并结合监督微调与强化学习进行训练,使模型能够根据上下文动态选择合适的行为动作;同时,构建基于大语言模型(LLM-based)的Cocktail-DialogGen数据生成管道,以合成多样化社会场景下具有真实角色分工的多说话人对话数据,从而实现对复杂社交环境中自然且有选择性的语音交互建模。
链接: https://arxiv.org/abs/2607.27756
作者: Xilin Jiang,Riki Shimizu,Sukru Samet Dindar,Junkai Wu,Zhongweiyang Xu,Nima Mesgarani
机构: 未知
类目: ound (cs.SD); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
备注:
Abstract:Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant’s behavior with three action tokens: |respond|, |listen|, and |ignore|, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in |respond| mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.
[NLP-53] Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
【速读】: 该论文旨在解决当前大型视觉语言模型(Large Vision Language Models, LVLMs)在评估其推理能力时存在的局限性问题,即现有评估体系多局限于单一模态感知或特定领域(如数学、编程),缺乏对开放世界环境中感知与推理协同能力的综合性评测。为此,论文提出以视觉错觉(visual illusions)作为诊断工具,构建了真实世界采集的“IllusionReasoning”基准数据集,包含多样化的错觉图像及对应的问答标注。该方案的关键在于利用人类视觉系统在面对客观信号时产生的认知偏差现象,检验模型在复杂、非理想化场景下将感知信息转化为合理推理的能力。实验表明,多数LVLMs的推理能力远未达到宣称水平,揭示了当前模型在跨模态理解与深层推理方面的不足,为未来优化提供了新方向。
链接: https://arxiv.org/abs/2607.27747
作者: Liangjie Zhao,Jiaqing Lyu,Kexin Tang,Zecheng Fang,Rong Yin,Yulan Hu,Da Li,Jianing Li
机构: Adelaide University; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Tsinghua University; Amap, Alibaba Group; Beihang University
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.
[NLP-54] A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding AAAI2027
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)推理过程中由生成式 AI(Generative AI)引发的内存带宽瓶颈问题,尤其关注推测解码(Speculative Decoding)在实际应用中因起草开销(drafting overhead)、token 接受率(token acceptance)与推测长度(speculation length)之间的权衡而带来的加速受限问题。其核心挑战在于:当边际接受概率低于相对起草成本时,延长推测范围反而会降低整体效率。为此,论文提出一种无需训练的自推测解码框架 SparseSpec-L,其关键创新在于通过动态稀疏化并可召回的键值缓存(KV cache),直接从目标模型生成轻量级草稿;同时利用全上下文验证阶段产生的每头注意力统计信息作为无额外前向传播的置信度信号,实现对关键历史 token 的高效召回,而无需永久丢弃密集的 KV 缓存;此外,引入基于在线熵的控制器,根据预期每步效率动态选择最优推测长度,从而在多任务、多模型规模下实现一致的端到端加速,最高可获得显著超越自回归解码的加速比,且保持目标模型输出分布的准确性。
链接: https://arxiv.org/abs/2607.27735
作者: Yuesong Liu,Yuan Zeng,Min Lyu,Ruilin Liu,Yu Guo,Yinlong Xu
机构: 未知
类目: Computation and Language (cs.CL)
备注: 9 pages, 4 figures, subbmited to AAAI 2027
Abstract:Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptance, and speculation length. We present a unified efficiency analysis showing that extending the speculation horizon can reduce rather than improve speedup when the marginal acceptance probability falls below the relative drafting cost. Guided by this analysis, we introduce SparseSpec-L, a training-free self-speculative decoding framework for long-context inference. SparseSpec-L generates lightweight drafts directly from the target model using a dynamically sparsified and recallable KV cache. It recycles per-head attention statistics produced during full-context verification as a no-extra-forward importance signal, allowing critical historical tokens to be recalled without permanently discarding the dense KV cache. An online entropy-based controller further selects the speculation length according to expected step-wise efficiency. Experiments across multiple long-context tasks and model scales show consistent end-to-end acceleration, with up to speedup over autoregressive decoding while preserving the target model’s output distribution.
[NLP-55] Baikal: Structured Search for Deep Research over Data Lakes
【速读】: 该论文旨在解决在数据湖(data lake)中进行深度研究时,现有基于大语言模型(LLM)的代理在有限预算下难以有效平衡探索(exploration)与利用(exploitation)的问题。具体而言,传统方法依赖迭代式检索与生成,易过度聚焦局部高回报证据,导致对不同语义区域覆盖不足,影响研究的全面性与多样性。为此,本文提出Baikal框架,将异构证据聚类为语义区域,并在此基础上进行自适应搜索,以实现探索与利用的动态平衡。其核心创新在于:在选定的语义区域内,生成并探究基于区域的子问题,利用发现质量作为奖励信号,更新各区域的价值估计,并结合多种策略(如随机选择、LLM引导、贝叶斯ε-贪心及上置信界UCB)指导后续搜索。实验在HybridQA和TAT-QA两个数据湖上进行,分别包含10,993和2,757张表格,以及数十万篇维基百科和金融报告文本。通过新提出的涵盖可证伪性、相关性、多样性和实用性四维度的评估体系,使用GPT-5-mini评分表明,Baikal在多种区域选择策略下均表现优异,其最优配置相比最强基线(DeepSearcher及OpenCode研究代理)在HybridQA和TAT-QA上分别提升报告得分28%和36%。分析显示,性能提升主要归因于对语义证据区域的结构化组织与系统性探索,显著增强了结果的可证伪性、多样性及实用性。该工作验证了结构化语义探索在异构数据湖中系统性研究与发现中的关键价值。
链接: https://arxiv.org/abs/2607.27726
作者: Dhruv Agarwal,Rishitha Guttapalle Mohan,Aarti Kumari,Ashi Sinha,Athulya Anil,Kavitha Srinivas,Horst Samulowitz,Andrew McCallum
机构: University of Massachusetts Amherst (马萨诸塞大学阿姆赫斯特分校); University of Maryland, College Park (马里兰大学学院帕克分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Deep research over data lakes requires an LLM agent to investigate evidence across thousands of heterogeneous tables and passages to synthesize a report. Existing methods perform iterative retrieval and generation, letting accumulated context determine what to investigate next, which can overexploit locally promising evidence and fail to cover distinct semantic regions under a fixed budget. To address this, we cast deep research over data lakes as a budgeted search problem and present Baikal - a framework that clusters heterogeneous evidence into semantic regions, then searches over them adaptively to balance exploration and exploitation. Within each selected region, Baikal generates and investigates region-grounded subquestions, using finding quality as rewards to update region-level value estimates and guide search under policies ranging from random and LLM-guided selection to Bayesian \epsilon -greedy and UCB. We evaluate Baikal on 15 queries each over HybridQA and TAT-QA data lakes containing 10,993 and 2,757 tables, respectively, together with 227K Wikipedia passages and 13K financial report passages. We assess research quality with a new rubric covering groundedness, relevance, diversity, and utility, and use GPT-5-mini to score Baikal and strong baselines, including DeepSearcher and an OpenCode research agent with retrieval and clustering variants. Across both data lakes, Baikal performs strongly under several region-selection policies; its best configuration improves report scores over the strongest baselines by 28% on HybridQA and 36% on TAT-QA. Our analyses attribute these gains to organizing and exploring semantic evidence regions, which improves groundedness and diversity and yields more useful findings under the same subquestion budget. These results demonstrate the value of structured semantic exploration for systematic research and discovery over heterogeneous data lakes.
[NLP-56] Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention
【速读】: 该论文旨在解决稀疏注意力(Sparse Attention)在长序列推理中因动态Top-K选择机制导致的计算开销问题。现有方法需对完整键值(KV)缓存进行全局评分与排序,使选择器成本随上下文长度线性增长,严重制约了稀疏注意力在长文本生成任务中的实际效率。其解决方案的关键在于提出一种无需训练的ReTopK方法,通过复用历史查询-支持(query-support)对的检索决策来加速动态Top-K注意力。ReTopK利用相似查询往往关注重叠支持区域的特性,为每个注意力头维护一个有限大小的历史缓存,并在新查询到来时,通过相似性检索召回最接近的历史查询,将它们的支持集与近期窗口内的支持合并,仅对紧凑的候选集重新排序,从而大幅减少当前查询的评分开销。同时引入基于相似性的回退机制以应对不可靠复用情况,并周期性刷新缓存以抑制缓存漂移。该方法不复用历史得分、注意力权重或输出,仅复用被选中的索引,保持了原始模型的精确性。实验表明,在16K–128K上下文长度下,ReTopK在PG19困惑度、NIAH和LongBench等指标上均优于其他近似方法;在128K上下文、K=512时,仅增加0.50%的困惑度,但实现3.07倍的注意力计算加速,显著提升了长序列生成的效率与质量。
链接: https://arxiv.org/abs/2607.27692
作者: Wenshuai Yao,Wenyong Zhou,Hanyong Shao,Yizhe Chen,Zhiyuan Ning,Yuannuo Feng,Ru Huang,Kechao Tang
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 9 pages, 9 figures, and 5 tables
Abstract:Top- K sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key–value (KV) entries. However, identifying this subset still requires scoring the current query against the full KV cache and performing global Top- K selection, leaving selector cost linear in context length and limiting the practical efficiency of sparse attention for long-context decoding. In this paper, we introduce ReTopK, a training-free method that accelerates dynamic Top- K attention by reusing historical retrieval decisions. ReTopK builds on the observation that similar queries often attend to overlapping supports and that partially overlapping supports can still preserve most of the Exact Top- K attention mass. For each attention head, it maintains a bounded cache of historical query–support pairs, retrieves the most similar cached queries for each new query, unions their stored supports with a recent window, and reranks only the resulting compact candidate set using exact current-query scores. A similarity-based fallback invokes full-history Exact Top- K when reuse is unreliable, while periodic exact refreshes limit cache drift. ReTopK retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs. Across 16K–128K contexts, ReTopK achieves the lowest PG19 perplexity and the highest NIAH and LongBench scores among the evaluated approximate methods. At 128K with K=512 , ReTopK incurs only a 0.50% perplexity increase over Exact Top- K while accelerating attention computation by 3.07\times .
[NLP-57] ght Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
【速读】: 该论文旨在解决低秩适应(LoRA)在微调大规模预训练模型时的统计特性不完整理解问题,特别是缺乏对泛化误差的紧致下界以及如何最优选择LoRA秩r的理论指导。其核心解决方案在于通过局部Rademacher论证建立了经验风险最小化器在目标适配秩不超过r时的上界为O~(rd/n),并基于Fano型的秩-r子空间填充构造出匹配的极小极大下界Ω(rd/n),证明该下界适用于任何输出位于秩-r LoRA类中的估计器。结合上下界分析,揭示了两种不同优化范式下的秩选择二分性:对于受约束的经验风险最小化器,最优秩等于内在秩r*,过高的秩会显著损害性能;而对于核范数-截断类的自适应估计器,过拟合秩无害,收敛速率在r ≥ r时稳定于Θ~(rd/n)。该理论结果首次在充分指定的局部二次框架下刻画了LoRA微调的统计复杂度,并将实践中观察到的过参数化惩罚归因于未正则化的经验风险最小化,而非LoRA结构本身。理论预测在合成迹回归基准及三个真实场景(DistilBERT/RoBERTa在SST-2/MRPC任务)的LoRA微调实验中得到验证,所有配置均呈现预测的U形验证损失曲线,其中两个配置在高秩时表现出显著的损失膨胀(配对置换检验p = 0.016)。
链接: https://arxiv.org/abs/2607.27680
作者: Arunan J
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Springer Nature Submission
Abstract:Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood. Existing generalization results provide upper bounds of the form O~(sqrt(rd/n)) or O~(rd/n), but a matching lower bound is missing, and the question of how to choose the LoRA rank r has no formal answer. Both gaps are closed here. A local Rademacher argument establishes an upper bound of O~(rd/n) on the excess risk of the empirical risk minimizer over rank-r LoRA, whenever the target adaptation has rank at most r. A matching minimax lower bound of Omega(rd/n) is then proved via a Fano-type packing of the rank-r subspace of R^d x d; the bound applies to any estimator whose output lies in the rank-r LoRA class. Combining the two yields a rank-selection dichotomy. For the constrained empirical risk minimizer, the optimal rank equals the intrinsic rank r*, and over-ranking strictly hurts. For adaptive estimators of the nuclear-norm-then-truncate type, over-ranking is harmless and the rate saturates at Theta~(r* d / n) regardless of r. Taken together, the three results characterize the statistical complexity of LoRA fine-tuning within the well-specified locally quadratic regime, and identify the empirically observed over-parameterization penalty as a property of unregularized empirical risk minimization rather than of the LoRA class itself. Predictions of the theory are verified on a synthetic trace-regression benchmark and on real LoRA fine-tuning across three (model, task) configurations covering DistilBERT and RoBERTa on SST-2 and MRPC. All configurations exhibit the predicted U-shape in validation loss, with two showing statistically significant loss inflation at large ranks (paired permutation p = 0.016).
[NLP-58] ICLE: Modeling Fine-Grained Traits for Holistic Essay Scoring NAACL2024
【速读】: 该论文旨在解决当前自动化作文评分(AES)模型评估过度依赖ASAP语料库所带来的泛化能力质疑问题。现有模型虽在ASAP上表现良好,但其在其他语料上的迁移性能尚不明确,限制了模型实际应用的可信度。为此,论文提出ICLE++语料库,该语料库包含经过整体评分(holistic scores)与特质维度评分(trait-specific scores)双重标注的论说性学生作文,不仅可用于检验基于ASAP训练的模型在跨数据集场景下的泛化能力,还可支持新型AES任务如多特质评分(multi-trait scoring)和跨提示评分(cross-prompt scoring)的评估。其关键贡献在于构建了一个高质量、多维度标注的标准化语料库,填补了当前AES研究中对多样化、可扩展评估资源的迫切需求。
链接: https://arxiv.org/abs/2607.27671
作者: Shengjie Li,Vincent Ng
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted as a long paper to NAACL 2024
Abstract:The majority of the recently-developed models for automated essay scoring (AES) are evaluated solely on the ASAP corpus. However, ASAP is not without its limitations. For instance, it is not clear whether models trained on ASAP can generalize well when evaluated on other corpora. In light of these limitations, we introduce ICLE++, a corpus of persuasive student essays annotated with both holistic scores and trait-specific scores. Not only can ICLE++ be used to test the generalizability of AES models trained on ASAP, but it can also facilitate the evaluation of models developed for newer AES problems such as multi-trait scoring and cross-prompt scoring. We believe that ICLE++, which represents a culmination of our long-term effort in annotating the essays in the ICLE corpus, contributes to the set of much-needed annotated corpora for AES research.
[NLP-59] Looped Transformers with Source-Centered State Evolution
【速读】: 该论文旨在解决循环型Transformer(Looped Transformers)中输入条件性与共享递归结构下锚点不变性之间的矛盾问题。在传统方法中,由于共享的Transformer块需在整个递归深度上处理不断变化的隐藏状态,且在加性注入机制下每一步都会重新引入输入相关信号,导致隐藏状态偏离初始锚点,从而破坏了输入依赖性的稳定保持。其核心解决方案是提出源中心状态演化(Source-Centered State Evolution, SCSE),通过引入可学习的锚点(anchor)和初始偏差(initial deviation),使模型在保持输入依赖性的同时,实现零偏差映射为零的精确锚点不变性。SCSE通过零偏差掩码强制实现零偏差下的恒定锚点状态,并将零偏差驱动偏置设为零,从而消除有害的“零偏差强迫偏置”——该偏置在理论上可能对任务性能产生负面影响。研究表明,该偏置是一个可调节的设计自由度,而SCSE选择将其置零以确保精确锚点不变性。实验表明,在WikiText-2、WikiText-103、直接网络语料预训练、未见网络文本迁移以及LAMBADA补全等任务中,SCSE均提升了受控递归质量边界;消融分析进一步确认,可学习锚点和锚坐标偏差递归是性能提升的主要来源,且对训练后模型的案例研究验证了锚点响应诊断与实际递归运动的一致性。
链接: https://arxiv.org/abs/2607.27656
作者: Bum Jun Kim,Kohei Hayashi,Shunsuke Kamiya,Masanori Koyama,Yusuke Iwasawa,Yutaka Matsuo
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 5 figures
Abstract:Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective depth at a fixed parameter count. However, that shared block must then govern an entire trajectory of varying hidden states over trained and extrapolated depths. Furthermore, in additive-injection looped Transformers, an input-conditioned signal is reintroduced at every recurrent step, so applying the shared transition at an input-conditioned reference can still move the hidden state. In this paper, we propose Source-Centered State Evolution (SCSE), which is designed to reconcile input conditioning with reference-preserving shared recurrence. Specifically, SCSE retains input dependence through its learned anchor and initial deviation, allows nonzero deviations to drive recurrent computation while mapping zero deviation to zero, and guarantees exact anchor invariance through its zero-deviation mask. The designated anchor is thereby a one-step fixed point by construction. The zero-deviation forcing bias is the next deviation produced from the anchor itself and vanishes in SCSE, while nonzero deviations remain active and support state-dependent recurrent computation. Our theory shows that the zero-deviation forcing bias is a design degree of freedom whose task effect can be harmful, neutral, or beneficial; SCSE resolves this choice in favor of exact anchor invariance by setting the bias to zero. Across WikiText-2, WikiText-103, direct web-corpus pretraining, held-out web-text transfer, and LAMBADA completion, SCSE improves the controlled recurrent quality frontier. Ablation studies identify the learned anchor and the anchor-coordinate deviation recurrence as the primary contributors to the gain, and a trained-model case study grounds the anchor-response diagnostic in observed recurrent motion.
[NLP-60] From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models SIGIR2026 SIGIR
【速读】: 该论文旨在解决现有大语言模型(LLM)在事件分析任务中因文档粒度限制、任务设计单一及数据来源局限而导致的综合理解能力不足的问题。其解决方案的关键在于构建首个系统性的多粒度事件分析基准——MiGUE-Bench,通过引入基于大语言模型的自校正标注框架MiGUE-Pipeline,实现高质量、自动标注事件数据的大规模获取;同时设计了事件检测、关系推理、结构归纳与未来预测四类核心任务,覆盖从原子级事件细节到跨文档复杂叙事的不同层次,全面评估模型在多粒度事件分析中的能力边界,揭示当前主流大模型与检索增强生成(RAG)方法在复杂事件理解上的关键缺陷,为后续提升提供了明确方向。
链接: https://arxiv.org/abs/2607.27654
作者: Tao Wen,Shuai Shao,Pei Ke,Xu Han,Jie Zou,Guannan Li,Tao Tian,Jinjie Qiu,Lan Wang,Ke Qin
机构: University of Electronic Science and Technology of China (电子科技大学); Tsinghua University (清华大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 9 pages. Published in the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)
Abstract:Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.
[NLP-61] Harness-G: A Graph-Structured Harness for Search Agents
【速读】: 该论文旨在解决强化学习(Reinforcement Learning, RL)搜索代理在多轮交互中因检索建模不当导致的“检索等价性坍塌”(retrieval-equivalence collapse)问题。具体而言,现有方法将检索过程建模为自由形式的自然语言查询生成,并依赖最终答案奖励进行优化,但在训练过程中发现,相同问题的不同推理轨迹虽生成不同的查询语句,其累积证据集却趋于高度重叠,导致检索决策间的有效区分度丧失,进而削弱了策略更新的信号质量。其解决方案的关键在于提出一种图结构化的检索框架Harness-G,重新设计策略与环境之间的接口:将自由形式的查询生成转化为有限动作选择,即策略从环境中预设的证据句子或实体中选择动作,或直接作答,而环境负责维护检索状态、验证并执行每一步选择。这一设计显著降低了语言上的歧义性(linguistic aliasing),使同一状态下的不同动作可直接比较。在此基础上,引入结构化非短期信用分配(Structured Non-myopic Credit, SNC),利用冻结的答案评分器评估所选动作与其替代动作的相对收益,并将下游收益回溯分配给促成该选择的早期动作,从而增强策略对关键检索步骤的学习能力。实验表明,Harness-G在六个问答基准上均取得最优性能,在1.5B和3B模型规模下分别优于最强基线Graph-R1 10.74和3.98个F1点。
链接: https://arxiv.org/abs/2607.27652
作者: Yanning Hou,Haoyuan Chen,Sihang Zhou,Xiaoshu Chen,Xirui Liu,Duanyang Yuan,Lingyuan Meng,Quan Liu,Jian Huang
机构: 未知
类目: Computation and Language (cs.CL)
备注: Code: this https URL
Abstract:Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more structured credit signals, but rarely examine whether retrieval is properly formulated at the policy-environment interface. We observe pronounced retrieval aliasing during Search-R1 training: rollouts for the same question continue to generate distinct query strings, yet their accumulated evidence sets increasingly overlap. We call this phenomenon retrieval-equivalence collapse; in this regime, trajectories approach utility equivalence with respect to retrieval decisions, leaving within-group returns with little effective retrieval contrast. To address this problem, we propose Harness-G, a graph-structured retrieval framework that redesigns this interface. It reformulates free-form query generation as finite action selection: the policy selects an evidence sentence or entity, or chooses to answer, while the environment constructs the menu, tracks retrieval state, and validates and executes each choice. This interface reduces linguistic aliasing and makes same-state alternatives directly comparable. Building on this interface, we introduce Structured Non-myopic Credit (SNC), which uses a frozen answer scorer to compare the selected action with its alternatives and assigns downstream gains to the earlier actions that enabled them. Across six QA benchmarks, Harness-G achieves the highest average F1 at both evaluated model scales, outperforming the strongest baseline, Graph-R1, by 10.74 points at 1.5B and 3.98 points at 3B.
[NLP-62] ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
【速读】: 该论文旨在解决大语言模型在数学推理任务中因长推理路径与稀疏结果奖励导致的令牌级信用分配不可靠问题。现有基于近端策略优化(Proximal Policy Optimization, PPO)的方法虽具备理论上实现细粒度信用分配的潜力,但在实际应用中,标准价值函数难以准确评估中间推理状态,进而产生噪声较大的优势估计,影响策略优化效果。其解决方案的关键在于提出一种参考引导且差异感知的PPO框架——ReDiPPO:引入一个基于参考答案的参考引导价值网络,利用训练阶段的先验知识提供更精确的价值估计;同时保留标准价值网络,并通过计算两者在令牌层面的估计差异(reference-standard discrepancy),量化推理过程中的困难状态。该差异被用作权重因子,动态调整相应令牌的优势值,从而提升策略更新的准确性。实验结果表明,ReDiPPO显著提升了价值估计精度,在多个数学推理基准上均优于PPO、DAPO和GSPO等先进基线方法。
链接: https://arxiv.org/abs/2607.27631
作者: Zhenrong Zhang,Fei Wu,Jun Du,Jianshu Zhang,Si Wei
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among existing policy optimization methods, Proximal Policy Optimization (PPO) remains particularly appealing because its learned critic can, in principle, provide token-level credit assignment. However, in mathematical reasoning tasks characterized by long reasoning horizons and sparse outcome rewards, reliable token-level credit assignment remains challenging. The standard critic often fails to accurately evaluate intermediate reasoning states, resulting in noisy advantage estimates and suboptimal policy updates. In this paper, we propose ReDiPPO, a Reference-guided and Discrepancy-aware PPO framework for mathematical reasoning. ReDiPPO introduces a reference-guided critic that uses reference answers as training-time privileged signals to provide more accurate value estimation. Meanwhile, it retains a standard critic and quantifies the token-level reference-standard discrepancy between the standard value estimate and the reference-guided value estimate. This discrepancy serves as an indicator of difficult reasoning states and is used to reweight the corresponding token-level advantages during PPO optimization. Extensive experiments on diverse mathematical reasoning benchmarks demonstrate that ReDiPPO improves value-estimation accuracy and consistently outperforms strong policy optimization baselines, including PPO, DAPO, and GSPO, in final reasoning performance. Our code is available on this https URL.
[NLP-63] DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
【速读】: 该论文旨在解决当前基于大语言模型(LLM)的手语翻译(SLT)方法中存在的两个关键问题:一是语言先验退化(language-prior degradation),即现有方法未能有效利用LLM强大的语言建模能力,导致生成文本不连贯;二是词汇保真度不足,因视频与文本通常仅在句子层面进行对齐,难以保证细粒度词汇级别的准确匹配。为应对上述挑战,论文提出DualAnchor框架,其核心创新在于引入两种互补的锚定机制:一为词元级语言先验锚定(Token-level Prior Anchoring, TPA),通过在每一步解码时将多模态解码器正则化至冻结LLM在相同自回归前缀下的下一个词元分布,以保留并强化语言先验;二为最优传输对齐(Optimal Transport Alignment, OTA),将视觉-文本匹配建模为熵正则化的部分最优传输问题,借助Sinkhorn优化实现视觉词元与文本内容词元间的软对齐,从而提升词汇级保真度。实验结果表明,DualAnchor在PHOENIX-2014T和CSL-Daily数据集上均取得优异性能,且消融分析证实了两种锚定机制的协同作用:TPA显著提升生成流畅性,OTA有效减少细粒度词汇错误。
链接: https://arxiv.org/abs/2607.27614
作者: Hongbin Zhang,Junhao Liu,Xuefeng Bai,Youcheng Pan,Yang Xiang,Kehai Chen
机构: 1. Tsinghua University (清华大学); 2. Institute for AI Industry Research (AIR), Tsinghua University (清华大学人工智能产业研究院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM’s language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.
[NLP-64] AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
【速读】: 该论文旨在解决企业年报中关于外汇风险管理、衍生品使用、自然对冲及显性非使用等信息呈现弱结构化导致的可量化性差的问题。现有方法难以从非结构化文本中准确提取与验证企业的对冲披露行为,从而限制了对外汇风险敞口管理实践的系统性研究。其解决方案的核心在于提出AWARE-FX——一个可审计的生成式AI/NLP决策支持系统,通过融合专业词典、否定句与会计状态逻辑、渠道特异性财务编码器、精确证据门控机制、保守聚合策略以及审计日志,实现对企业年报文本中对冲相关陈述的可追溯、可验证的度量。该系统在2008–2025年期间覆盖24,909个香港企业-年度样本,成功提取并评分543,527条证据片段。关键创新点在于将信息检索、状态判断、分类识别、不确定性处理、聚合方式与外部验证等环节设计为独立可审计模块,确保整个流程的透明性与可靠性。实证结果表明,尽管通用大模型(如Qwen3-8B)在部分任务上表现不佳,但基于FinBERT的编码器在多数任务中展现出更高且稳定的F1分数(0.702–0.872),并通过置信度阈值剔除低置信度样本后进一步提升性能。更重要的是,严格定义的外汇风险评分与基线及压力期的外汇暴露呈负相关,提供了外部效度验证,证实了该系统的测量有效性,而非因果推断。因此,AWARE-FX不仅提供了一套可复现的自动化分析框架,更构建了一个兼顾准确性、可解释性与可审计性的生成式智能分析范式。
链接: https://arxiv.org/abs/2607.27611
作者: Qi Wang
机构: University of Nottingham (诺丁汉大学)
类目: Computation and Language (cs.CL); Risk Management (q-fin.RM)
备注: 40 pages, 4 figures, 12 tables. Preprint; not peer reviewed
Abstract:Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-use. This study develops AWARE-FX, an auditable AI/NLP decision-support system that converts report text into traceable firm-year hedging-disclosure measures. The system combines a professional-source lexicon, negation and accounting-status logic, channel-specific financial encoders, exact evidence gates, conservative aggregation, and an audit ledger. Across 24,909 Hong Kong firm-years from 2008-2025, it retrieves and scores 543,527 snippets. Reliability is evaluated through ablations, a stratified 300-snippet human audit, three-seed FinBERT-ModernBERT comparisons, strict 2023-2025 temporal tests, probability calibration, selective prediction, and fixed-prompt generative-model benchmarks. FinBERT has the higher mean F1 in seven of eight encoder task-split comparisons; its temporal F1 ranges from 0.702 to 0.872. Abstaining on the 20% least-confident temporal observations raises retained-sample F1 by 0.050-0.077. Deterministic Qwen3-8B performs strongly on commodity and negation evidence but poorly on foreign-debt and accounting-context labels, showing that a general-purpose LLM does not uniformly replace domain constraints. The strict FX score is negatively associated with linked baseline and stress-period FX exposure, whereas the generic broad score is not. These associations provide external construct validation, not causal estimates of hedging effectiveness. AWARE-FX contributes a tested decision-support architecture in which retrieval, status logic, classification, uncertainty handling, aggregation, and external validation remain separately auditable.
[NLP-65] Beyond Similarity: Grounded Agent ic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
【速读】: 该论文旨在解决现有计算型互文性分析方法在识别文本复用行为时仅能提供相似度评分或平行段落列表,而无法阐明复用的方式(how)与动机(why)这一核心缺陷。其解决方案的关键在于将细粒度互文性提取重构为一项代理任务(agentic task),利用大语言模型(LLM)通过受约束的工具接口,对两个文本单元进行完整阅读,并精确标注复用内容在原文中的字符跨度,同时依据五维分类体系(形式、方面、源标记、功能、立场)对复用行为进行语义化归类。该方法在《论语》与《汉书》的全面对比中得到验证,由三位领域专家对多模型生成的候选对进行仲裁,构建出包含2,533个互文对的专家标注基准。基于此基准,研究评估了12个LLM的精度(56%–93%)、成本差异(达51倍)及置信度校准水平,揭示出表层可观察维度(如形式与来源标记)具有较高一致性,而涉及意图推断的维度(如功能与立场)则存在争议,从而界定出标注结果的有效边界。进一步将经验证的提取器扩展至《二十四史》全集(65,380组比较,5,766对互文),发现相似度分数无法捕捉的语料级结构得以显现:尽管整体引述模式在十八个世纪中保持稳定,但具体引用的字面忠实度呈下降趋势,符合文化吸引理论所预期的“集体稳定、个体漂移”特征。研究已公开提取协议与专家标注基准,为后续互文性研究提供可复现的范式。
链接: https://arxiv.org/abs/2607.27595
作者: Zhaoji Wang,Wanyu Si,Jun Wang
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
备注: 9 pages, 4 figures, 3 tables
Abstract:Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why. We recast fine-grained intertextuality extraction as an agentic task in which a large language model (LLM) reads two text units in full and, through a constrained tool interface, must ground each proposed reuse in exact character spans on both sides and label it under a five-dimension typology of reuse (form, aspect, source-marking, function, stance). We validate the approach on an exhaustive comparison of the Analects with the Book of Han, where three domain experts adjudicate a pooled multi-model candidate set into a benchmark of 2,533 intertextual pairs. Against this standard we study twelve LLMs, reporting precision (56%-93%), a 51 \times cost spread at comparable quality, and how well their confidence is calibrated. Expert agreement traces a reliability gradient: dimensions legible on the textual surface are annotated consistently, while those requiring inference of intent are contested, delimiting the claims such annotation supports. Scaling the validated extractor to the full Twenty-Four Histories (65,380 comparisons, 5,766 pairs) recovers corpus-level structure a similarity score cannot express. The interpretive composition of citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. Stability in the aggregate with drift in the individual case is what a cultural-attraction account expects. We release the extraction protocol and the expert-adjudicated benchmark.
[NLP-66] Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLM s
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)推理过程中前馈网络(Feed-forward Networks, FFNs)带来的高内存访问与计算开销问题,尤其是针对无训练(training-free)激活稀疏化方法在高稀疏度下模型质量显著下降的瓶颈。其核心挑战在于现有方法因通道选择策略受限,难以在保持模型性能的同时实现高效稀疏化。解决方案的关键创新在于提出一种两阶段无训练框架Prox,其核心洞察是:稀疏执行仅依赖于中间状态(如SwiGLU中间态)所诱导的通道掩码,而该掩码可通过其元素的相对大小排序(magnitude ranking)而非精确值构建。具体而言,第一阶段利用输入稀疏性和量化代理权重生成共享掩码;第二阶段仅对选定通道进行精确计算,从而实现三个投影层的稀疏执行。该方法在六个模型家族的十种LLM上均优于现有无训练基线,在70%的FFN稀疏度下实现了最高1.99倍的端到端解码加速,并兼容量化和稀疏注意力机制。
链接: https://arxiv.org/abs/2607.27591
作者: Jinyi Liu,Wei Chen,Pengyu Chen,Xinyi Yuan,Minghe Bai,Guoquan Wu,Jun Wei
机构: 1. Institute of Automation, Chinese Academy of Sciences (中国科学院自动化研究所); 2. University of Chinese Academy of Sciences (中国科学院大学); 3. School of Artificial Intelligence, Renmin University of China (中国人民大学人工智能学院)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activation sparsification. However, existing training-free methods suffer substantial model-quality degradation at high sparsity due to limitations in their channel-selection strategies. We observe that the SwiGLU intermediate state provides a highly effective channel-selection signal, but obtaining it requires costly dense computation. To address this, we present \emphProx, a two-stage training-free framework for sparse SwiGLU FFNs. Prox hinges on the key insight: sparse execution requires only the channel mask induced by the intermediate state, which can be constructed from the magnitude ranking of its entries rather than their exact values. Specifically, Stage 1 uses input sparsity and quantized proxy weights to construct a shared mask; Stage 2 computes the selected channels exactly, enabling sparse execution of all three projections. Across ten LLMs from six model families, Prox outperforms training-free baselines at all sparsity levels, achieves up to a 1.99\times end-to-end decoding speedup at 70% FFN sparsity, and is compatible with quantization and sparse attention.
[NLP-67] raining Skills Like Parameters via Self-Supervised Semantic Diffusion
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在高度专业化、开放式领域(如创意剧本创作)中难以达到人类专家水平的问题。现有方法多依赖于后训练策略,如监督微调或强化学习,但这些方法不仅需要对闭源前沿模型的权重访问权限,且计算开销巨大,同时所学知识局限于单一检查点,无法被人工审查。近期基于智能体持续学习的方法虽尝试通过积累外部文本技能来弥补差距,却仍严重依赖昂贵的人工标注或不可靠的“大模型作为裁判”反馈进行反思。为突破这一瓶颈,本文提出一种受扩散模型“噪声-重建”范式启发的无监督自演化智能体框架,其核心创新在于利用高质量的人类创作成果构建自监督信号,替代显式的外部评分机制。训练过程遵循神经网络的经典流程:前向传播、损失计算与反向传播,其中损失函数由智能体生成结果与原始人类作品之间的对比决定。关键在于,被更新的并非模型权重,而是外部文本技能库。实验在短剧剧本创作这一高难度任务上验证了该方法的有效性,结果表明,该框架可使智能体自主提取并内化高度泛化的创作技能,显著提升其领域特定生成能力。该自对比反思机制为智能体实现复杂高质量人类产物的自我教学提供了可扩展的路径,无需外部监督。
链接: https://arxiv.org/abs/2607.27557
作者: Mo Li,Zixin Yin,Ting Cao,Yunxin Liu
机构: Tsinghua University (清华大学); Shanghai AI Laboratory; The Hong Kong University of Science and Technology (香港科技大学); Xiaobing.AI
类目: Computation and Language (cs.CL)
备注: Preprint, work in progress
Abstract:While Large Language Models (LLMs) demonstrate remarkable general instruction-following capabilities, they often fall short of human experts in highly specialized, open-ended domains such as creative screenwriting. Prior approaches typically adopt post-training, yet both supervised fine-tuning and reinforcement learning require weight access that closed-source frontier models do not offer, and demand heavy compute. Moreover, what is learned is tied to a single checkpoint and cannot be inspected by humans. Recent advancements in agentic continual learning instead attempt to bridge this gap by accumulating external textual skills. However, these methods heavily rely on costly human expert annotations or unreliable LLM-as-a-judge feedback for reflection. To overcome this bottleneck, we propose a novel, unsupervised self-evolving agent framework inspired by the corruption-and-reconstruction paradigm of diffusion models. Instead of relying on explicit external scoring, we leverage existing high-quality human artifacts to construct self-supervised signals. Training then follows the familiar loop of neural network training, forward, loss, and backward, with the loss coming from contrasting the agent’s reconstruction against the human original. What is updated is not model weights but an external library of textual skills. We evaluate our framework on the challenging task of short drama screenwriting. Experimental results demonstrate that our method enables the agent to autonomously extract and internalize highly generalizable skills, significantly enhancing its domain-specific generation capabilities. Furthermore, this self-contrastive reflection paradigm offers a scalable pathway for agents to teach themselves the production of complex, high-quality human artifacts, without requiring external supervision.
[NLP-68] Using Large Language Models for Idea Generation in Innovation
【速读】: 该论文旨在解决生成式 AI(Generative AI)在新产品创意生成中的有效性问题,特别是评估其相对于人类创意的优劣。研究聚焦于针对大学生群体、定价不超过50美元的新产品创意,对比了三类创意池:一是未使用生成式AI前大学生在产品设计课程中产生的创意;二是基于零样本提示(zero-shot prompting)生成的GPT-4创意;三是基于少样本提示(few-shot prompting)生成的GPT-4创意。解决方案的关键在于通过市场调研标准方法预测购买意愿概率,结合文本挖掘分析创意相似性,并由人工评分员评估创意新颖性。研究发现,尽管AI生成的创意在平均购买意愿上显著优于人类创意,且少样本提示略优于零样本提示,但其创意在新颖性感知和配对相似性方面表现较差,表明解决方案多样性较低。然而,在顶尖创意层面,AI生成的创意进入前10%的可能性是人类创意的七倍,这一优势被认定为保守估计,因未计入AI更高的生成效率。研究结论表明,尽管存在局限性,生成式AI在新产品开发中仍具有显著的高质量创意生成能力。
链接: https://arxiv.org/abs/2607.27553
作者: Lennart Meincke,Karan Girotra,Gideon Nave,Christian Terwiesch,Karl T. Ulrich
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Economics (econ.GN)
备注:
Abstract:This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for new products targeted toward college students and priced at 50 dollars or less. The first pool of ideas was created by university students in a product design course before the availability of LLMs. The second and third pools of ideas were generated by GPT-4 from OpenAI using zero-shot and few-shot prompting, respectively. We evaluated idea quality using standard market research techniques to predict average purchase intent probability. We used text mining to assess idea similarity and human raters to evaluate idea novelty. We find that AI-generated ideas outperform human-generated ideas in terms of average purchase intent, with few-shot prompting yielding slightly higher intent than zero-shot prompting. However, AI-generated ideas are perceived as less novel and exhibit higher pairwise similarity, particularly with few-shot prompting, indicating a less diverse solution landscape. When focusing on the quality of the best ideas rather than the average ideas, we find that AI-generated ideas are seven times more likely to rank among the top 10 percent of ideas, demonstrating a significant advantage over human-generated ideas. We propose that this seven-to-one advantage is a conservative estimate because it does not account for the greater productivity of AI. Our findings suggest that despite some drawbacks, AI creativity presents a substantial benefit in generating high-quality ideas for new product development.
[NLP-69] Subtract or Replay? Exact Deletion from Language-Model Memory
【速读】: 该论文旨在解决持续语言模型(Persistent Language Model)中对记忆记录进行精确删除(Exact Deletion)的难题,核心挑战在于记忆表征方式决定了删除操作的可行性。现有方法中,仅可通过代数减法(algebraic decrement)移除可定位(addressable)的影响,而由后续写入共享循环状态(shared recurrent state)所引发的隐式影响,则需通过重建(rebuilding)此前状态才能消除。论文的关键解决方案在于:基于记忆表征结构区分可直接减除与需重构的部分,并通过“重播”(replay)机制实现精确删除。具体而言,在Gemma 3中用支持向量记忆(support-vector memory)替代全局注意力层后,经低秩恢复(low-rank recovery)与键值重拟合(retained-key refit),在128亿参数规模下实现了与未摄入数据基线几乎无差异的输出(中位数KL散度达5.4×10⁻¹⁵),且在扰动测试、重学习、采样及LiRA攻击下表现一致;在480亿参数的Kimi Linear混合模型中,通过检查点回滚-重播(checkpointed rewind-and-replay)机制,可在长达18,842个词元的上下文范围内精确删除真实临床记录,确保所有递归状态与从未摄入数据的基线比特级一致,从而实现真正意义上的精确删除。因此,精确删除的本质是记忆表征的属性——即对可寻址记录进行减法操作,并对纠缠写入(entangled writes)执行重播重构。
链接: https://arxiv.org/abs/2607.27539
作者: Vishwajith Ramesh
机构: Vy Labs, Inc.(Vy实验室)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 22 pages, 8 figures
Abstract:Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic decrement; influence transformed by later writes inside shared recurrent state requires rebuilding from before the write. We test this distinction in two pretrained models against explicit record-omitted references. First, we replace Gemma 3’s global-attention layers with support-vector memory. After low-rank recovery at 1B, decrement and retained-key refit agree at the next-token output to median KL 5.4\times10^-15 over 31 support-token deletions, with +2.0% perplexity relative to a matched fine-tune. A masked-refit proxy is indistinguishable from the never-ingested floor under elicitation, relearning, sampling, and LiRA attacks. At 4B and 12B, certificate ordering persists but utility cost rises to 11.2% and 44.3% . Second, in a 48B Kimi Linear hybrid, additive writes admit a fixed decrement and diagonal decay a corrected one, whereas the delta rule makes 12 – 49% of a record’s contribution suffix-dependent. Checkpointed rewind-and-replay deletes real clinical records at contexts up to 18,842 tokens, matching never-ingested logits and all recurrent states bit for bit within a deterministic MLX implementation; replaying a correction provides exact amendment. Exact deletion is therefore a property of memory representation: subtract addressable records and replay entangled writes.
[NLP-70] hreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
【速读】: 该论文旨在解决云原生软件架构中威胁建模依赖人工分析所导致的效率低下与安全专家资源稀缺的问题。其核心挑战在于如何实现从源代码仓库到结构化攻击树的自动化生成,并将攻击步骤精准映射至多类威胁框架(如MITRE ATT&CK、CAPEC及云特定威胁矩阵)中的战术、技术与程序(TTP),同时输出可操作的缓解措施。解决方案的关键在于提出ThreatForest——一个基于多智能体系统的端到端自动化威胁建模框架,通过分阶段的代理流水线(包括仓库分析、上下文精炼、威胁生成、并行攻击树构建与TTP映射、缓解措施合成及报告生成)实现流程化控制,采用有向图编排方式确保确定性验证节点、有限重试机制与三处人机协同验证点。系统引入领域专用句向量模型(sentence-transformer)基于余弦相似度匹配攻击步骤与候选技术,实证表明该嵌入阶段是整体准确率的主要瓶颈,而非多智能体架构本身;对比单次调用基线模型,其在映射可防御性上显著提升,进一步证实限制因素在于嵌入编码器性能。因此,该研究不仅实现了跨框架的证据驱动型攻击树与缓解策略生成,还建立了可复用的评估基准体系,为后续自动化威胁建模系统提供了方法论与评价标准。
链接: https://arxiv.org/abs/2607.27528
作者: Cristian Leo,Anton Dykyi,Danny Cortegaca,Daniel Begimher,Prakash Jha
机构: Amazon(亚马逊)
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
备注: 20 pages, 12 tables, 1 figure
Abstract:Threat modeling is essential for secure software development, yet manual analysis of cloud-native architectures is slow and demands scarce security expertise. We present ThreatForest, a multi-agent system that generates structured attack trees from source code repositories, maps attack steps to adversary tactics, techniques, and procedures (TTPs) from a pluggable set of frameworks (MITRE ATTCK, CAPEC, and cloud-specific threat matrices), and synthesizes actionable mitigations. ThreatForest decomposes threat modeling into a multi-stage agent pipeline – repository analysis, context refinement, threat generation, parallel attack-tree construction with TTP mapping and mitigation synthesis, and report generation – orchestrated as a directed graph with deterministic verification gates, bounded retries, and three human-in-the-loop validation points. A domain-specific sentence-transformer maps each attack step to candidate techniques by cosine similarity; we show empirically that this embedding stage, not the surrounding pipeline, is the dominant accuracy bottleneck. We evaluate ThreatForest across seven application domains on a sixteen-dimension rubric, scored by a panel of independent LLM raters with an adversarial verification pass and expert review. Panel-measured quality reaches 0.63-0.68 (on a 0-1 scale) for threat statements, attack trees, and mitigations, but only 0.29 for embedding-only TTP mapping – a gap stable across all seven domains that isolates the binding constraint. A controlled single-call baseline on the same model more than doubles mapping defensibility, pinning the limitation on the embedding encoder rather than the multi-agent design. To our knowledge, ThreatForest is the first end-to-end system that turns a code repository into TTP-mapped attack trees with evidence-based mitigations across adversary frameworks, with a reusable framework for benchmarking such systems.
[NLP-71] Models for minimalist RAG : B1ade 335M Embedding and 1B Parameter Small Language Models
【速读】: 该论文旨在解决传统检索增强生成(Retrieval-Augmented Generation, RAG)系统对大规模预训练和显式接地监督的依赖问题,提出了一种高效且资源节约的RAG架构B1ade。其核心解决方案在于通过战略性的模型组合与奖励设计实现高性能的接地行为:B1ade-embed是一个仅335M参数的紧凑嵌入模型,通过无参融合五个预训练编码器,在无需额外训练的情况下达到子500M参数模型在MTEB榜单中的顶尖表现;B1ade-1B则是一个经过轻量级训练的自回归小语言模型(SLM),利用低成本GPU基于723M标记(220万样本)的精选上下文-问题对,采用组相对策略优化(Group Relative Policy Optimization, GRPO)进行强化学习训练,奖励函数仅优化答案相似性。关键发现为“涌现引用”(emergent attribution)——尽管未接受任何显式的来源引用监督,B1ade-1B仍能在42.4%的回答中引用检索到的段落,超出其训练分布的引用率5.5个百分点,表明在强化学习驱动下,接地行为可作为提升准确性的最优策略自发产生。在标准问答基准测试中,B1ade-1B在PopQA、PubMedQA和FEVER上分别取得81.82%、65.8%和51.09%的准确率;在端到端RAG评估中,其综合得分(涵盖正确性、完整性、连贯性和忠实性)平均达0.654,相较监督微调(SFT)模型提升10.8%,并显著缩小了与规模为其1.5倍模型之间的差距。研究结果表明,通过精心设计的模型结构与奖励机制,即可实现高效、低资源消耗的高质量RAG系统,无需依赖大规模预训练或复杂的显式监督。
链接: https://arxiv.org/abs/2607.27506
作者: Shreyas Subramanian,Mecit Gungor,Vikram Elango
机构: Amazon(亚马逊)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 28 pages, 3 figures. Submitted to COLM 2026
Abstract:Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built components: a compact embedding model and a purpose-built SLM. B1ade-embed, a 335M parameter retrieval model constructed via parameter-free fusion of five pretrained encoders achieves top MTEB scores among sub-500M models with zero additional training, and B1ade-1B, an SLM trained on low-cost GPUs using Group Relative Policy Optimization (GRPO) on 723M tokens (2.2M examples) of curated context-question pairs with rewards that optimize only answer similarity. Our central finding is emergent attribution: despite receiving no explicit supervision for source citation, B1ade-1B cites retrieved passages in 42.4% of responses, exceeding the attribution rate of its training distribution by 5.5 percentage points. This demonstrates that grounding behavior can emerge as an accuracy-maximizing strategy under RL training, without explicit reward engineering. On standard QA benchmarks, B1ade-1B achieves 81.82% on PopQA, 65.8% on PubMedQA, and 51.09% on FEVER. In end-to-end RAG evaluation, B1ade-1B achieves an average score of 0.654 across correctness, completeness, coherence, and faithfulness, a 10.8% improvement over the SFT, while closing the gap with models 1.5x its size. These results show that strategic model composition and reward design suffice for resource-efficient RAG, without large-scale pretraining.
[NLP-72] SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
【速读】: 该论文旨在解决生成式AI系统中文本知识与参数化技能(weight-space skills)之间难以无缝融合的问题,即如何实现文本信息与模型权重的跨模态协同以提升复杂任务的自主求解能力。当前研究通常将文本知识的组合与反思、以及参数空间中的技能合并视为相互独立的路径,导致无法充分发挥多模态输入的潜力。本文的关键解决方案在于将模型权重本身视为一种可被大语言模型(LLM)原生推理的额外模态,并通过前缀调优(prefix-tuning)实现参数化学习的显式控制。作者提出名为SkillSmith的增强型LLM架构,能够同时处理前缀权重和丰富的文本数据,基于指令驱动的方式合成目标技能对应的新型前缀权重。这一方法实现了文本与参数空间的深度融合,显著优于仅依赖文本或仅依赖权重空间的基线模型,从而在性能上突破单一模态适应的局限性。
链接: https://arxiv.org/abs/2607.27497
作者: Lucio M. Dery,Benedict Aaron Tjandra,Siavash Samiei,Adhiguna Kuncoro,Zohar Yahav,Jiajun Shen,Arthur Szlam
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Agentic systems driven by large language models (LLMs) regularly feature two key mechanisms to autonomously solve complex problems: synthesizing text-based knowledge and procedures from past experiences and building parametric (weight-space) skill libraries for recurring sub-goals. To date, research has largely treated these as orthogonal pursuits: either organizing textual knowledge through composition and reflection, or consolidating parametric skills via weight-space merging. Consequently, the seamless integration of text and model weights for targeted performance improvements remains largely unexplored. This work bridges this modality gap by treating model weights as an additional modality that an LLM can natively reason over. We instantiate parametric learning via prefix-tuning and augment an LLM to ingest both prefix weights and rich textual data which capture relationships to a target capability. Our augmented LLM, which we call SkillSmith, synthesizes these inputs to perform instruction-steered parametric synthesis, directly outputting new prefix weights that manifest the target skill. We demonstrate that our approach significantly outperforms both text-only and weight-space-only baselines, unlocking performance gains that are out of reach for uni-modal (text-only or weight-only) adaptations.
[NLP-73] Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights
【速读】: 该论文旨在解决时间序列数据流中存在非连续的、离散的动态变化阶段(regimes)时,如何从模型权重随时间演化的轨迹中恢复这些潜在阶段的问题。其核心挑战在于,传统方法往往假设数据分布随时间连续变化,而实际中可能存在突变的、结构化的阶段性转变,这些转变影响模型泛化能力,但难以被直接观测。解决方案的关键在于:利用隐马尔可夫模型(Hidden Markov Model, HMM)对在连续时间窗口上训练的分类器权重所构成的时序轨迹进行建模,从而识别出隐藏的状态(latent states),并据此将时间线划分为语义上一致的相位(coherent phases)。研究发现,同一状态内的模型在跨时间窗口迁移时表现出显著优于跨状态边界的泛化性能,且该优势在控制时间邻近性后依然存在,并超越了基于等长连续区间划分的基准方法。更重要的是,这些恢复出的状态与数据类别分布的变化高度相关,而非仅依赖于权重空间的几何结构;在去除类别分布偏移和滞后效应后,该迁移优势仍显著高于随机置换的零模型,表明所恢复的状态捕捉到了超越数据分布本身、与模型可迁移性相关的深层结构。该结论在多模态虚假信息检测(Fakeddit)和情感分析(Yelp)两个具有时间漂移特性的任务上均得到验证,尽管在标签分布更稳定的Yelp数据上效应有所减弱。
链接: https://arxiv.org/abs/2607.27482
作者: Kevin Guan
机构: Princeton University (普林斯顿大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:A temporally drifting data stream may pass through discrete regimes rather than changing continuously. We ask whether such regimes are recoverable from the weights of models trained on the stream, using a hidden Markov model (HMM) fit to the chronologically ordered trajectory of those weights. We study this question in two domains known to drift over time: multimodal misinformation detection, using the Fakeddit dataset; and sentiment analysis, using the Yelp dataset. We train classifiers on consecutive temporal windows and fit an HMM to the trajectory of their aligned weights, recovering latent states that partition each timeline into coherent phases. On both datasets, classifiers generalize better to data from windows sharing the state of their training window than to windows across state boundaries. This within-state transfer advantage survives a control for temporal proximity and modestly exceeds the advantage recovered by a naive partition into contiguous states of equal size. Although the states are estimated solely from model weights, they correlate more strongly with shifts in the data’s class distribution than with the weight-space geometry used to estimate them. After class divergence and lag are residualized out, the within-state advantage exceeds its permutation null on both tasks, indicating that the states recover structure relevant to transfer beyond the data distribution. Every effect replicates on both tasks but is attenuated on Yelp, whose label distribution is more temporally stable.
[NLP-74] Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
【速读】: 该论文旨在解决在计算资源、延迟和鲁棒性约束下,任务导向对话系统中如何系统性选择可部署的开源大语言模型(open-weight language models)这一关键问题。其解决方案的关键在于开展一项涵盖41个开源语言模型(覆盖15个模型家族,参数量范围为135M至9B)的零样本评估,覆盖8个英文单标签意图分类数据集及一个包含5个标注示例的ATIS五样本辅助评估。评估不仅包括标准基准测试,还引入大规模语音助手语料库与生产环境衍生的电商数据集,全面考察了精确匹配准确率、置信度校准、对现实输入扰动的鲁棒性、模型排名的统计可靠性、部署效率以及基准饱和度等多维度指标。研究发现,经过指令微调的30亿参数模型可超越部分70亿参数的基础模型;在MASSIVE数据集上,领先模型间的性能差异在成对麦克尼马尔检验下无统计显著性;而如SNIPS等常用基准已出现饱和现象,无法有效区分当前主流开源模型。此外,指令微调对置信度校准的影响并非一致负面,提示需谨慎解读其影响。这些发现为实际场景中开源语言模型在意图分类任务中的选型与评估提供了实证依据。
链接: https://arxiv.org/abs/2607.27421
作者: Parishruthi Ganesh,Gerry Dozier,Cheryl Seals
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints. We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M–9B parameter range across eight English single-label intent-classification datasets. A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result. The evaluation includes standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy, we analyze confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Our results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models. Instruction tuning’s effect on confidence calibration is inconsistent rather than uniformly harmful. These findings provide practical guidance for selecting and evaluating open-weight language models for intent classification.
[NLP-75] Dimensionality and Measurement Precision in HLEs Multiple-Choice Subset
【速读】: 该论文旨在解决当前前沿语言模型评估中一个关键问题:人类最后考试(Humanity’s Last Exam, HLE)所划分的八大学科领域子分数是否真正反映了可分离的潜在能力维度,以及该基准测试能否有效区分能力相近的先进模型。其解决方案的关键在于采用心理测量学方法对HLE文本类单项选择题子集(共428个题目)进行系统分析。研究通过拟合双参数逻辑斯蒂克项目反应理论(Two-Parameter Logistic IRT)模型,发现HLE实际上主要衡量单一通用推理因子(general reasoning factor),其结构效度指标显示麦克唐纳ω_h高达0.998,各领域标签仅解释了3.5%的题目作答方差;领域内与跨领域残差相关性几乎无差异(Cohen’s d = 0.016),且各领域能力估计值与总分高度相关(相关系数r ≥ 0.81),表明领域特定得分具有高度冗余性。此外,基于测验信息函数的分析揭示,该基准的测量精度集中于中等能力水平,在θ > 0的前沿模型所在区域急剧下降。因此,研究结论指出,HLE的领域子分数不宜被解释为独立能力表现,且该基准在区分顶尖模型方面存在显著局限性。
链接: https://arxiv.org/abs/2607.27420
作者: Mayank Sharma,Savira Nadela,Tyler Matteson
机构: Stanford University (斯坦福大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Humanity’s Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE ( J = 428 items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald’s \omega_h = 0.998 , domain labels explain only 3.5% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen’s d = 0.016 ), and domain-specific ability estimates are near-redundant with the total score ( r \geq 0.81 ). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above \theta = 0 , where frontier models sit. These findings suggest that HLE’s domain subscores do not warrant distinct capability interpretations and that the benchmark’s ability to discriminate among the strongest models is limited.
[NLP-76] Benchmarking LLM Competence on Logical Inference over Probability Operators
【速读】: 该论文旨在解决自然语言中不确定性表达与推理之间的逻辑一致性问题,尤其关注在包含可变认知模态词(如“probably”“might”“must”)的句子上进行有效推理的挑战。这类推理对于医疗、法律等高风险领域至关重要。其核心解决方案是构建一个针对概率算子推理的基准测试(benchmark),涵盖14,320个程序生成的英文提示,覆盖十五种推理模板,系统性地变化问题形式、否定策略和表面内容。研究发现,多数大语言模型表现出与逻辑形式无关的答案偏差(如对“是”或“否”的系统性偏好),并提出“能力下限”(competence floor)作为衡量模型真实推理能力的指标——即模型在“是正确”和“否正确”样本上的最低准确率。结果显示,仅有9/29的模型表现优于随机水平,且在问题形式、动词短语、姓名性别与来源等多个维度均存在显著偏见,揭示了当前模型在脱离表面模式匹配实现符号化、原则性推理方面仍存在根本性缺陷。
链接: https://arxiv.org/abs/2607.27405
作者: Nayera Hasan,Jack Greff,Alvin Grissom II
机构: Haverford College; Haverford College; Haverford College
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Under review
Abstract:Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators–inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model’s accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
[NLP-77] AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
【速读】: 该论文旨在解决阿拉伯语语境下多模态仇恨表情包(hateful memes)的检测问题,尤其针对现有研究在阿拉伯语资源中缺乏细粒度、多标签标注的现状。当前多数阿拉伯语表情包数据集仅聚焦于宣传内容或粗粒度有害内容分类,无法有效捕捉仇恨表达中的文化隐喻、文本与图像协同语义及攻击策略等复杂特征。为此,论文提出AHA-Memes——首个大规模、细粒度标注的阿拉伯语仇恨表情包基准数据集,包含5,000条人工标注的表情包,采用涵盖多种仇恨类型及其攻击策略的分类体系,并额外提供约6.6万条银标签(silver-labeled)数据以支持后续研究。解决方案的关键在于构建具有文化敏感性与多模态对齐能力的标注体系,同时在多种模型架构(包括仅文本、仅图像、晚期融合的多模态模型,以及零样本与微调设置下的少样本上下文学习、开放权重与闭源视觉-语言模型)上建立全面基准,揭示了阿拉伯语仇恨表情包检测中由文化背景驱动的深层挑战,为未来跨模态、跨语言的仇恨内容识别研究提供了重要基础。
链接: https://arxiv.org/abs/2607.27393
作者: Mohamed Bayan Kmainasi,Ali Ezzat Shahroor,Abul Hasnat,Md. Rafiul Biswas,Wajdi Zaghouani,Firoj Alam
机构: Qatar Computing Research Institute, Qatar; APAVI.AI, France; Hamad Bin Khalifa University, Qatar; Northwestern University in Qatar, Qatar
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 26 pages, 14 figures, 15 tables
Abstract:Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels. We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations. The dataset includes 5K manually annotated memes using a taxonomy that captures hate types, i.e., attack strategies. We further provide ~66K silver-labeled memes to support future studies. We benchmark text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning (ICL) and open- and closed-weight Vision-Language Models (VLMs) under zero-shot and fine-tuning settings. Our results establish strong baselines and highlight key challenges in culturally grounded Arabic hateful meme detection. We release the dataset, annotation guidelines, and evaluation scripts to support future research. WARNING: This paper contains examples that may be disturbing to readers.
[NLP-78] Same Facts Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
【速读】: 该论文旨在解决生成式 AI 在临床诊断推理中对社会语言风格(sociolinguistic register)的敏感性问题,即“叙事锚定”(Narrative Anchoring)现象:相同临床事实以不同语言风格表述时,模型输出的诊断结果出现显著偏差。其核心挑战在于,这种偏差并非源于显式的身份标记(如种族、收入等),而是由语言表达方式的差异引发,且在无任何人口统计学特征的情况下依然存在。解决方案的关键在于提出一种名为 NarrativeShield 的三代理论框架,通过结构化地提取并验证临床事实,在诊断推理前消除语言风格干扰,从而将叙事锚定差距从0.064至0.151降低至-0.004至0.037,接近零水平,并显著减少严重决策不稳定性(DSS ≤ 0.8)。该方法在保持合理准确率的前提下,有效解耦语言风格与临床推理,为构建鲁棒、公平的临床辅助诊断系统提供了可推广的技术路径。
链接: https://arxiv.org/abs/2607.27384
作者: Prabhjot Singh,Pritam Deka,Vijay Chennareddy
机构: University of Texas at Austin (德克萨斯大学奥斯汀分校); Queen’s University Belfast (贝尔法斯特女王大学); Middlesex University (密德萨斯大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ( -0.004 to 0.037 ) and achieving the lowest rate of severely unstable decisions (DSS 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
[NLP-79] HSS-Synth: Humanities and Social Sciences Data Synthesis for LLM s ACL
【速读】: 该论文旨在解决人文与社会科学(HSS)领域高质量、多样化数据稀缺且成本高昂的问题,尤其针对现有数据合成方法在开放性任务中难以有效应用的挑战。传统数据合成方法多聚焦于能力导向、碎片化的方法,未能系统覆盖HSS领域的复杂性和多样性。为此,本文提出一种以主题为中心(subject-centric)的新范式,首次构建涵盖14个主流学科的HSS领域体系,并设计了首个专为HSS定制的数据合成流水线——HSS-Synth。其核心解决方案包括:(1)通过多阶段过滤与文本精炼从网络语料中构建种子文档,并由评判模型评估质量;(2)采用“需求+角色”设定进行逆向翻译,生成兼具多样性与忠实性的指令,并通过严格的问答对齐检查确保一致性;(3)引入教师强制回答(teacher-forced Answering)机制,在生成过程中注入种子文档以锚定语义、抑制幻觉并保持语气与完整性,突破大语言模型(LLM)的响应长度限制。实验表明,HSS-Synth生成的23.7万条高质量指令微调样本在16项基准测试中超越14个主流基线模型,基于Qwen3-8B-Base微调后的模型达到新SOTA水平,接近官方Qwen3-8B性能,同时在人类偏好和知识能力上均实现提升,无性能退化现象。大量实验证明该方法具备强鲁棒性与良好的可迁移性。
链接: https://arxiv.org/abs/2607.27379
作者: Ru Peng,Tianyu Zhao,Xijun Gu,Zhiting Fan,Haokai Xu,Jinyang Zhang,Yawen Zeng,Yihong Zhuang,Kexin Yang,Junyang Lin,Dayiheng Liu,Junbo Zhao
机构: Zhejiang University (浙江大学); Inclusion AI, Ant Group (蚂蚁集团); Peking University (北京大学); Qwen Team, Alibaba Group (阿里巴巴集团)
类目: Computation and Language (cs.CL)
备注: ACL Findings 2026 Paper
Abstract:High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying “requirements + persona” to backtranslate seed documents into diverse yet faithful instructions with a strict QA alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth’s robustness and transferability. Our code is publicly available at this https URL.
[NLP-80] Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
【速读】: 该论文旨在解决生成式建模中长期存在的非端到端训练问题,即尽管生成模型具备强大能力,但其训练过程仍依赖于分解式的生成流程,无法实现真正的端到端优化。其核心挑战在于生成任务涉及多模态分布,现有可扩展方法通过因子分解生成过程来应对,但这限制了模型对模式的精准捕捉。本文提出“探索性建模(Explorative Modeling, XM)”这一新范式,关键在于将因子分解从生成过程转向训练循环:在每轮训练中探索K个候选生成结果与真实数据的匹配,并仅基于最优匹配进行参数更新,从而促使模型聚焦于特定模式而非模糊平均。该方法不仅为现有生成模型引入了“探索度”作为继参数量和数据规模之后的第三个可扩展维度,且在连续与离散域(图像、视频、语言)上均表现出随规模增长而单调提升的性能增益;在大规模场景下,探索带来的效率提升显著——计算效率提高4.1倍、采样效率提升6.2倍、参数效率提升47%,并在无引导条件下使图像生成达到接近前沿的1.43 FID(ImageNet)。此外,XM实现了真正意义上的端到端重建生成建模,在控制任务中以16–256倍更少的推理步数媲美扩散模型,验证了其作为现有模型预训练新轴线及独立端到端生成范式双重价值。
链接: https://arxiv.org/abs/2607.27372
作者: Alexi Gladstone,Heng Ji,Yilun Du
机构: UIUC (伊利诺伊大学厄本那-香槟分校); Harvard (哈佛大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being remarkably capable, they are still not trained end-to-end. This is because, at its core, generative modeling is about handling distributions with many modes, and existing scalable approaches handle this the same way, by factoring the generation procedure, which prevents end-to-end generation. In this work, we introduce Explorative Modeling, a new paradigm that instead factors the training loop, exploring K candidate matches between model generations and data, and training on the best, so predictions commit to modes rather than blurring them. We find Explorative Models (XMs) useful in two settings. First, increasing exploration adds a third pretraining axis beyond parameters and data for existing generative models-where scaling exploration monotonically improves performance across both continuous and discrete domains (images, video, and language). Notably, gains from exploration increase with scale, climbing from 7% to 36% as data scales and from 13% to 23% as models grow, with efficiency gains more than doubling at 3x the compute. Concretely, exploration improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, parameter efficiency by 47%, lifts the strongest of image-generation recipes to a near-state-of-the-art 1.43 FID on ImageNet without guidance, enables scaling how end-to-end existing models are, and unlocks scaling generalization. Second, XMs enable end-to-end reconstructive generative modeling, matching diffusion on control tasks with 16-256x fewer inference steps. Together, these results establish XMs as both a new pretraining axis for existing generative models and a standalone end-to-end generative modeling paradigm.
[NLP-81] BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在人文与社会科学(Humanities and Social Sciences, HSS)领域中数据合成缺乏针对性的问题,尤其针对开放性任务中难以量化客观正确性、而更依赖于细微质量判断的特性。现有偏好对齐方法要么成本过高,要么不适用于广泛多样的HSS学科。为此,论文提出BridgeAlign,作为首个面向广义HSS领域的偏好对齐框架,其核心创新在于三阶段设计:首先通过启发式与生成式大模型(Generative AI)结合的过滤及文本优化策略,精选高质量的HSS种子文档;其次利用基于角色的指令逆向生成偏好三元组,并引入问答一致性校验以保障质量;最后突破传统“人类-模型”粗粒度对比的局限,将偏好判断锚定于HSS领域质量评估量规(quality rubric),并通过受控的质量退化生成近边界偏好对,实现对细微质量差异的精细化区分。该方案通过构建并对超过21万条合成偏好样本进行对齐,使Qwen3-8B在17个基准测试上平均表现优于11个主流基线模型,且在人类偏好与知识能力之间实现了无权衡的协同提升,验证了其有效性与理论适配性。
链接: https://arxiv.org/abs/2607.27366
作者: Ru Peng,Haokai Xu,Xijun Gu,Tianyu Zhao,Zhiting Fan,Yawen Zeng,Yihong Zhuang,Jinyang Zhang,Kexin Yang,Jian Wu,Hao Chen,Junyang Lin,Dayiheng Liu,Junbo Zhao
机构: Alibaba Group (阿里巴巴); Ant Group (蚂蚁集团); Qwen Team (通义千问团队)
类目: Computation and Language (cs.CL)
备注:
Abstract:While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than objective correctness. This makes preference alignment a natural paradigm for broad HSS tasks. Yet existing methods are either costly or not tailored to broad HSS disciplines. We thus propose BridgeAlign, among the first preference-alignment pipelines for broad HSS disciplines, with three phases: i) Seed Curation: curating HSS seed documents from web corpora via heuristic/LLM-based filtering and text refinement; ii) Preference Data Synthesis: generating preference triplets via persona-based instruction inversion with QA consistency checks; iii) Preference Optimization: moving beyond naive human-vs-model heuristics by first grounding preferences in HSS quality rubric, then generating transitional responses via controlled quality degradation to form near-boundary preference pairs for finer-grained quality discrimination. Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines; importantly, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them, as supported by extensive experiments and contextualized by existing theories.
[NLP-82] LayerRAG -Bench: A Cross-Layer Reliability Benchmark for Agent ic Retrieval-Augmented Generation
【速读】: 该论文旨在解决生成式AI系统在多层可靠性(如证据层、工具调用层、授权层及会话状态层)中出现的“表面可信但实际失效”问题,即系统生成的答案看似基于证据,实则在关键底层环节存在错误。其核心解决方案是提出一个名为LayerRAG-Bench的受控跨层可靠性基准测试体系,涵盖8个企业级领域、240项任务、9种故障场景、2种合约模式以及来自OpenAI、Anthropic和Gemini等九个模型的38,880条实时任务记录。研究发现,模式归一化(schema normalization)可显著提升模式漂移(schema-drift)场景下的成功率(从0.000提升至0.913),但对过时证据、工具输出缺失、权限拒绝及错误会话上下文等问题无法修复;同时,仅依赖“地面性”(groundedness)评估会产生大量假阳性结果,尤其在面对过时或错误会话证据时。因此,该研究强调应遵循“分层评估原则”:任何可靠性干预措施只能被认定为对其目标层级的有效修复,不可误认为是通用万能解法。
链接: https://arxiv.org/abs/2607.27353
作者: Musa Shams(Independent Researcher)
机构: Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL)
备注: 10 pages, 9 tables. Code and data: this https URL
Abstract:Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are not recovered by schema normalization. Groundedness-only evaluation also produces substantial false positives under stale and wrong-session evidence. These results support a layer-specific evaluation principle: a reliability intervention should be credited for repairing its target layer without being mistaken for a universal fix.
[NLP-83] HGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
【速读】: 该论文旨在解决动态异构图(Temporal Heterogeneous Graphs)中跨类型结构异质性与关系交互时序动态性联合建模的挑战,现有方法在实现参数高效的跨类型迁移与关系感知的专属性化之间存在矛盾,且通常将时间信息以附加特征形式引入注意力机制之外,难以有效捕捉相对时间对关系演化的影响。其解决方案的关键在于提出一种统一的双路径架构——THGFM(Temporal Heterogeneous Graph Fusion Model),通过并行设计“共享空间时序注意力”分支实现参数高效跨类型迁移,以及“关系类型分片时序注意力”分支实现关系感知的专属性建模,并采用“双路径关系-共享融合”机制,具体以“类型条件非竞争门控求和融合”实现对共享与专用分支的自适应加权,避免零和竞争;同时引入“旋转时序注意力”(Rotary Temporal Attention),在注意力计算前对查询与键进行相对时间半相位旋转,从而直接将相对时间信息嵌入注意力评分过程。该框架在多个学术图基准上显著优于基线图变换器模型,六任务平均提升达+3.25%,在部分任务上达到最高+12.37%的相对增益,验证了其在复杂动态异构图建模中的有效性。
链接: https://arxiv.org/abs/2607.27303
作者: Yixin Peng,Diego Collarana,Er Jin,Stefan Decker
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)
备注:
Abstract:Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve over time. Learning on such graphs requires jointly modeling cross-type structural heterogeneity and the temporal dynamics of interactions, yet existing methods still struggle to reconcile parameter-efficient cross-type transfer with relation-aware specialization, and typically inject time only as additive features outside the attention kernel. We propose \textbfTHGFM, a web-scale temporal heterogeneous graph fusion model that addresses both limitations within a unified dual-path architecture. THGFM couples a \textitShared-Space Temporal Attention branch for parameter-efficient cross-type transfer with a \textitRelational Type-Partitioned Temporal Attention branch for relation-aware specialization, and integrates them through \textitDual-Path Relational–Shared Fusion, instantiated with \textitType-Conditioned Non-Competitive Gated Sum Fusion: a adaptive mechanism that assigns independent, type-conditioned feature-wise gates to the shared and specialized branches, allowing both to be amplified or suppressed without zero-sum competition. To directly incorporate relative time into the attention score, THGFM further introduces \textitRotary Temporal Attention, which rotates queries and keys by half-phases of relative time before matching. THGFM consistently outperforms baseline graph transformer models on academic graphs benchmarks, delivering a +3.25% six-task mean gain, with peak relative gains of +12.37% on OAG-CS PV, +4.87% on PF- L_2 , and +1.18% on PF- L_1 , and +4.24% , +3.73% , and +4.61% on OGBN-MAG, HTAG-ArXiv, and HTAG-DBLP, respectively.
[NLP-84] Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
【速读】: 该论文旨在解决生成式 AI(Generative AI)在理解新闻文本中情感框架方面的能力问题,具体聚焦于大语言模型(LLM)是否能够准确捕捉人类对政治与地缘政治冲突相关标题所引发的情感共鸣。研究的关键在于通过一个大规模、人口统计学多样化的实证评估,系统比较7个主流大语言模型与英国代表性成人样本(n=3011)在判断新闻标题是否引发特定立场同情心方面的表现。结果显示,尽管领先模型整体上与人类情感判断高度一致(如GPT-5.2的相关系数达0.789),但不同模型间存在显著差异,且即使在总体对齐度较高的情况下,模型与人类判断的契合度仍随年龄、性别、教育水平、地缘政治知识及个体立场等群体特征呈现差异化表现。因此,该研究的核心贡献在于揭示了“非普遍性对齐”这一关键问题——即高均值性能掩盖了潜在的群体偏差,强调在开发伦理化、普适性的智能系统时,必须考虑并应对这种基于人口统计与文化背景的差异化对齐需求。
链接: https://arxiv.org/abs/2607.27232
作者: Haran Shani-Narkiss,Michael Fire,Oren Tsur
机构: University College London (伦敦大学学院); Ben Gurion University of the Negev (本古里安大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG)
备注:
Abstract:Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants’ predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs’ comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal – it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.
[NLP-85] AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
【速读】: 该论文旨在解决随着生成式AI的普及,学术会议面临投稿数量激增而审稿人力不足的挑战。其核心问题是:如何在不降低评审质量的前提下,提升同行评审的效率。解决方案的关键在于构建一个由人工智能驱动的预审系统——bosc-pre-review,该系统基于已有的详细评分标准(rubric),对提交的摘要进行自动化评估,涵盖开放性(openness)、合规开源许可证(valid open source license)以及可运行性(runnability)等六项关键指标;同时引入Runabilly模块,在隔离的Docker容器中自动构建并测试项目,确保安全性与可复现性。整个过程中,AI仅负责收集和呈现证据,所有最终接受决定均由人类审稿人独立做出。调研结果显示,多数审稿人认为该预审工具具有实用价值,但更倾向于自主验证AI结论,而非完全依赖其判断,体现了人机协同审稿中“增强而非替代”的设计原则。
链接: https://arxiv.org/abs/2607.27228
作者: Tazro Ohta,Nomi L. Harris,Seth Carbon
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
备注: 18 pages, 4 figures
Abstract:Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences are seeing an overwhelming surge of submissions. We wanted to see if generative AI could help our conference’s volunteer reviewers by pre-reviewing abstracts for certain criteria. The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate submitted abstracts on multiple criteria, including openness (public availability of the code or other content associated with the project), valid open source license, and “runnability” (how easy it is to download, build, and run the project - an important measure of reusability). For BOSC 2026, we built bosc-pre-review, an agentic skill that assessed six review criteria, and Runabilly, which builds and tests each project in a disposable Docker container for safety. The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts. After the review period, we surveyed the reviewers to determine how useful they found the pre-review. Most of those who responded said they found it useful, but they preferred to check the AI’s conclusions against their own, rather than accepting the AI results unquestioningly.
[NLP-86] Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
【速读】: 该论文旨在解决学术文献爆炸式增长背景下,如何实现高效、可靠的信息综合问题。传统单次提示(single-shot prompting)方法在处理复杂合成任务时往往表现出可靠性不足与生成质量不稳定的问题。其解决方案的关键在于提出并实证验证了一种多阶段提示链(prompt chaining)的方法论,通过将复杂的生成任务分解为一系列有序的子任务,并以链式结构串联各阶段提示,从而提升整体系统的鲁棒性与输出一致性。实验结果表明,相较于经过精心优化的单次提示基线,该方法在“教育”领域任务中实现了100%的成功率,而基线失败率达50%,同时在质量指标上也显著优于基线(ROUGE-L F1-score:0.507 vs. 0.486),主要得益于更高的精确率。研究结论指出,提示链是一种更可靠且高效的工程范式,能够有效缓解单体提示固有的失败风险与不一致性问题,适用于复杂、多步骤的生成任务。
链接: https://arxiv.org/abs/2607.27210
作者: Andrei Lazarev
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Digital Libraries (cs.DL); Software Engineering (cs.SE)
备注: This is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: this https URL
Abstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for such tasks. This approach is implemented in our system, AI SciBrief, which automatically generates scholarly digests. We conducted a comparative experiment, measuring the performance of our prompt chaining method against a carefully optimized single-shot baseline. Both systems were evaluated against a human-authored “gold standard” report for the “Education” domain. The results demonstrate a significant difference in reliability: our prompt chaining method achieved a 100% success rate, whereas the optimized baseline failed in 50% of its runs. In terms of quality, the proposed method also demonstrated a clear advantage, achieving a superior ROUGE-L F1-score (0.507 vs. 0.486), driven primarily by higher precision. We conclude that prompt chaining is a more dependable and effective engineering approach for complex, multi-step generative tasks, significantly mitigating the risks of failure and inconsistency inherent in monolithic prompts.
信息检索
[IR-0] AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
链接: https://arxiv.org/abs/2607.28618
作者: Bing Yan,Gregory Wolfe,Stefano Martiniani,Kyunghyun Cho
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemble cross-paper answers manually. We present AskChem, a claim-centered infrastructure for cross-paper chemistry search. AskChem changes the unit of retrieval from the paper to the provenance-carrying claim: each paper is converted into atomic, typed claims, each grounded by a source DOI and a verbatim quote or an explicit evidence locator. Over this shared claim store, AskChem exposes complementary structures for search and synthesis: a stabilized faceted taxonomy for hierarchical retrieval and browsing, an evidence graph linking claims through relations, and an exploratory living taxonomy that situates indexed papers under scientific principles. AskChem currently indexes 2.4M claims from 147K papers and provides a web interface, as well as REST, SDK, and MCP access for AI agents. On AskChem-Bench, grounding a GPT-5.5 reader in AskChem yields 100% resolvable DOIs, compared with 88.3% without retrieval, and the highest citation density among five tested systems. AskChem is live at this https URL.
[IR-1] Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently
链接: https://arxiv.org/abs/2607.28571
作者: Simon Roy,Mark Bong,Giovanni Beltrame
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注: 10 pages, 3 figures
Abstract:Operational Earth observation increasingly calls for answering queries such as find the image pairs where a new building appeared.'' This means searching an archive of before-and-after (bi-temporal) satellite image pairs and ranking each pair by how well it matches a natural-language description of the change. The component that performs this match, the fusion module that combines the before’’ and ``after’’ views, must be run at query time across many candidate pairs, so its speed largely sets the cost of every search. We present a controlled comparison of how to build that module. Using one fixed image encoder (a frozen CLIP model) and one training recipe for all variants, we evaluate eight designs drawn from three families: attention, state-space models (Mamba), and learned compression (our Temporal Bottleneck Fusion, TBF). Each design is tested on two benchmarks (LEVIR-CC and Dubai-CC) with ten random seeds, so the reported differences are statistically grounded. We outline three findings: first, a training-free two-stage search (a cheap difference model that shortlists candidates, followed by attention fusion that re-ranks them) matches or exceeds full-fusion recall on LEVIR-CC while cutting query cost 10 - 15\times , with comparable R@1/R@5 on Dubai-CC; second, the linear-time scan of Mamba, attractive on paper, gives no speed benefit at the patch counts typical of vision transformers ( L=196 ): the scan is limited by memory bandwidth, whereas attention maps cleanly onto parallel hardware; and third, compressing the fused representation (TBF) reduces parameters by 2.3\times and latency by 1.6\times for a change-only BLEU-1 cost of 0.007 , although more aggressive compression quietly discards change-relevant detail that aggregate metrics fail to reveal.
[IR-2] CA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
链接: https://arxiv.org/abs/2607.28498
作者: Yuto Suzuki,Farnoush Banaei-Kashani
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Existing SIR methods rank papers by topical similarity and do not explicitly represent how a candidate inspiration transfers to a target problem. This is especially limiting for remote inspirations, whose value often lies in reusable problem-solving principles rather than topical overlap. Motivated by how humans abstract transferable aspects of a source and remap them to a new target, we reformulate SIR as target-conditioned abstraction (TCA). The retrieval object is a transferable abstract principle extracted from a candidate specifically for the target. We present TCA-SIR, which learns to generate target-conditioned abstractions and uses their representations to predict transferability. On ResearchBench, TCA-SIR outperforms prior SIR methods and direct LLM retrieval, improving HitRate@top4% over MOOSE-Chem by more than 10 percentage points. Learned abstractions also recover target-relevant mechanisms more clearly than an untrained TCA prompt, yielding both stronger retrieval and an interpretable rationale for scientific inspiration.
[IR-3] GLM-RAG : Graph Language Models for Graph-Based Retrieval-Augmented Generation
链接: https://arxiv.org/abs/2607.28397
作者: Maya Arseven,Anette Frank,Beni Egressy,Johann Higl,Moritz Plenz
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 10 pages, 19 figures
Abstract:Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model graph topology in multi-hop reasoning tasks. In parallel, graph language models (GLMs) have emerged as a promising paradigm that integrates graph reasoning and the semantic capabilities of language models. In this work, we introduce a GLM-based retriever and investigate the comparative strengths of GLM-based, GNN-based, and traditional vector-search-based retrievers in single- and multi-hop RAG settings, and with a particular focus on transferability to unseen domains. Our findings suggest that finetuned GLM retrievers generalize better out of domain, achieving SOTA on two multi-hop benchmarks. On in-domain multi-hop QA datasets they remain comparable to prior work, with promising scaling as parameters and subgraph coverage increase. GNN-based retrievers achieve higher graph coverage with an efficient training setup, whereas the vector-search baseline excels at single-hop datasets.
[IR-4] EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
链接: https://arxiv.org/abs/2607.28229
作者: Luigi Sigillo,Matteo Silvestri,Francesco Tabaro,Rajat Bhatnagar,Syed Irtaza Mubashar,Matt Jeffryes,Daljit Nijjer,Vittorio Perera,Ola Spjuth,Julio Saez-Rodriguez,Melissa Harrison,Fabio Petroni
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than 16 points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about 8 points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at this https URL
[IR-5] Extended Depth-First Representations of k2-trees
链接: https://arxiv.org/abs/2607.28136
作者: Gabriel Carmona,Paolo Ferragina,Giovanni Manzini,Francesco Tosoni
类目: Data Structures and Algorithms (cs.DS); Information Retrieval (cs.IR); Performance (cs.PF)
备注: 44 pages, 7 figures, 18 tables
Abstract:In this paper, we study static, computation-friendly, lossless compression formats for graphs, focusing on memory locality and operational efficiency of k^2 -trees. We observe that their traditional level-wise layouts suffer from poor cache performance due to weak locality, especially in operations such as matrix-vector and matrix-matrix operations. To address this limitation, we propose four depth-first representations of k^2 -trees: a plain depth-first layout (EDF-1), a balanced-parenthesis representation (BP), and their compressed variants (CEDF and CBP). We further introduce a linear-time compression method based on suffix and LCP arrays to identify and compress identical subtrees. We experimentally evaluate the execution time, the disk space, and the peak-memory usage of our approaches against classical level-wise k^2 -trees and DFUDS-based representations across two real and one synthetic dataset (i.e., Web Graphs, Wikidata, and random adjacency matrices) over the above linear-algebra operations. Results show that our depth-first layouts are competitive and often superior than known approaches: CEDF achieves the best compression in most settings, EDF-1 and CEDF reduce the peak memory usage consistently, and performance varies by workload, with different layouts excelling in different operations and data regimes. Overall, this work demonstrates that depth-first layouts of k^2 -trees provide a practical and efficient alternative to traditional layouts, improving both compression and computational performance in matrix operations. Comments: 44 pages, 7 figures, 18 tables Subjects: Data Structures and Algorithms (cs.DS); Information Retrieval (cs.IR); Performance (cs.PF) MSC classes: 68P05, 68P30, 68R10, 65F50 ACMclasses: E.1; E.2; E.4; F.2.2 Cite as: arXiv:2607.28136 [cs.DS] (or arXiv:2607.28136v1 [cs.DS] for this version) https://doi.org/10.48550/arXiv.2607.28136 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[IR-6] Face and Voice Cross-modal Association with Learning Convex Feature Embedding
链接: https://arxiv.org/abs/2607.28129
作者: Taewan Kim,Jiwoo Kang
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注:
Abstract:Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
[IR-7] CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent
链接: https://arxiv.org/abs/2607.28070
作者: Yunlong Wang,Huizhe Zhang,Haonan Hu,Yudong Li,Bing Wen,Jianchao Tu,Chengxiang Zhuo,Zang Li
类目: Information Retrieval (cs.IR)
备注:
Abstract:Recent studies in industrial recommendation systems have demonstrated that sequential recommendation models built upon self-attention can benefit from predictable scaling laws by increasing sequence length and model capacity. However, practical recommender systems impose strict latency and resource constraints, making it challenging to balance computational overhead with fine-grained feature interaction. In this paper, we propose CCFormer, an efficient Transformer backbone that unifies cross-field feature interaction and compressed long-sequence modeling for industrial recommendation. Specifically, CCFormer combines feature-field separated cross attention with long-sequence subspace token mixing to exploit long-term preference signals across heterogeneous feature domains. A hierarchical sequence compression strategy with progressively expanded receptive fields enables efficient long-sequence modeling with reduced information loss. Extensive experiments on two public benchmarks and a large-scale industrial dataset demonstrate that CCFormer consistently outperforms state-of-the-art baselines. Online A/B tests in a video recommendation scenario and an advertising ranking scenario at Tencent further validate its industrial practicality, yielding a 3.57% CTR gain and a 1.71% advertising revenue lift, respectively, while accelerating model training by 2.21x over the strong HSTU baseline. CCFormer has been fully deployed in Tencent’s production recommendation system, serving the main traffic of both scenarios.
[IR-8] VIG-RL: Learning to Search and Insert for Verified Image Grounding
链接: https://arxiv.org/abs/2607.28055
作者: Qinhan Yu,Jun Guang,Chong Chen,Wentao Zhang
类目: Information Retrieval (cs.IR)
备注:
Abstract:In knowledge-intensive scenarios, providing reliable interleaved text-image responses requires Verified Image Grounding (VIG): the precise integration of retrieved authentic visual evidence. Existing retrieval-augmented frameworks predominantly rely on decoupled, static pipelines, inherently failing to dynamically reason about when external knowledge is required and where visual assets should be contextually inserted. To bridge this gap, we propose VIG-RL, an autonomous agentic framework that formulates the search-selection-insertion workflow as an active decision-making process. Operating within a dynamic ReAct-style loop, VIG-RL is optimized via reinforcement learning, guided by a composite reward system that holistically evaluates the agent’s step-by-step tool execution and final multimodal alignment. Extensive evaluations demonstrate that VIG-RL establishes a new state-of-the-art, significantly outperforming existing static baselines.
[IR-9] FiRE: Enhancing MLLM s with Fine-Grained Context Learning for Complex Image Retrieval
链接: https://arxiv.org/abs/2607.27959
作者: Bohan Hou,Haoqiang Lin,Xuemeng Song,Haokun Wen,Meng Liu,Yupeng Hu,Xiangyu Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
备注:
Abstract:Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs’ retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model’s context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.
[IR-10] SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
链接: https://arxiv.org/abs/2607.27955
作者: Jennifer D’Souza,Sameer Sadruddin,Anisa Rula,Ana Bossler,Andrés Fullana,Enric Bas,Syed Ather,Defne Circi,Anlan Chen,L. Catherine Brinson,Alyssa Columbus,George Demetriou,Dongjun Jeong,Tarun Kumar,Frank Krüger,Sascha Genehr,Kai Budde-Sagert,Anamaria Leonescu,Francesco Lodola,Chiara Florindi,Gagana Balasubramanya Murthy,Samson Oluwapelumi Olagbile,Nazia Riasat,Yan Sha,Kevin Shen,Shaokai Yang
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注: 25 pages, 9 figures, Submitted for peer review to Nature Scientific Data
Abstract:Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of this http URL, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology Biotechnology, Materials Chemistry, Imaging Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.
[IR-11] Interpretable Representation via LLM -Driven Generative Disentanglement for Local-Life Service Recommendation
链接: https://arxiv.org/abs/2607.27944
作者: Long Zhang,Hao Jiang,Sheng Yu,Fei Pan,Peng Jiang,Kun Gai
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:
Abstract:While large language models (LLMs) have advanced ID-based recommendation through Semantic ID (SID) modeling, existing SID generation frameworks largely follow a single-representation-then-quantization paradigm. This design faces two bottlenecks: semantic entanglement mixes heterogeneous attributes, such as geography, brand, and category, causing information loss during quantization, low-quality SIDs, and severe collisions; moreover, black-box representation learning provides neither explicit attribute semantics nor clear geographic or semantic meanings for SID positions. These limitations weaken both retrieval reliability and the ability to diagnose or control SID generation. We propose Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation (LGRID). LGRID introduces a generative disentanglement paradigm through an Encode - Disentangle - Align - Quantize pipeline. It first uses joint LLM encoding to preserve cross-attribute geographic-semantic dependencies, rather than encoding fields independently. A Structured Disentangled Block then routes hidden states into attribute-aligned slots for geographic and semantic factors. Synergistic Alignment Learning makes these slots both generatively decodable and discriminative for retrieval, while Dual-Stream Residual Quantization separately discretizes the two streams into compact SIDs with explicit attribute correspondence. This design yields interpretable SIDs with positions grounded in item attributes and local-service semantics. Experiments on Kuaishou and Foursquare show that LGRID consistently outperforms strong SID baselines, achieving up to a 5.44 percent relative AUC gain. It also achieves over 99 percent attribute-decoding accuracy for coarse geographic fields and reduces the full-SID collision rate to 39.9 percent, compared with 97.0 percent for LGSID.
[IR-12] From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation
链接: https://arxiv.org/abs/2607.27789
作者: Zhi Chen,Minmao Wang,Xingchen Liu,Haoqiang Liang,Huihuang Lin,Likang Wu,Hongke Zhao,Yulong Wang,Shijie Yi,Fei Pan,Peng Jiang
类目: Information Retrieval (cs.IR)
备注:
Abstract:Semantic-ID-based generative recommenders enable efficient next-item generation, but their item-level supervision mainly captures behavioral co-occurrence and local transitions. Large language models (LLMs) can complement these models by reasoning over heterogeneous interaction histories to understand the user’s current demand. However, LLMs are not inherently trained with recommendation-specific outcome feedback, and linguistically plausible reasoning therefore does not necessarily lead to effective recommendation decisions. We term this mismatch the Understanding-Action Gap. Accordingly, we distinguish intent knowledge, which captures the user’s current demand, from policy knowledge, which specifies the recommendation direction and rejection boundary under that demand. To bridge this gap, we propose a feedback-driven agent framework that first induces task-oriented intent and then discovers recommendation policies according to their incremental utility over an intent-only baseline. Candidate policies are evaluated and refined using outcome-derived feedback rather than linguistic plausibility. We further transfer the resulting intent and policy knowledge into two latent tokens of a lightweight Semantic-ID generator through dual-space relational distillation, enabling LLM-free online inference. Experiments on public benchmarks show consistent improvements over baselines, while large-scale online A/B tests achieve gains of 4.506% in Revenue and 4.621% in ADVV.
[IR-13] Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
链接: https://arxiv.org/abs/2607.27766
作者: Xinyu Luo,Hui Liu,Yihua Shao,Junyi Yang,Arindam Basu,Haoliang Li
类目: Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Under review
Abstract:On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA’s rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.
[IR-14] DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
链接: https://arxiv.org/abs/2607.27763
作者: Bowen Wang,Youwen Zhang,Ritesh Mehta
类目: Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: 21 pages, 9 figures
Abstract:We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ‘‘Honest Threshold Tuning’’ procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary F_1 of 0.5790 and a secondary F_1 of 0.9657 . In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary F_1 of 0.5780 and a secondary F_1 of 0.9599 -essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall 0.3571 , ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging ( 0.3564 ), and a zero-shot MedGemma-4B run with a PubMed-style prompt ( 0.3186 ), spanning a wide range of model scales and training costs. Code: this https URL.
[IR-15] Hierarchical Latent Reasoning for LLM -based Recommendation
链接: https://arxiv.org/abs/2607.27760
作者: Peiyu Hu,Siying Gu,Weihai Lu,Zhuodong Liu,Yuntian Tang,Jiahao Liang,Yiying Xie,Jiang Rong,Zhaokai Luo,Zhiyong Wang,Jia Wang
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注:
Abstract:Large Language Models (LLMs) have shown strong potential for recommendation by leveraging their semantic understanding and contextual modeling capabilities. Recent studies further introduce reasoning mechanisms to improve user preference modeling. However, explicit natural-language reasoning incurs substantial inference overhead, whereas existing latent reasoning methods mainly focus on generating or verifying intermediate states, leaving their layer-wise preference roles and contributions insufficiently characterized. We propose HiLaR, a Hierarchical Latent Reasoning framework with layer-aware reinforcement optimization for LLM-based recommendation. HiLaR constructs temporal-guided hierarchical user preference representations, aligns them with multiple LLM latent reasoning states, and organizes the reasoning process from broad preferences to fine-grained current intents. To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state. Experiments on four Amazon benchmark datasets show that HiLaR generally outperforms strong sequential, generative, and LLM-based recommendation baselines. Ablation and sensitivity analyses further verify the contribution of hierarchical representation learning, latent alignment, and process-level optimization. Our code is available in this https URL.
[IR-16] A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery
链接: https://arxiv.org/abs/2607.27748
作者: Mengdi Chen,Yuanxin Huang,Yulin Jiang,Wei Sun
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Databases (cs.DB)
备注: 6 pages, 2 figures, 2 tables. Submitted to DAI 2026 Industry Track
Abstract:Enterprise data analytics agents face two structural failures: generic RAG retrieves the wrong asset (Hit@10=19.1%) and delivers no usage knowledge to prevent metric misinterpretation—stemming from four root causes (C1–C4) ranging from semantic gap and entity ambiguity to schema drift and asset-usage gap. We present a two-layer solution deployed in the commercial advertising data warehouse at Xiaohongshu (5,300+ Hive tables, 14 domains). A three-tier dual-purpose knowledge base (179 documents, eight-section annotation template) serves both retrieval and generation, with a closed-loop refresh pipeline maintaining day-level freshness (one yes/no approval, 30s hot-reload). The Graph-Guided Retriever (GGR) uses a 2,859-node knowledge graph as a candidate gate with intent routing to deliver 71.6x token reduction. The Scene-Aware Ranker (SAR) applies 19-class entity recognition and explicit scenario annotations; negative knowledge alone contributes 25 percentage points of Hit@10 gain. On two 100-question benchmarks, Hit@10 rises from 19.1% to 96.6% (+77.5pp) and knowledge coverage from 56% to 77%, at 4.84–5.33s end-to-end latency.
[IR-17] ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
链接: https://arxiv.org/abs/2607.27744
作者: Yuxin Chen,Liang Luo,Buyun Zhang,Jian Jiao,Boda Li,Haoyu Wang,Tongyi Tang,Ao Cai,Zijian Shen,Zhengkai Zhang,Wenyi Xie,Ryan Dick,Han Liu,Neng Shi,Bin Yu,Jianbo Xiao,Shuyao Bi,Hongtao Yu,Yuanwei Fang,Zhuoran Zhao,Sijia Chen,Yang Chen,Shuqi Yang,Qianru Li,Zikun Liu,Wei Ling,Sihan Zeng,Longhao Jin,Jiaxin Lu,Yinbin Ma,Jiawei Li,Yichen Ruan,Yong Ler Lee,Birmingham Guan,Zijian Li,Jianbo Sun,Zhengyu Zhang,Zeliang Chen,Xiaohan Wei,Yuchen Hao,GP Musumeci,Venkatesh Ranganathan,Yantao Yao,Chunqiang Tang,Wenlin Chen,Santanu Kolay,Ellie Dingqiao Wen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
备注:
Abstract:Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) Cite as: arXiv:2607.27744 [cs.LG] (or arXiv:2607.27744v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27744 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[IR-18] Measuring Alignment With Reader Highlights Net of Position and Length
链接: https://arxiv.org/abs/2607.27739
作者: Kazuki Nakayashiki,Keisuke Watanabe
类目: Information Retrieval (cs.IR); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
备注: 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper’s numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A
Abstract:Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting offers a non-circular reference: many people independently marking passages on the same page. But the obvious metric, the fraction of crowd-marked sentences a compressor keeps, is confounded twice: crowd marks are front-loaded and crowd-marked sentences are longer, so any method favouring early or long sentences scores well regardless of readers. We remove both by matching each marked sentence against unmarked sentences of the same document at equal relative depth and equal within-document length rank, and we calibrate every estimator on synthetic nulls built from position and length alone - a step that matters, since depth-only stratification returns a false positive on 20-36% of nulls containing no effect. On 120 web documents (at least 12 independent readers each), a language-model importance ranking keeps 38.4% of crowd-marked sentences against 19.9% of their matched neighbours: an enrichment of +0.196 [+0.148, +0.239], at p = 0.0005 under an exact randomization test that assumes nothing about clustering, and replicated cross-vendor. Naive truncation, whose keep rule is position, correctly falls to +0.003. To give the number a scale: scored identically, on the same budget, against a crowd label recomputed to exclude them, a single human reader reaches +0.182 - indistinguishable from GPT-5.4 (+0.002 [-0.081, +0.088]) and below Claude Opus 5. Classical methods are not null - Luhn’s 1958 heuristic reaches +0.088 - so reader selection is partly recoverable by counting words; conditioning additionally on lexical centrality removes only 0.010, so the agreement is not centrality. We also report that a claim in our own prior work does not reproduce on this corpus.
[IR-19] Restoring Collaborative Signals in Semantic-ID Generative Recommendation via Personalized Natural Language
链接: https://arxiv.org/abs/2607.27682
作者: Changjiang Han,Qingyang Li,Yaqiang Zang,Jikun Kang,Pinghua Gong,Xue Liu,Bowei He
类目: Information Retrieval (cs.IR)
备注: 8 pages, 4 figures
Abstract:Making LLM-based generative recommendation models stronger and more personalized through natural language and explicit reasoning is a widely anticipated yet still unsolved goal. Such models cast recommendation as autoregressively generating an item’s semantic-ID (SID), a short tuple of discrete codes, so that recommending well reduces to emitting the right SID. In this setting the model verbalizes its knowledge poorly, and text and SID tokens live in misaligned embedding spaces. Deep reasoning therefore rarely turns into a correct SID, and enabling explicit “thinking” often gives no gain or even hurts. The deeper cause is that a compact SID cannot hold content and collaborative signal at once: the two compete, and collaboration loses. Because a mis-predicted SID is a wrong recommendation, this caps accuracy directly. Costly multi-round training barely helps, and few methods try to supply the missing signal at inference time. What is missing is a reliable channel that carries collaborative signal into SID generation. We therefore propose a framework, guided by personalized natural language, that adds hierarchical collaborative cues as the model generates, without altering the backbone or retraining the SIDs. Rather than mapping language onto SIDs directly, it uses language to attach analyzable links between collaborative patterns and their audiences, restoring the collaborative signal that SIDs miss. The result is consistent gains in recommendation accuracy, grounding generation in collaborative structure at inference time rather than relying on explicit reasoning or retraining.
[IR-20] LoopMemGR: From Behavior Logs to Evolving Memory for Generative Recommendation
链接: https://arxiv.org/abs/2607.27647
作者: Hui Qian,Changfa Wu,Chang Liu,Binbin Cao,Jian Wu,Yuliang Yan,Han Zhu,Bo Zheng
类目: Information Retrieval (cs.IR)
备注:
Abstract:Generative recommendation formulates next-item prediction as conditional autoregressive generation over discrete Semantic IDs, enabling end-to-end recommendation over large-scale item spaces. However, most existing methods follow a history-as-context paradigm that repeatedly reconstructs user preference from behavior history while discarding system-side recommendation decisions after each request. This creates an asymmetric memory: the system remembers what the user has done, but not what it has previously recommended or learned from the resulting feedback. Consequently, useful preference-validation signals, potential negative evidence, and historical exploration information cannot be directly reused across requests. To address these limitations, we propose LoopMemGR, a closed-loop recommendation experience memory framework for generative recommendation. In addition to the conventional behavior log, LoopMemGR maintains a recommendation experience log that records past recommendation–feedback trajectories. It extracts request-relevant evidence through three complementary views: the recency view captures short-term interaction dynamics, the frequency view summarizes recurring recommendation patterns, and the global view distills transferable regularities shared across users. These signals are compressed into a fixed number of experience tokens to condition the generative backbone under a bounded input budget. Extensive experiments on an industrial Taobao dataset demonstrate the effectiveness of closed-loop experience accumulation and multi-view experience extraction.
[IR-21] Dynamic Exploration Graph: A Novel Approach for Efficient Nearest Neighbor Search in Evolving Multimedia Datasets
链接: https://arxiv.org/abs/2607.27640
作者: Nico Hezel,Kai Uwe Barthel,Bruno Schilling,Konstantin Schall,Klaus Jung
类目: Information Retrieval (cs.IR)
备注:
Abstract:Approximate Nearest Neighbor Search (ANNS) represents a fundamental problem in various applications (image-search, recommendation systems). While graph-based algorithms have demonstrated a good balance between search accuracy and time, handling dynamic datasets, where data points are continuously added or removed, remains a challenge. This paper introduces the Dynamic Exploration Graph (DEG), an extension of the continuous refining Exploration Graph, which retains high search efficiency for static dataset while adding essential support for dynamic data. At the core of the DEG design are two key innovations: a novel vertex deletion algorithm which guarantees graph connectivity and a data distribution-agnostic method for graph expansion. Through these mechanisms, the DEG maintains a balanced and well-connected structure, even under continuous data alterations. Empirical experiments in both streaming and online scenarios demonstrate the superior performance of the DEG, surpassing existing dynamic graph algorithms in terms of construction time and search efficiency. Although optimized for dynamic datasets, the DEG delivers results as good as current state-of-the-art approaches for static dataset, underscoring its broad applicability.
[IR-22] An Exploration Graph with Continuous Refinement for Efficient Multimedia Retrieval
链接: https://arxiv.org/abs/2607.27623
作者: Nico Hezel,Kai Uwe Barthel,Konstantin Schall,Klaus Jung
类目: Information Retrieval (cs.IR)
备注:
Abstract:As datasets and the dimensionality of feature vectors continue to grow, Approximate Nearest Neighbor Search (ANNS) in large multimedia databases becomes increasingly relevant. Graph-based approaches have demonstrated to offer the best trade-off between retrieval precision and search time. Despite their ability to deliver search times several orders of magnitude faster than exact search techniques, existing methods suffer from slow constructions speeds or high memory requirements. This paper presents a “continuous refining Exploration Graph” (crEG), a novel approach for rapidly constructing a compact exploration graph with state-of-the-art search performance. Additionally, it provides the ability to enhance its effectiveness even further through an optional edge optimization algorithm. Both algorithms are specifically designed to produce and operate on undirected graphs with even degrees and guarantee graph connectivity at any time, a property particularly valuable for “exploratory search”, where the query is part of the database elements. Although such queries provide an advantageous starting point for graph search algorithms, they have been rarely considered in the context of ANNS, yet are crucial for recommendation and exploration systems. Our experiments demonstrate high efficiency in ANNS does not necessarily translate to a good performance in “exploratory search”.
[IR-23] Heterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study RECSYS
链接: https://arxiv.org/abs/2607.27577
作者: Di Bai,Jintao Liu,Zhenwei Tang,Peifan Wu,Nada Al-Thawr,Luoshu Wang
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注: Accepted to ACM RecSys Industry Track 2026
Abstract:Heterogeneous recommendation feeds present complex challenges that extend beyond those found in highly homogeneous environments (e.g., music-only or video-only closed-ecosystem platforms). In Google Discover, a unified feed integrates diverse content sourced from the decentralized open web, including web articles, long-form and short-form videos, user-generated content (UGC), and beyond. Different content types exhibit distinct feature densities and user interaction patterns. Building a unified ranking model that sustains high performance across such heterogeneity, while avoiding negative transfer or majority bias, remains a significant industrial challenge. This paper presents an end-to-end case study on the industrial-scale multi-task ranking of heterogeneous feeds, grounded in real-world deployment. We introduce HA-MoE, a heterogeneity-adaptive multi-gated mixture-of-experts architecture that incorporates explicit heterogeneity context into both gating networks and expert representations. This approach enables effective specialization without significantly increasing operational overhead. To support reliable deployment, we introduce LENS, a lightweight observability framework that provides interpretable diagnostics of expert specialization and tracks this functional heterogeneity across continuous retraining. We evaluate our method using Dual-Level AUC (DL-AUC), a heterogeneity-aware evaluation metric that combines global ranking performance with cross-segment ranking correctness. Offline evaluations on a large-scale industrial dataset demonstrate consistent improvements over baseline models. Furthermore, online A/B testing confirms gains in feed activity and exploration metrics. Together, offline and online results validate the effectiveness of our approach for managing heterogeneity in industrial-scale recommender systems. Comments: Accepted to ACM RecSys Industry Track 2026 Subjects: Information Retrieval (cs.IR); Machine Learning (cs.LG) Cite as: arXiv:2607.27577 [cs.IR] (or arXiv:2607.27577v1 [cs.IR] for this version) https://doi.org/10.48550/arXiv.2607.27577 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Journalreference: Proceedings of ACM RecSys 2026 Related DOI: https://doi.org/10.1145/3773078.3831848 Focus to learn more DOI(s) linking to related resources
[IR-24] Hierarchical Reranking for Scalable Financial RAG System IJCAI ECAI2026
链接: https://arxiv.org/abs/2607.27523
作者: Joohyun Lee,Sungwoo Hong
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
备注: 8 pages, 1 figure, 3 tables. Accepted at FinLLM @ IJCAI-ECAI 2026
Abstract:Analyzing financial documents such as 10-K filings, tabular disclosures, and macroeconomic reports demands expert reasoning and extensive time. However, existing Retrieval-Augmented Generation systems often struggle to process hybrid text-table structures or the massive scale of financial documents. To address these challenges, we propose Hierarchical Reranker, a RAG framework designed to improve retrieval performance and generative reliability across large-scale financial datasets. The system integrates three key innovations: Pre-Retrieval Optimization, enhancing query clarity and search efficiency through normalization, keyword expansion, and table transformation; Hierarchical Reranker Architecture, improving retrieval precision through a two-stage ranking mechanism; and Long-Context Management, preserving reasoning accuracy through adaptive input partitioning and fusion under extensive contexts. Across multiple benchmarks, including FinQA, FinanceBench, and ConvFinQA, the proposed system achieved an NDCG@20 score of 0.7918 and demonstrated superior factual consistency. Its robustness was further validated by achieving second place in the ACM-ICAIF '24 FinanceRAG Challenge. This work presents a deployable, domain-optimized RAG pipeline that enhances both the accuracy and scalability of financial reasoning, paving the way for automated audit reporting and quantitative investment analysis. The source code will be made publicly available on GitHub upon acceptance.
[IR-25] OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
链接: https://arxiv.org/abs/2607.27475
作者: Ziwei Li,Shuyao Li,Xufeng Cai,Xue Zou,Yiming Ma,Huiting Lu,Wujie Yan,Zhichen Zhao,Yang Lu,Zhe Wang,Rui Luo,Zhengyu Su,Dan Zhang,Ji Liu
类目: Information Retrieval (cs.IR); Machine Learning (cs.LG)
备注:
Abstract:In modern recommendation systems, retrieval serves as a primary stage responsible for filtering billions of candidate items down to thousands prior to refined ranking. To make this massive search effective and efficient, the system relies on ranking accuracy and indexing efficiency. However, these two objectives are traditionally misaligned: while the former optimizes for the alignment between ranking predictions and user behavior, the latter optimizes for a structural grouping of item representations which enables fast search among billions of candidates. Thus, despite extensive efforts to scale up interaction modeling for retrieval, they remain fundamentally limited by the structural misalignment between the ranking objectives and the proximity-learned index. In this work, we address this long-standing dichotomy by proposing a new holistic retrieval framework, OneShot. It is an end-to-end, in-model index learning framework that natively aligns index learning with ranking objectives. Using this joint learning as a structural foundation, OneShot pushes the boundaries of retrieval expressiveness by scaling interaction modeling with neural scoring beyond the persistent dot-product bottleneck. OneShot is fully deployed in Instagram’s industrial short-video recommendation system, driving significant wins in user daily sessions, engagement, and time-spent. Additionally, OneShot achieves a 20% recall gain at the operational ranking volume and a 10x efficiency improvement at an equivalent recall level.
[IR-26] MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
链接: https://arxiv.org/abs/2607.17751
作者: HONOR Agentic Search Team:Zhengzong Chen,Lei Tang,Lijun Liu,Chuandi Jiang,Fan Yang,Keyun Chu,Chu Zhao,Shihao Liu,Minghang Li,Bo Liang,Can Wen,Hailong Wu,Jingnan Ju,Mian Liu,Nengbin Zhang,Peiqiang Wang,Penghe Nie,Qinhui Gu,Sijia Lv,Siqi Chen,Wei Zhang,Yang Xu,Yuhao Qian,Yuxiang Zhang,Zeng Cheng,Zhen Wang,Zuan Chen,Yuanyuan Zhao,Fei Huang
类目: Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.
人机交互
[HC-0] AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
链接: https://arxiv.org/abs/2607.28617
作者: Xiangning Lin,Shenzhe Zhu,Shu Yang,Zhenyu Zhang,Haoqian Zhang,Yipeng Zhao,Chengxuan Qian,Tianwei Wang,Ziheng Zhang,Zhenlong Yuan,Dingcheng Wang,Juncheng Wu,Yuan Si,Jiaxin Liu,Baolong Bi,Robert Mahari,Tobin South,Dazza Greenwood,Zexue He,Rishi Bommasani,Sophia Kazinnik,Andreas Haupt,Samuele Marro,Erik Brynjolfsson,Alex Pentland,Jiaxin Pei
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:
Abstract:System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.
[HC-1] CrossAtlas: Evaluating Projection Techniques for Spatial Referencing in Cross-Reality Collaboration
链接: https://arxiv.org/abs/2607.28583
作者: Haoyang Yang,Chenyang Zhang,Elliott H. Faa,Weijian Liu,Lily Seika Chisholm,Benjamin Lee,David Saffo,Feiyu Lu,Blair MacIntyre,Yalong Yang
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 8 figures, Accepted to IEEE ISMAR 2026 (TVCG journal track)
Abstract:Cross-reality collaboration increasingly connects immersive and desktop users within synchronized workspaces, yet little is known about how bidirectional projection techniques between immersive 3D layouts and desktop 2D views influence communication. Spatial referencing depends on shared spatial understanding, but different mappings preserve and distort geometric relationships in different ways, altering perceived adjacency, orientation, and coverage across collaborators’ views. We present CrossAtlas, a synchronized PC-VR collaboration platform that integrates multiple bidirectional projection techniques, including three planar projection variants and equirectangular, a spherical projection variant, across layouts of varying curvature. In a controlled study with 24 dyads, collaborators completed spatial referencing tasks under different projection-layout conditions while we collected performance and subjective measures. Our results show that projection choice strongly shaped collaboration, with the spherical variant often outperforming planar projections and remaining robust across object layouts.
[HC-2] Multi-Session User Experience Assessments of Computationally Optimized Automated Vehicle Functionality Visualizations
链接: https://arxiv.org/abs/2607.28552
作者: Mark Colley,Pascal Jansen,Svenja Krauss,Enrico Rukzio
类目: Human-Computer Interaction (cs.HC)
备注: accepted to AutomotiveUI 2026
Abstract:Understanding automated vehicles (AVs) is crucial to improving their acceptance. Numerous approaches to visualizing relevant traffic information to passengers have been proposed and empirically evaluated. As this is time-consuming, costly, and reduces the possible design parameters, we employed multi-objective Bayesian optimization to optimize the design of visualizations in AVs. In particular, we evaluated multi-session aspects involving iterative optimization. We optimized the design for passenger trust and perceived safety while minimizing cognitive load. Results from an online study (N=74) show that this method effectively identifies visualization design parameter values that improve trust, safety, and predictability while making the design process more efficient and scalable. However, shortcomings of the computational approach when optimizing for subjective measurements are highlighted and discussed.
[HC-3] Effects of Auditory Information for People With Visual Impairments in Highly Automated Vehicles
链接: https://arxiv.org/abs/2607.28544
作者: Mark Colley,Tobias Aescht,Omid Rajabi,Max Rädler,Pascal Jansen,Enrico Rukzio
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to AutomotiveUI 2026
Abstract:Automated vehicles promise to improve accessibility and access to personal mobility for everyone. However, their design and current research trends in visualizing relevant information do not reflect this commitment to accessibility for users with visual impairments. Therefore, we designed and implemented a visual and auditory communication concept for people with visual impairments seated inside fully automated vehicles. Furthermore, in an online video-based study (N=35, 12 with visual impairments), we compared three levels of auditory information communication: low (safety-relevant information only), medium (additionally including vehicle control and route updates), and high (additionally including sightseeing and destination information). Results showed that trust and user experience significantly improved with additional information, with a corresponding, albeit less robust, effect on perceived safety. However, they also revealed that a potential information saturation was reached with medium information. Our work helps to improve the accessibility of automated vehicles by guiding designers towards adequate information communication.
[HC-4] Identifying a Level-up Pathway for AI-assisted Counterspeech through Elaboration
链接: https://arxiv.org/abs/2607.28239
作者: Han Li,Inhwan Bae,Natalie Bazarova,Drew Margolin
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Given the profound societal impact of vaccine-skeptical content on social media, community-driven counterspeech has emerged as a promising participatory response to contest and curb such objectionable content. Yet crafting effective counterspeech remains challenging for ordinary users, limiting their willingness and ability to engage constructively. We designed and evaluated three generative AI-assisted counterspeech writing systems that vary by assistance stage (co-writing vs. re-writing) and mode (guided vs. unguided) to support lay users’ responses to vaccine-skeptical content. We ask whether AI can help users craft counterspeech perceived as both effective and authentic, which forms of AI support work best, and through what mechanisms. In a randomized controlled trial with social media users, participants wrote counterspeech responses to both statistical and narrative vaccine-skeptical content. Across evidence types, AI-assisted writing increased perceived counterspeech effectiveness while largely preserving authentic self-expression, and perceived effectiveness was the strongest predictor of willingness to counterspeak publicly. AI’s primary benefit was facilitating more elaborate writing, producing messages that were more informative, analytical, and lexically sophisticated. These findings suggest a level-up pathway for AI-assisted, community-driven counterspeech, which helps cultivate more effective and motivated counterspeakers, contributing to higher-quality public discourse on pressing societal issues.
[HC-5] oward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony
链接: https://arxiv.org/abs/2607.28204
作者: Guandong Pan,Yaqian Yang,Shi Chen,Yi Zheng,Yi Zhen,Hongwei Zheng,Shaoting Tang
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Continuous emotional arousal quantification remains bottlenecked by time-consuming and labor-intensive manual annotation. This work investigates group-level EEG dynamic neural synchrony (DNS) as a principled signal for continuous arousal quantification that bypasses per-subject manual labeling. Using Correlated Component Analysis (CorrCA) with sliding-window computation across four EEG datasets spanning 142 subjects and over 207 hours, we systematically evaluate DNS as a group-level marker for emotional arousal dynamics. Three key findings emerge. First, DNS exhibits significant emotion information from valence-dependent differences (all p0.003), with positive emotions eliciting higher synchrony. Second, DNS correlates more strongly with the first-order derivative of arousal than with raw arousal values, revealing that neural synchrony captures the rate of emotional change rather than static intensity. Third, we provide the first systematic characterization of how DNS-arousal coupling depends on key methodological choices, finding that moderate windows (10-30 s), positive lags (0-10 steps), and First-order Difference feature of EEG from the dominant CorrCA component yield consistently strong coupling. Subject-split replication and block permutation tests confirm these associations are not statistical artifacts. Our findings establish DNS as an empirically validated group-level marker toward annotation-efficient continuous emotional arousal quantification.
[HC-6] Student Perceptions and Preferences Regarding AI-Generated Instructional Videos in Computing Education
链接: https://arxiv.org/abs/2607.28203
作者: Esse Ciego,Shubbhi Taneja,Wilson Wong,Amanpreet Kapoor
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Students differ in how they prefer to engage with learning resources, with some favoring textual materials and others visual or video-based content. Recent advances in generative AI have led CS education research to focus on text-based AI tools for developing learning resources. However, advances in AI video models and the rapid proliferation of AI video generation tools have made it possible for instructors to create high-quality personalized educational videos efficiently and cost-effectively. Understanding students’ perceptions of AI-generated videos is thus critical for helping CS instructors know when and how to use them purposefully. To address this gap, we conducted a descriptive post-test survey study in which 170 computing students at two U.S. institutions watched three 3-minute AI videos on the Markdown markup language created with Knowlify. Students then completed a survey about their perceptions of the Markdown videos and their broader views on the use of AI-generated videos in education. Students rated the Markdown videos as high-quality, accurate, and usable, with nearly half unable to determine whether the videos were AI-generated. At the same time, students expressed limited comfort with the widespread adoption of AI videos in the classroom. They preferred AI videos for simple, supplemental, and visual use cases, while expressing concerns about lower-quality or inaccurate content, reduced instructor interaction, and diminished educational value.
[HC-7] A Mathematical Framework for Reading the Autopsias Meta - Compositional System
链接: https://arxiv.org/abs/2607.28155
作者: Patricio F. Calatayud,Pablo Padilla Longoria,Álvaro Martínez Ramírez
类目: Human-Computer Interaction (cs.HC)
备注: 13 pages, 12 figures, 11 equations
Abstract:Background. New forms of music writing using computers have arisen in the past 20 years. Most of them use the capacities of digital manipulation of data like animation, algorithmic processing, cinematic view, and much more. These scores use dynamic musicography, and all of them share a problem. They have readability problems. We will argue that this problem can be addressed by mathematical tools. Aims. Take the Autopsias [Autopsies] meta-compositional system as a study case for starting the construction of a mathematical framework that can overview the readability for cynetic musicography. The Autopsias system is the process of transforming a musical score dynamically. Methodology. We will start to build a mathematical framework by taking a group of basic topological concepts, and applying them after a bridge between a printed score and a dynamic computational re-appropriation of it has taken place. We will observe the writing and the performing of the several performances of the Autopsias composition, and analyze them from a mathematical perspective, in order to get the ability to start constructing a new set of orthographical tools for the performance of dynamic musicography. Main Contribution. The performances of Autopsias show that mathematics, as with pitch-class sets, can be helpful for understanding music information, as well as its performance. Also show the necessity of augmenting the set of orthographic rules, in order to improve the dialog between composers and performers.
[HC-8] Investigating Effective Uncertainty Visualizations for Ordinal Crowdsourced Data of Crowding Conditions
链接: https://arxiv.org/abs/2607.28072
作者: Bea Alexis Arcega,Annika Dominique S. Campos,Kathleen Therese Cruz,Aaron Ace Toledo,Briane Paul V. Samson
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Commuters often encounter crowding in railway systems, particularly in queues where passenger density varies throughout the day. This introduces uncertainty in crowdedness, making it difficult for individuals to anticipate conditions and plan their trips effectively. Crowdsourcing has been a valuable method for collecting localized user data. But the unpredictability of crowds and the uncertainty of crowdsourced information pose new challenges for decision-making. However, we know little about how to effectively visualize uncertainty in crowdedness to support informed commuting decisions, particularly when using crowdsourced ordinal data. Here, we investigated different uncertainty visualizations and their effectiveness in representing the variability and reliability of crowdsourced crowding data. They were evaluated through an online study, and we found that cluster visualization is best suited to reduce cognitive load while maximizing user confidence and trust. On the other hand, participants showed higher accuracy in determining crowd levels when using bubble treemaps.
[HC-9] VizPilot: Automated Onboarding for SVG-based Composite Visualizations using Multimodal LLM s IEEE-VIS2026
链接: https://arxiv.org/abs/2607.27938
作者: Nishaanthini Gnanavel,Yong Wang
类目: Human-Computer Interaction (cs.HC)
备注: Accepted by IEEE VIS 2026
Abstract:Composite visualizations integrate multiple visualizations to represent complex datasets effectively, but their intrinsic composite designs often impose a high initial cognitive load on novice users. Existing visualization onboarding approaches are typically platform-dependent, require substantial manual authoring effort, and struggle with the structural complexity of composite visualizations, limiting their general applicability. We present VizPilot, an automated visualization onboarding approach that reverse-engineers composite visualization structure to generate interactive onboarding experiences directly from raw visualization artifacts. VizPilot consists of two modules: a Composite Visualization Analyzer and an Onboarding Interface. Leveraging Multimodal Large Language Models (MLLMs), the Analyzer employs a two-stage pipeline that decomposes a visualization into visual components, extracts structured explanations, and maps them to precise SVG elements for reliable highlighting and interaction. Implemented as a browser extension, VizPilot requires only a brief visualization description and optional interaction source code from the visualization developer to automatically generate onboarding content. The Onboarding Interface supports both guided narrative scrollytelling and free exploration, enabling users to learn visualization components progressively or on demand. We evaluate VizPilot through a comparative analysis of different input modalities, a usage scenario demonstrating reduced authoring effort, and a user study assessing its impact on users’ cognitive load. The results demonstrate that VizPilot effectively automates the authoring of onboarding experiences while improving the usability and accessibility of composite visualizations.
[HC-10] Creative Task Cards for Reflection Self-Efficiency and Self-Regulation in CS1 Introductory Programming: Initial Insights
链接: https://arxiv.org/abs/2607.27863
作者: Corey Ford,Yinmiao Li,Rosa Van Koningsbruggen
类目: Human-Computer Interaction (cs.HC)
备注: In Proceedings of The First Reflection in Creative Experience (RiCE) Workshop (RiCE W1) arXiv:2607.24558
Abstract:Computer Science students in introductory programming courses struggle with emotional challenges such as low levels of sustained interest and a lack of self-belief. However, students are typically introduced to disciplinary strategies (such as debugging strategies) and not the metacognitive nor self-regulation strategies that could help them overcome such emotional challenges. This paper addresses this issue by introducing creative task cards designed to support students’ reflection, self-efficacy, and self-regulation in programming assignments. To evaluate the cards, twenty-nine students used them to complete a mock coding assignment in-class and participated in surveys and focus groups. Our initial qualitative insights suggest that students: found value in taking breaks, particularly when they could leave the classroom to reflect with peers; could use drawing to better reflect on feelings of cognitive overload; and felt more capable progressing with their work once breaking it into smaller steps.
[HC-11] Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
链接: https://arxiv.org/abs/2607.27851
作者: Ming Wang,Jiaqi Wu Young,Wenfang Wu,Daling Wang,Shi Feng
类目: Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Social and Information Networks (cs.SI)
备注:
Abstract:Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker’s emotional experience. Emotional support conversation selects and sequences support for the seeker’s current needs. Sustained use introduces a further goal. Effective support should sustain users’ capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non-use, transition, and termination. A targeted literature-and-corpus audit motivates this position. In a PRISMA-ScR-guided sample, 95% of 60 system-building papers pursue relief-oriented goals. None evaluates capability or longitudinal outcomes, and only 1 considers dependency, autonomy, or termination risk. In 300 ESConv supporter turns, capability-relevant functions appear in 43.0%, while generic suggestions account for 22.0%, compared with 4.0% reappraisal, 6.7% self-efficacy support, and 0.3% boundary behavior. We release a protocol for extending the audit to model behavior. An illustrative process model connects latent user capability to six design commitments, four evaluation timescales, and lifecycle constraints. The resulting agenda makes CSED testable across data, policy design, training, evaluation, and governance.
[HC-12] DP-LENS: A Density-Aware Polyfocal Lens with Topology-Driven Auto-Routing for Occlusion Management in Immersive 3D Analytics
链接: https://arxiv.org/abs/2607.27697
作者: Nieyu Cao,Xian Wang,Lik-Hang Lee
类目: Human-Computer Interaction (cs.HC)
备注: Accepted to IEEE International Symposium on Mixed and Augmented Reality (ISMAR) 2026, to appear in IEEE Transactions on Visualization and Computer Graphics (TVCG). 11 pages
Abstract:Immersive environments, e.g., virtual reality (VR), offer a unique approach to exploring complex 3D datasets, where data is often heavily occluded and exploration incurs a high cognitive load. We propose DP-LENS, a density-aware polyfocal fisheye lens equipped with topology-driven auto-routing. While preserving peripheral context through geometric deformation and 3D perspective techniques, it enables users to explore 3D data with a lower cognitive load. To facilitate hands-free macro-navigation, we integrate a Large Language Model (LLM) to serve as a supplementary voice-based target selection tool that initiates the auto-routing algorithm. Two user studies with 34 participants investigate the potential benefits of this system. Our first study (N=18) compared the manual DP-LENS against two industry-standard baselines (i.e., World-in-Miniature and volumetric slicing) in heavily occluded 3D datasets. The results show that DP-LENS significantly reduced cognitive load, decreased completion time, and improved user preference. The second study (N=16) compared the topology-driven auto-routing system (initiated via voice commands) with a fully manual DP-LENS. The results show that the auto-routing system improved task efficiency, further reduced cognitive load, and garnered higher user preference. Furthermore, the auto-routing partially decoupled exploration efficiency from the physical dimensions of the data and mitigated physical fatigue to some extent. Based on the findings, we proposed design implications to inform the development of more spatially scalable and low-fatigue interactions for future 3D visual analytics systems.
[HC-13] Recognition and Label-Free Adaptation Across Recording Sessions in Surface-EMG Gesture Decoding
链接: https://arxiv.org/abs/2607.27568
作者: Jethro Odeyemi,W. J. Zhang
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 23 pages, 6 figures, 6 tables
Abstract:Recognition accuracy obtained during a recording session does not persist when a user puts on the electrodes again after the electrodes had previously been removed. The electrodes may have moved slightly, the skin may be drier or wetter, or the elbow may be positioned differently; these factors all contribute to day-to-day variability and therefore represent a major obstacle to implementing successful pattern-recognition based myoelectric control systems in daily practice. However, simply recalibrating a user’s hand for 20 min at every doff/don event is a clearly unrealistic expectation. A montage-agnostic encoder built for cross-user, cross-montage transfer is trained here using data collected during a particular recording session, and then applied to data collected later in a different recording session without adjusting anything, on the ten intact subjects of NinaPro DB6. The performance of this approach is compared to that of a per-user LDA classification pipeline, and to that of two published approaches that only rely on source data collected from the same recording session. Carried unchanged across recording sessions, the encoder retains 0.688 macro-F1 against 0.540 for the per-user pipeline, and, on the per-window metric the published baselines use, sits above both published source-only results, a band of two points that locates the encoder rather than ranking it. Of five label-free test-time adaptations, only feature-statistic alignment improves every subject; batch-normalisation re-estimation, a standard method in the domain-adaptation literature, collapses this architecture entirely. Aligning the encoder’s feature statistics to the new session recovers about what a single labelled calibration repetition would.
[HC-14] A Montage-Agnostic Encoder for Calibration-Light Cross-User Gesture Recognition from Surface Electromyography
链接: https://arxiv.org/abs/2607.27565
作者: Jethro Odeyemi,W. J. Zhang
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 22 pages, 5 figures, 7 tables
Abstract:Pattern-recognition control promises a myoelectric prosthesis that responds to many intended gestures rather than one or two, but the promise has stayed in the laboratory. A recogniser trained on one person rarely transfers to the next, and useful performance usually demands a fresh round of labelled calibration from the end user. A montage-agnostic encoder is introduced that reads each electrode with shared weights and locates it by its physical coordinate rather than its index, so one architecture ingests any channel count without montage-specific parameters. Trained across users, it exceeds a per-user Hudgins and linear-discriminant classifier by 0.234 macro-F1 on DB1 for every held-out subject and by 0.108 on DB2, and falls below it on the ten-subject DB5. Each of the encoder’s three key components individually accounts for more than half of its 3-shot macro F1 in an otherwise budget-matched ablation study. A controlled subject-count sweep shows the margin is close to flat from nine training subjects to thirty-nine, so the training pool binds only as a stability floor below which cross-user training fails to converge; what tracks the direction of the comparison across the three databases is instead the strength of the per-user baseline, which signal fidelity sets. Comparing against an LDA baseline depends on budget spent training models and on how good that baseline is, and self-supervised pretraining had no benefits once a supervised model was adequately trained.
[HC-15] FADEx: Feature Attribution and Distortion-based Explanation of Dimensionality Reduction
链接: https://arxiv.org/abs/2607.27463
作者: Lucas Greff Meneses,Evandro S. Ortigossa,Claudio Silva,Luis Gustavo Nonato
类目: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
备注: 18 pages, 17 figures, to be published in IEEE Transactions on Visualization and Computer Graphics
Abstract:Dimensionality Reduction (DR) is a fundamental tool for high-dimensional data exploration, reducing the complexity of latent spaces of machine learning models, and assisting in the explanation of complex opaque models. However, non-linear DR techniques often function as opaque transformations themselves, making it challenging to understand how individual features influence instance positioning in the reduced space. This lack of transparency complicates the analysis and interpretation of structural patterns, hindering the ability to reason about the organization of high-dimensional data based on the projected layout. In order to address this challenge, dimensionality reduction explanation methods have shown promise in improving the understanding of the observed groups and cluster structures. Unfortunately, existing DR explanation approaches tend to suffer from limitations such as multiple attributions per feature and restricted applicability to specific dimensionality reduction methods, which hinder their use. In this work, we propose FADEx, a novel local per-instance feature attribution method that leverages local linear approximation via first-order Taylor expansion and Singular Value Decomposition to provide explanations. FADEx computes the local linear models via weighted least squares, eliminating the need for out-of-sample data mapping, making it agnostic to the DR method, while simultaneously providing local feature attributions and distortion analysis. Through qualitative and quantitative evaluations, comparisons with existing methods, and case studies, we demonstrate FADEx’s effectiveness and versatility in providing explanations and analytical resources for analyzing the behavior of DR methods. The results indicate FADEx yields robust and reliable explanations, outperforming existing approaches in several aspects.
[HC-16] Evaluating the Vergence-Accommodation Conflict in Gaze-Based 3D Target Selection
链接: https://arxiv.org/abs/2607.27369
作者: Mohammad Raihanul Bashar,Mohammadreza Amini,Aunnoy K Mutasim,Mayra Donaji Barrera Machuca,Wolfgang Stuerzlinger,Anil Ufuk Batmaz
类目: Human-Computer Interaction (cs.HC)
备注: 9 Pages, 7 Figures, IEEE ISMAR 2026 (TVCG)
Abstract:State-of-the-art head-mounted displays (HMDs) enable gaze-based selection in virtual environments. Yet, these HMDs suffer from the vergence-accommodation conflict (VAC), which is known to affect interaction performance. The VAC might influence gaze-based selection performance because it directly affects eye-movement behavior. Thus, in this paper, we investigate how the VAC influences gaze-based 3D target selection across varying depth conditions. Our results show that as the (visual) depth increases, user performance significantly decreases with gaze-based selection. Moreover, a previously suggested Variation in Diopter Fitts’ law model captured this performance change better relative to a linear model. These findings provide evidence that gaze-based pointing is negatively affected by the VAC and highlight the importance of accounting for depth-dependent factors when designing gaze-based interaction in 3D environments.
[HC-17] Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy
链接: https://arxiv.org/abs/2607.27212
作者: Asif Azad,Mohammad Sadat Hossain,MD Sadik Hossain Shanto,Sabri Boughorbel,Abdulrhman Aljouie,Bdour Alwuqaysi,Yahya Bokhari,Ayah Othman Sindi,Ehsan Hoque
类目: Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
备注:
Abstract:Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a near-total absence of culturally grounded digital therapy materials. We present Digital Harf, a pervasive, multimodal AI platform that extends clinician-led speech and language therapy into the home. The system integrates three therapeutic modules - language therapy, speech intelligibility, and picture description - within a unified workflow that adapts to each child’s performance over time. To address the critical shortage of Arabic therapy content, we introduce an Agentic Synthetic Data Engine (ASDE) that automatically generates culturally relevant images, prompts, and language tasks guided by explicit therapeutic and cultural criteria. Expert evaluation with 13 licensed Speech-Language Pathologists yielded a 90.1% clinical acceptance rate for ASDE-generated content without any manual curation or selection, and strong ratings for cultural and linguistic alignment across the full platform. Digital Harf demonstrates that AI-driven therapeutic systems can be built from the ground up for underrepresented linguistic settings, treating cultural grounding as core infrastructure rather than an adaptation afterthought.
[HC-18] Snapshot plots: displaying summary tables as parallel univariate plots with consistent color highlighting
链接: https://arxiv.org/abs/2607.28302
作者: Matthias Schonlau,Sandra Huang,Tiancheng Yang
类目: Applications (stat.AP); Graphics (cs.GR); Human-Computer Interaction (cs.HC)
备注:
Abstract:For empirical studies, social and health scientists give background characteristics of their sample and summarize them in the famous “Table 1”. When treatment/ control groups are present, this table gives summary statistics by group to see whether the background characteristics differ by group. We propose snapshot plots — parallel univariate plots with consistent highlighting — to visualize such tables. Compared to “Table 1”, such plots are designed to facilitate comparisons of background characteristics — in particular among groups — and give more detail on numerical variables. We provide a web app as well as a python implementation of snapshot plots. Snapshot plots arise as edge cases of hammock plots (parallel coordinate plots for mixed categorical/ numerical data). We demonstrate the usefulness of snapshot plots for two "Table 1"s.
计算机视觉
[CV-0] ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
链接: https://arxiv.org/abs/2607.28627
作者: Yao Xiao,Reuben Tan,Zhen Zhu,Yuqun Wu,Jianfeng Gao,Derek Hoiem
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code: this https URL
Abstract:Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: this https URL
[CV-1] ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
链接: https://arxiv.org/abs/2607.28625
作者: Yukang Cao,Haozhe Xie,Beichen Wen,Runmao Yao,Yinghao Liu,Yue Huang,Zhichao Liao,Yunxiang Wang,Haiheng Liu,Xingshun Tian,Dawei Su,Long Zhuo,Dacheng Tao,Xiaogang Wang,Liang Pan,Ziwei Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL
Abstract:Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
[CV-2] PhiZero: A World Model Built Around Physical Language
链接: https://arxiv.org/abs/2607.28624
作者: Shuyao Shang,Yuqi Wang,Ruopeng Gao,Xu Chen,Tieniu Tan,Lue Fan,Zhaoxiang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans’ ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
[CV-3] Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
链接: https://arxiv.org/abs/2607.28611
作者: Chongjian Ge,Hanwen Jiang,Tianyu Wang,Jiuxiang Gu,Yiran Xu,Ziwen Chen,Shaoteng Liu,Jing Shi,Yicong Hong,Zefan Cai,Hailin Jin,Hao Tan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 40 pages
Abstract:Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor’s functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.
[CV-4] Beacon: Knowing When and How to Perform Agent ic Visual Reasoning
链接: https://arxiv.org/abs/2607.28595
作者: Qixun Wang,Yang Shi,Letian Cheng,Zhuoran Zhang,Yan He,Yuqi Tang,Qi Zhang,Xinlei Yu,Ruizhe Chen,Tianrun Xu,Yuanxing Zhang,Pengfei Wan,Haotian Wang,Xianghua Ying
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 33 pages
Abstract:The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model’s capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hint-Guided Capability Expansion mechanism in the reinforcement learning stage, which respectively encourage adaptive tool invocation based on task necessity and strengthen the model’s tool-use capability on the most challenging problems. Extensive experiments across diverse benchmarks demonstrate the strong overall performance of Beacon and its substantial improvements in both Mode Adaptiveness and Tool Effect.
[CV-5] MixFrag : Frag ility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
链接: https://arxiv.org/abs/2607.28589
作者: Md. Mehrab Hossain Opi,Robiul Islam Ryad,Md. Umar Faruk
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However, existing PTQ methods typically employ uniform bit-widths across transformer components, overlooking their heterogeneous sensitivity to quantization and leading to inefficient precision allocation. In this paper, we propose MixFrag, a fragility-guided mixed-precision PTQ framework for Vision Transformers. MixFrag first estimates component-level quantization fragility by measuring the Kullback–Leibler (KL) divergence between full-precision and isolated quantized output distributions using a small calibration set. It then formulates bit allocation as a Multiple-Choice Knapsack Problem (MCKP), enabling adaptive layer-wise precision assignment under a target bit budget. Extensive experiments on ImageNet-1K across multiple Vision Transformer architectures demonstrate that MixFrag achieves competitive classification performance under practical mixed-precision settings. Furthermore, evaluations on COCO object detection and instance segmentation show that MixFrag achieves state-of-the-art performance among existing mixed-precision PTQ methods, improving the previous best method by up to 9.6 AP under the challenging MP3/MP3 setting. Additional analyses validate the proposed fragility metric and demonstrate its strong correlation with the learned bit allocation. These results establish MixFrag as an effective framework for mixed-precision post-training quantization of Vision Transformers.
[CV-6] ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
链接: https://arxiv.org/abs/2607.28581
作者: Xiao Luo,Mingyang Du,Xin Zhou,Tianrui Feng,Xiwu Chen,Xiaofan Li,Jiangning Zhang,Dingkang Liang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requires learning geometry from scratch and overlooks the rich semantic and structural priors already encapsulated in discriminative 3D foundation models. We contend that leveraging the profound understanding of the 3D world possessed by these discriminative models can significantly reduce generative cost. To this end, we propose ROAD, a framework that reduces the training cost of 3D generation by transferring these rich discriminative priors into diffusion transformers. To address the inherent semantic-structural heterogeneity between generative and discriminative latents, we introduce a reciprocal-objective alignment strategy. This method synergizes Holistic Semantic Condensing to enforce global semantic coherence and Structural Optimal Alignment, which is formulated as a bipartite matching problem to rigorously align microscopic geometric details between disparate latent spaces. The 3D foundation model is only used for training-time supervision of alignment and is not used at inference, incurring no additional inference cost. Compared with the industrial baseline Step1X-3D, the proposed ROAD achieves highly competitive generation performance with only 1.5% of the training data and significantly reduces training costs, effectively reducing the computational overhead of high-fidelity 3D generation. Code is available at this https URL.
[CV-7] MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion ACM-MM2026
链接: https://arxiv.org/abs/2607.28565
作者: Yunzhan Fu,Xiangyu Shen,Yifei Sun,Yuhan Chen,Jian Wu,Hongxia Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14pages, 14 figures, accepted by ACM MM2026
Abstract:Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter. This module explicitly extracts source image features before serialization, injecting them into the network via strict dimensional alignment to effectively supplement image features. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss. This loss ensures deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion, holding significant promise for intent-driven intelligent clinical decision support systems.
[CV-8] ScaFE: Data-Efficient Scar Classification with LLM -Generated Clinical Feature Programs
链接: https://arxiv.org/abs/2607.28538
作者: Ruman Wang,Hangting Ye
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and feature-level SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Iterative refinement also raises the executable-program rate from 66.7% to 95.0%, with verified evidence for 91.7% of the final features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.
[CV-9] MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition
链接: https://arxiv.org/abs/2607.28532
作者: Alex Andonian,Samuel G Rodriques,Andrew D White,Siddharth M Narayanan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing machine learning model training sets, they must be transformed into line notations. The two common forms of this task are translating an image of a single molecule (optical chemical structure recognition - OCSR) and translating a Markush structure that represents a family of molecules. While prior work in the former case is quite mature, Markush structure parsing remains a challenging task. In this work, we treat both tasks as an image-to-text translation problem. We then propose OCSRGlyph, a state-of-the-art OCSR model, improving performance over prior methods by carefully considering stereochemistry. For the Markush task, we introduce MarkushGlyph, a vision-language model that reads the entire Markush structure as an image. This contrasts with prior systems, which often use multiple stages to separately process visual and text input content. Finally, we introduce a new metric for determining the accuracy of Markush structure translations, handling failure modes present in prior metrics.
[CV-10] What to Remove What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
链接: https://arxiv.org/abs/2607.28526
作者: Cencen Liu(1),Wen Yin(1),Dongyang Zhang(1),Dongmin Li(1),Shan Zhao(2),Bing Su(2),Tao He(1),Jielei Wang(1),Guoming Lu(1) ((1) University of Electronic Science and Technology of China, (2) Jiigan Technology)
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonly encode heterogeneous degradation conditions in a shared latent space, where degradation-related cues and scene content can remain entangled. We characterize the resulting challenge as dual ambiguity: semantic ambiguity in channel-wise modulation and spatial ambiguity in restoration responses, which can lead to content corruption and residual artifacts. To mitigate this issue, we propose DAR-Net, a Dual-Ambiguity Rectification Network for all-in-one image restoration. DAR-Net first introduces a Degradation Archetype Representation (DAR) module to construct a structured degradation state through simplex-constrained archetype mixture modeling. Based on this state, a Semantic Ambiguity Rectification (SeAR) module generates degradation-aware prompts to improve channel-wise conditioning in the decoder. A Spatial Ambiguity Rectification (SpAR) module further regularizes degradation-aware and complementary features toward orthogonal response subspaces, reducing spatial interference between removal and preservation cues. Extensive experiments on standard all-in-one restoration benchmarks show that DAR-Net achieves the best overall performance under both three-degradation and five-degradation settings, improving the average PSNR over the strongest competitor by 0.14 dB and 0.34 dB, respectively; it additionally shows superior performance on CDD-11 and WeatherBench.
[CV-11] Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
链接: https://arxiv.org/abs/2607.28516
作者: Bowen Liu,Shuning Wang,Xinpeng Ding,Zhiheng Wu,Bodong Du,Xiaomeng Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answering. Our key idea is to organize selected frames into query-relevant cross-frame evidence before generation. We formulate this post-selection stage as a latent evidence interface and instantiate it with GenEvA ( \textbfGenerative Latent \textbfEvidence \textbfAggregation ), a distribution-guided latent evidence aggregation framework. Specifically, GenEvA uses a query-conditioned evidence distribution to focus aggregation on relevant frames, forming compact cross-frame latent evidence from their frame-specific information. Since cross-frame integration is not always needed, the same distribution determines whether to insert this latent complement. Across four benchmarks and two Video-MLLM backbones, GenEvA consistently improves matched-frame baselines. At 8 frames, it raises the four-benchmark LLaVA-Video average by +5.2 points and Qwen2.5-VL accuracy on LVBench by +10.1 points. These gains require only 0.11% – 0.40% average video-token overhead; analyses further show task-aware allocation and benefits from Adaptive Evidence Invocation.
[CV-12] RefCaptioner: Multi-Reference Image-Grounded Video Captioning
链接: https://arxiv.org/abs/2607.28509
作者: Tengfei Liu,Yang Shi,Yuran Wang,Xiaohan Zhang,Yuqing Wen,Yuqi Tang,Qixun Wang,Zhuoran Zhang,Xuanyu Zhu,Weihong Lin,Xinlei Yu,Yujie Wei,Xinwei Long,Fengxiang Wang,Xinlong Chen,Yue Ding,Jialu Chen,Haotian Wang,Yuanxing Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: this https URL
Abstract:Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing 20,000 videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
[CV-13] AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
链接: https://arxiv.org/abs/2607.28487
作者: Jingwen Yang,Senmao Wang,Luoyao Kang,Runmeng Cui,Keying Zhang,Yunjia Bao,Haifan Gong,Lin Lin,Haiyue Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.
[CV-14] owards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles ICANN2026
链接: https://arxiv.org/abs/2607.28483
作者: Luca de Martino,Federico Aromolo,Federico Nesti,Giorgio Buttazzo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 2 figures, 3 tables. Accepted at the Efficient Deep Learning: Methods and Applications workshop, 35th International Conference on Artificial Neural Networks (ICANN 2026)
Abstract:Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational cost limits their deployment on embedded hardware. This work presents an efficient and accelerated pipeline designed for both embedded and desktop platforms, targeting the autonomous driving and railway domains. The proposed approach reformulates the Neyman-Pearson scoring stage of PixOOD, a state-of-the-art out-of-distribution detection method, and deploys the full pipeline through hardware-optimized TensorRT compilation, reaching up to 182 FPS on a desktop NVIDIA RTX 4060 GPU and 75 FPS on the NVIDIA Jetson AGX Orin embedded platform, respectively 20x and 18x faster than the original baseline. The achieved results demonstrate that advanced anomaly segmentation can be efficiently deployed for onboard processing in autonomous driving and railway applications.
[CV-15] owards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
链接: https://arxiv.org/abs/2607.28470
作者: Antonio Delgado-Rosa,David Muñoz-Valero,Enrique Adrian Villarrubia-Martin,Juan Moreno-Garcia
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 43 pages, 14 figures
Abstract:Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conventional approach still transmits terabytes of raw imagery to the ground for processing, and open satellite datasets for aircraft are scarce and severely class-imbalanced. These limitations either delay timely decision-making or prevent standard detectors from learning robust representations of rare aircraft classes. In this paper, a workflow that combines on-board inference with generative data augmentation is proposed to address both limitations jointly. Inference is executed on a 6U CubeSat equipped with a low-power edge tensor accelerator, while a diffusion model fine-tuned through low-rank adaptation generates synthetic minority-class imagery. This synthetic output is automatically annotated, pseudo-labelled, by an intermediate detector and merged with classically augmented samples. The results show that the balanced dataset increases global mean average precision from 77.9% to 82.2%, with the minority class rising from F1=0.683 to F1=0.811, and that the quantised detector fits the on-chip memory and projects 25-30 frames per second on orbit. This approach contrasts with the conventional bent-pipe architecture, in which the satellite acts as a passive data collector. Therefore, the computational tests support the proposed workflow as a decision-support tool for real-time, autonomous airborne surveillance from nanosatellites.
[CV-16] Can Vision-Language Models Reason about AI Edits in Images?
链接: https://arxiv.org/abs/2607.28464
作者: Darsha Udayanga,Pin-Yu Chen,Payel Das,Qiang Ji
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Language Models (VLMs) offer a promising alternative due to their strong visual understanding and reasoning capabilities; however, existing approaches typically rely on supervised finetuning with curated explanations rather than exploiting their inherent reasoning capabilities. In this work, we investigate whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision. Motivated by the success in Group Relative Policy Optimization (GRPO), an RL technique that incentivizes the model to reason by generating thinking traces prior to giving the final answer, we propose a GRPO-based training framework that utilizes simple accuracy and format rewards. Given an input image, the model produces a structured reasoning trace and predicts whether the image has been tampered with. A lightweight segmentation model is then guided by the reasoning output to generate pixel-level localization masks. Experiments across multiple image manipulation datasets demonstrate that our approach achieves competitive detection and localization performance compared to state-of-the-art image forgery detectors, despite requiring substantially weaker supervision. We introduce effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization. These results suggest that reinforcement learning provides an effective and scalable mechanism for teaching VLMs to reason about AI-generated content.
[CV-17] VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
链接: https://arxiv.org/abs/2607.28463
作者: Haiyue Zhang,Yi Bin,Xun Jiang,Zeyu Ma,Duo Peng,Guoqing Wang,Yang Yang,Heng Tao Shen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to the large number of visual tokens and limited context windows. Visual sampling provides a practical solution by selecting an informative subset of frames. However, existing methods typically either rely on relevance-aware sampling, leading to redundant frame selection and insufficient temporal coverage, or adopt a fixed sampling strategy regardless of query type. In this paper, we propose VisualRouter, a training-free and plug-and-play framework for query-grounded visual sampling. VisualRouter first classifies each query as either global or local and then applies the corresponding sampling strategy. For global queries, it employs a relevance-coverage hybrid strategy that preserves temporal coverage while retaining query-relevant visual evidence. For local queries, it adopts an event-aware frame selection strategy that performs event partitioning, segment-level frame allocation, and intra-event frame selection, jointly balancing relevance, coverage, and diversity with a limited number of input frames. Experiments show that VisualRouter consistently improves multiple LVLMs over uniform sampling, achieving gains of 5.2%, 7.7%, and 11.6% on Video-MME, LongVideoBench, and MLVU with Qwen2.5-VL-7B, and outperforming existing training-free visual sampling methods under the same setting.
[CV-18] ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
链接: https://arxiv.org/abs/2607.28442
作者: Ping-Kun Chiang,Kun-Ru Wu,Po-han Li,Sandeep Chinchali,Ufuk Topcu,Yu-Chee Tseng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbfViewMind3D, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird’s-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What’’ questions in SQA3D, while maintaining strong overall accuracy (50.8%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.
[CV-19] Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
链接: https://arxiv.org/abs/2607.28428
作者: V.S. Usatyuk,D. A. Sapozhnikov,S. I. Egorov
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Information Theory (cs.IT)
备注: 42 pages, 10 figures, 5 tables, was presented at the 10th International Conference ‘Deep Learning on Computational Physics (DLCP2026)’, under review for the Moscow University Physics Bulletin, Physics series
Abstract:We introduce Kohn–Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral embedding evaluated at the Nishimori temperature of an associated Random-Bond Ising Model. By mapping pre-trained features onto quasi-cyclic low-density parity-check graphs and constructing a regularized Laplacian acting as a Kohn–Sham Hamiltonian, we solve D independent channel spectral problems in \mathcalO(N\log N + k^2_\textmode N) time via FFT on circulant blocks (leveraging Pontryagin self-duality of \mathbbZ/p\mathbbZ ) and low-order Rayleigh refinement. Graph topology is optimized using \emphstar-domain surgery: rather than destroying information-carrying codewords by removing frustrated cycles, we construct edge shifts creating local convexity around codewords while bounding residual frustration to \rho(B_\gamma)\leq 1+\delta . Multi-scale fractal analysis ( D_2 spectrum) and fractal learning-rate landscape certifies a landscape transition from rough regimes ( D_23 ) to star-domain basins ( D_21 ), enabling Rayleigh refinement with k_\textmode=5 modes. We prove six theoretical results: a generalized Ihara–Bass identity linking belief propagation to the Laplacian; trapping-set eigenvalue correspondence; additive channel separability with an explicit exchange-correlation bound; a surgery theorem bounding frustration with attractor width \Omega(1/\sqrtd_\min) ; a quasi-stationarity perturbation bound; and a fixed-point convergence theorem. In a transductive protocol on ImageNet-1000 with frozen EfficientNet-B4 features ( D=1792 ), KSSE achieves \textbf88.93% Top-1 accuracy using \approx 21.24 M parameters, outperforming Swin-L (197M, 86.4–87.3%) and matching ViT-H/14 (632M, 88.0–89.5%) under standard inductive setups, while reducing model footprint by 10\times and 30\times , respectively.
[CV-20] Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features
链接: https://arxiv.org/abs/2607.28423
作者: Katy L. Scott,Sejin Kim,Joshua Siraj,Caryn Geady,Matthew Boccalon,Mattea Welch,Mogtaba Alim,Andrew J. Hope,Benjamin Haibe-Kains
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 22 pages (including supplementary), 6 figures, 2 supplementary tables, 5 supplementary figures
Abstract:Radiomics and imaging foundation models promise non-invasive biomarkers of tumour biology, yet predictive signatures may reflect tumour volume or acquisition artifacts rather than meaningful image structure. We introduce READII-2-ROQC, an open-source framework that uses volume-preserving negative controls to assess whether radiomic and deep imaging features capture independent spatial signals. READII-2-ROQC generates voxel-perturbed images across tumour, background and whole-image regions using configurable randomization strategies, then compares feature behaviour and model performance between original and control images. Applied to three public cancer imaging cohorts, the framework processed 3,552 tumour volumes and extracted PyRadiomics and foundation-model features from original images and nine matched controls. Reproducing published survival and HPV-status signatures, we show that multiple models retain performance after spatial structure is destroyed, revealing volume-driven or contextual confounding, whereas others show perturbation-sensitive signal. READII-2-ROQC provides a scalable quality-control strategy for developing interpretable, biologically grounded imaging biomarkers and reproducible radiomics workflows.
[CV-21] QQWorld: Quantile-Quantile Matching for World Model Regularization
链接: https://arxiv.org/abs/2607.28415
作者: Zhoushun Yu,Xiaoyu Hu,Xiangyu Xu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Robotics (cs.RO)
备注:
Abstract:Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.
[CV-22] Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
链接: https://arxiv.org/abs/2607.28401
作者: Ilya Novikov,Svetlana Illarionova,Ruslan Dzharkinov,Maria Smirnova,Ayrat Abdullin,Anna Korotkova,Mariia Ulianova,Dmitrii Shadrin,Evgeny Burnaev
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 37 pages, 11 figures, 9 tables. Preprint submitted to Earth Systems and Environment. This version has not been peer reviewed
Abstract:Effective flood monitoring is critical for minimizing the impacts of flood disasters on populations and infrastructure. Yet reliable remote sensing across extensive and environmentally diverse regions remains challenging, as most segmentation algorithms lack the generalisation capacity required for large-scale application, while annotated flood data are scarce and unevenly distributed. This study presents an end-to-end multimodal framework for Russian Federation territories sustainable flood monitoring and damage assessment based on synthetic aperture radar data, multispectral imagery, and digital elevation models with their derivatives, forming a 21-channel input. Using a self-collected multimodal dataset covering seven Russian regions, two strategies for water surface detection under limited data conditions were compared: a supervised U-Net++ model and the self-supervised AnySat architecture pre-trained and fine-tuned for the segmentation task. Under the data conditions of this study, supervised learning proved more effective, while the AnySat-based approach offered greater stability and retains advantages for settings where larger unlabelled data or missing modalities at inference are expected. The best flood area predictions were used to estimate flood impact in urban areas in terms of the area affected, material damage, casualties, and ecological and agricultural impact. The estimations were conducted following the official methodology of the Russian Ministry of Emergency Situations. Applied to the 2019 Tulun flood, the obtained results closely matched official assessments, except for material damage, due to the open-source databases usage. The results demonstrate the potential of deep learning and multimodal satellite data integration for scalable, reliable flood monitoring across diverse environmental and data-limited conditions. Comments: 37 pages, 11 figures, 9 tables. Preprint submitted to Earth Systems and Environment. This version has not been peer reviewed Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.28401 [cs.CV] (or arXiv:2607.28401v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.28401 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-23] Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction Generation and Embodied Transfer
链接: https://arxiv.org/abs/2607.28394
作者: Weiquan Lin,Yu Deng,Shiyang Liu,Luping Xiao,Xu Tang,Junzhi Yu,Jiaolong Yang,Lei Zhang,Xingyu Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models’’ without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.
[CV-24] Explaining Image Similarity with Automatically Extracted Concept Activation Vectors
链接: https://arxiv.org/abs/2607.28386
作者: Isaac Roberts,Petra Bevandic,Alexander Schulz,Barbara Hammer
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing explainability methods often rely on gradient-based attribution maps to provide local justifications for similarity. These approaches struggle to provide global insights into what specifically drives similarity in regions of an embedding space, such as texture, shape, or color. We introduce a model- and metric-agnostic framework that explains image similarity using Concept Activation Vectors (CAVs) extracted automatically via Sparse Autoencoders (SAEs). Given a pair of images, we perturb their embeddings along discovered concept directions and measure the resulting change in a chosen similarity function, yielding concept importances. For image pairs, we provide localization with concept attribution maps. We extend this procedure to group-level settings, explaining what drives similarity across a cluster of images rather than a single pair, and further, we introduce Exemplar Retrieval, aiming to recover samples with similar reasons contributing to similarity. Our experiments show that our latent perturbations are more faithful to the underlying data distribution than pixel-space baselines, and that concept importances linearly recover the true similarity score. Qualitative results further confirm the usefulness of our methods in understanding a model’s individual and group similarity judgments.
[CV-25] ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
链接: https://arxiv.org/abs/2607.28362
作者: Jin Cao,Zian Meng,Kaipeng Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: this https URL
Abstract:We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at this https URL
[CV-26] Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
链接: https://arxiv.org/abs/2607.28341
作者: Jie Ma,Zhike Qiu,Jie Gao,Jiayi Ji,Qian Chen,Xiaoshuai Sun,Rongrong Ji
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates “late-blooming” tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.
[CV-27] Same Branches Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement MICCAI
链接: https://arxiv.org/abs/2607.28327
作者: Maame Owusu-Ansah,Kelvin Lee,Dr Vinod Venugopal,Muhammad Moazzam Jawaid,Wenting Duan,James Brown
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at STACOM 2026 (MICCAI workshop). 11 pages, 3 figures, 2 tables
Abstract:Fractional flow reserve derived from CT angiography (FFR-CT) simulates flow through a patient-specific vessel model, so its accuracy depends on the connectedness of the segmented tree, not only on volumetric overlap: a segmentation can reach high Dice yet sever a bifurcation, dropping the downstream subtree and reversing the treatment decision. Topology-aware losses such as clDice and Skeleton Recall act on the global centreline and can miss localised breaks. We study the Bifurcation Connectedness Score (BCS), which scores connectedness at each ground-truth bifurcation, and soft-BCS, its differentiable training surrogate. BCS captures a property of segmentation quality the standard metrics miss: it responds strongly to breaks in connectedness while staying largely unchanged under connectedness-preserving narrowing. Higher BCS accompanies closer agreement between the FFR-CT decisions a solver makes on predicted versus ground-truth geometry, most clearly in severe disease (OR 2.16, CI [1.23, 4.18]). Both decisions come from the same solver, so this reflects geometric, not clinical, fidelity. In training, soft-BCS and Skeleton Recall recover the same branches but build different trees. Recovering branches and keeping them connected are separable properties, so we recommend reporting a measure of each.
[CV-28] AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction
链接: https://arxiv.org/abs/2607.28320
作者: Peiyi Xu,Junpeng Zhang,Guanbin Li,Ronghua Shang,Mingtao Feng,Le Dong,Weisheng Dong,Guangming Shi,Jie Feng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 4 figures
Abstract:Monocular UAV videos provide valuable observations for dynamic reconstruction of complex urban scenes. However, such scenes exhibit pronounced spatiotemporal heterogeneity: different regions follow distinct temporal activity patterns, while the motion states of some dynamic regions may further evolve over time. Although dynamic Gaussian methods based on decomposed shared spatiotemporal feature fields have achieved efficient and accurate reconstruction in object-centric or relatively compact scenes, their commonly adopted fixed plane-wise feature combination mechanisms are less suited to the heterogeneous local dynamics of UAV scenes, often leading to ghosting artifacts and blurred dynamic details. To address this challenge, we propose AdaAnchor4D, an adaptive anchor deformation framework for monocular UAV dynamic scene reconstruction. At its core, Anchor-Conditioned Feature Aggregation (ACFA) adaptively aggregates shared spatiotemporal features using anchor-specific aggregation embeddings and temporal information, allowing different local units to obtain dynamic representations tailored to their local and temporal states. Decoupled Local Geometry Deformation (DLGD) separates anchor-state deformation from local Gaussian geometry deformation, while Density-Adaptive Coordinate Warping (DACW) reparameterizes feature-query coordinates according to the axis-wise anchor distributions, alleviating the mismatch between non-uniform geometric sampling and uniform grid parameterization. Experiments on UAV-Arc4D, VisDrone, and UAVDT show that AdaAnchor4D achieves higher rendering quality than representative dynamic Gaussian methods while maintaining real-time rendering performance. The code will be made publicly available.
[CV-29] ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
链接: https://arxiv.org/abs/2607.28312
作者: Mingkang Dong,Muxin Pu,Jie Li,Bohan Guo,Songruo Chen,Bin Ren,Xu Zheng,Chen Zhao,Tianwen Qian,Mohamed Elhoseiny,Yuqian Fu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 9 pages
Abstract:Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understanding. ObjectStream induces spatially coherent latent objects directly from frozen Video-LLM representations, links them across frames into persistent anchors, and maintains their histories under a bounded memory budget, without requiring external object detectors or segmentation models. Built on these anchors, ObjectStream preserves three complementary forms of evidence: persistent object histories, transient object changes, and recent visual context. This design enables existing Video Large Language Models (Video-LLMs) to reason over object identities, interactions, and state changes while leaving the underlying model unchanged. Extensive experiments on online streaming and offline long-video benchmarks demonstrate both effectiveness and efficiency. In online streaming evaluation, ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, while reducing peak GPU mem-ory and TTFT by approximately 50%. On offline long-video benchmarks, it surpasses the full-token baseline while discarding 82.5% of visual tokens. These results highlight latent objects as a practical and effective organizing principle for compact streaming video memory.
[CV-30] MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians
链接: https://arxiv.org/abs/2607.28300
作者: Pouya Ardekhani,Zahra Dehghanian,Morteza Abolghasemi,Hamid R. Rabiee
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features. We present a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration. Given a standard monocular video sequence as input, our method efficiently outputs a compact, highly interpretable, and fully searchable object-level semantic Gaussian map. Rather than entangling heavy language embeddings within the mapping loop, we extract geometry independently and ground semantics through a lightweight, modular post-processing framework. Extensive evaluations on the Replica dataset demonstrate that this decoupled architecture preserves strong rendering fidelity and competitive segmentation accuracy. Crucially, by replacing dense per-Gaussian storage with modular, object-level semantic embeddings, our approach delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines. This provides a highly efficient, scalable, and practical solution for open-vocabulary 3D retrieval and question answering directly from everyday monocular video.
[CV-31] Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras
链接: https://arxiv.org/abs/2607.28293
作者: Edoardo Ragusa,Giovanni Paolo Canuti,Simone Lugani,Rodolfo Zunino,Paolo Gastaldo
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
备注:
Abstract:While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored. The paper proposes two approaches: a reformulated version of hardware-aware neural architecture search, endowed with a newly designed search space to integrate depth (D) information into small-sized deep networks, and a dedicated fine-tuning approach, including a preprocessing layer to merge depth information with RGB data and make it compatible with conventional architectures. In both cases, those methods aim to generate solutions that benefit from modern (portable) hardware accelerators and overcome existing tiny-like approaches, which often fail to tackle critical scenarios due to the severe constraints set by the supporting hardware. Extensive experiments on a pair of real-world datasets demonstrate the effectiveness of the proposed method as compared with existing solutions. The approach presented in the paper generates, in most cases, solutions that identify the Pareto optimal front to balance generalization performance and hardware requirements. The paper also describes the supporting prototype, including a Jetson Nano board and a RealSense RGB-D camera. When considering the energy profile of the device, the overall system can attain real-time performances within an energy budget that is compatible with standard batteries, such as those used in smartphones.
[CV-32] ycho: Active Abstraction with Programmatic World Models for ARC-AGI-3
链接: https://arxiv.org/abs/2607.28287
作者: Jens Lehmann,Andrei Aioanei,Sahar Vahdati
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Symbolic Computation (cs.SC)
备注: 52 pages, 18 figures, 17 tables. Open-source implementation: this https URL
Abstract:ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game’s rules, hidden state, and goal while maintaining action efficiency because every move counts. We formalize these environments as parameterized rendered deterministic Moore machines and introduce Tycho, a coding-agent system that constructs and uses game-specific models during interaction. Tycho separates actionable observations from intermediate animation, level-completion, and game-over frames. From this structured history, an agent can model, test, plan with, repair, or bypass a free-form executable hypothesis. In one matched public-set run per policy, we compare four orchestration policies on all 25 public games using Claude Opus 4.8 under matched inference budgets. Actor-requested delegation to a model builder obtains the highest observed mean Relative Human Action Efficiency (RHAE), 88.49. With this selected policy, GPT-5.6 Sol and Opus 5 both reach 100.00 RHAE and complete all 183 levels. Their game-balanced first-run human-replay midranks are 98.5 and 100.0. Opus 5 uses 61% fewer scored actions than the aggregate official human baselines. Automatic repair after verification failures produces models that reproduce observed transitions much more accurately, yet reaches only 83.07 RHAE. Transition match indicates whether a simulator reproduces observed dynamics, not whether it has identified the objective or improves the next action. Strong play also requires deciding when to construct, repair, use, or bypass a model. We call this joint problem active abstraction: generating a testable model from costly interaction and deciding when acquiring or using it is worth its cost. Comments: 52 pages, 18 figures, 17 tables. Open-source implementation: this https URL Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Symbolic Computation (cs.SC) ACMclasses: I.2.6; I.2.8 Cite as: arXiv:2607.28287 [cs.AI] (or arXiv:2607.28287v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28287 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-33] Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions ACM-MM2026
链接: https://arxiv.org/abs/2607.28285
作者: Junrui Zhang,Jiaqi Li,Yiran Wang,Liao Shen,Zhiguo Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ACM MM 2026
Abstract:Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in isolation via image inpainting or augmentation, yielding limited robustness gains. Language, as a powerful complementary modality to vision, is demonstrated to enhance the visual perception capabilities of vision-language models (VLMs) via detailed long captions. However, prior language-integrated MDE methods fail to fully harness this potential due to short text input with limited information, coarse global text feature learning, and limited language guidance during depth decoding. To address these limitations, we propose CapDepth, a novel framework for robust MDE that leverages guidance from detailed long captions to alleviate visual ambiguities in both challenging scenarios. First, we design a detailed long caption input template that explicitly conveys rich spatial relationships among multiple atom sentences. Second, a dynamic caption encoder is introduced to extract fine-grained depth-relevant text features via progressive masked attention. Finally, we propose a text-adaptive decoder that guides enhanced depth decoding with text features via stable adaptive layer normalization. Extensive experiments validate the efficacy of CapDepth, which outperforms state-of-the-art methods, achieving depth error reductions of 25.0% on non-Lambertian surfaces and 22.0% under adverse weather conditions.
[CV-34] MSCM-net: A hyperspectral image classiffcation method based on multi-scale convolution and Mamba
链接: https://arxiv.org/abs/2607.28277
作者: Jianjun Chen,Linlin Wang,Lifang Chang,Limin Huo,Shujiang Song,Yanjia Zhao,Mingwei Shao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Hyperspectral imaging is widely used in remote sensing and engineering. Therefore, research on its classification methods is crucial. While CNN and Transformer-based methods have advanced, they still face locality constraints and high computational complexity. To address these issues, we propose an innovative hyperspectral image classification model, MSCM-net. Specifically, first of all, a model architecture combining multi-scale CNN and Mamba is proposed. It consists of a multi-scale feature extraction module (MCSE) and multiple stacked Mamba blocks, which integrates the local feature extraction capability of multi-scale CNN and the long sequence modeling advantage of Mamba. Secondly, the proposed MCSE module consists of multi-scale convolution and SENet. Convolution kernels of different scales extract local information with different receptive fields, enhancing the fusion of spatial and spectral information. Meanwhile, the SENet enables the model to automatically learn the importance of each channel in the multi-scale features. Furthermore, we also propose a dual-branch feature aggregation module, which further effectively extracts and integrates the spectral information contained in the central pixel and the spatial information in the surrounding pixels. Our model has undergone numerous experiments on three widely used benchmark datasets. The experimental results show that MSCM-net can achieve advanced classification performance while reducing computational complexity.
[CV-35] heia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
链接: https://arxiv.org/abs/2607.28269
作者: Simone Giano,Lorenzo Severini,Alessandro Galdelli,Adriano Mancini
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注:
Abstract:The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Data-Free Knowledge Distillation (DFKD). However, existing datasets in this domain either entirely lack descriptive text, such as Incidents1M, or suffer from severe text-image semantic misalignment, such as CrisisMMD. In this work, we present a novel methodology to construct and automatically validate a large-scale multimodal dataset for disaster response. Starting from the vision-only Incidents1M, we successfully recovered 100,000 images and generated high-fidelity textual descriptions using two distinct Qwen3.5 architectures: a 4B dense model and a 35B Mixture-of-Experts (MoE) model. To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline leveraging Qwen3.5-9B. By intentionally obscuring the original image from the judge, this evaluator accurately simulates the modality gap of the student model during data-free distillation. Our evaluation across 173,179 label pairs demonstrates a high semantic agreement (78.65/100) between the two architectures. Furthermore, the automated evaluation reveals a conservative captioning behaviour, characterized by a high Precision (77.6%) and low Recall (46.0%). This minimizes the false positive noise, while simultaneously exposing underlying human annotation inconsistencies in the original ground truth. This work provides a scalable, LLM-validated multimodal dataset and a reproducible framework to advance cross-modal knowledge distillation.
[CV-36] ARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
链接: https://arxiv.org/abs/2607.28261
作者: Jiwen Liu,Shujuan Li,Xiaohan Li,Zijie Meng,Xinyue Liu,Yulong Xu,Yan Zhou,Guoxin Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 5 figures
Abstract:Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: this https URL
[CV-37] Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery
链接: https://arxiv.org/abs/2607.28247
作者: Iason Tsardanidis,Alkiviadis Koukos,George Choumos,Vasileios Sitokontantinou,Charalampos Kontoes
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: This paper has been accepted for presentation at the 45th EARSeL Symposium, Athens, Greece
Abstract:Accurate and scalable parcel-level agricultural monitoring remains challenging because satellite Earth Observation alone provides only an overhead perspective of agricultural parcels, while optical observations are further affected by cloud-induced temporal gaps. This paper presents Space2Ground 2.0, a multi-source framework integrating Sentinel-1 SAR and Sentinel-2 multispectral time series with geo-tagged street-level imagery acquired using vehicle-mounted cameras and shared through the Mapillary platform. A largely automated processing pipeline performs semantic filtering, image quality assessment, viewpoint-based parcel association, and dataset refinement, transforming large volumes of crowdsourced imagery into parcel-linked, analysis-ready data. Applied over Cyprus during the 2022 growing season, the pipeline produced a curated dataset of 46,050 annotated street-level images, selected from an initial collection exceeding 900,000 images and linked with satellite information for 8,581 agricultural parcels. The practical value of the dataset was assessed through parcel-level crop classification experiments using both single- and multi-source observations. The results demonstrate that street-level imagery provides complementary fine-scale visual information that enhances classification when integrated with satellite time series. Overall, Space2Ground 2.0 provides an openly available benchmark dataset and a reproducible methodology for multimodal agricultural monitoring, with potential applications in visual verification, reduced reliance on costly field inspections, and data-driven agricultural policy implementation.
[CV-38] EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
链接: https://arxiv.org/abs/2607.28243
作者: Zexuan Yan,Yuzhou Wu,Yue Ma,Zonghang He,Kaibo Yin,Xiaobing Tu,Yinggui Wang,Jinkui Ren,Xiantao Zhang,Shijian Wang,Jinghong Liu,Linfeng Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: project page: this https URL
Abstract:Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
[CV-39] Qwen -UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
链接: https://arxiv.org/abs/2607.28227
作者: Hanzhang Zhou,Panrong Tong,Xu Zhang,Quyu Kong,Chenglin Cai,Tianyu Xia,Gongjie Zhang,Jianan Zhang,Long Li,Long Chen,Lei Wang,Gaole Dai,Pengxiang Li,Liangyu Chen,Yue Wang,Steven Hoi
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively. Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.28227 [cs.AI] (or arXiv:2607.28227v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28227 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-40] FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
链接: https://arxiv.org/abs/2607.28225
作者: Haoqing Wang,Xingrun Xing,Wei Xia,Ziheng Li,Yehui Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multimodal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the tool crops the wrong region or misses the queried target), yet the call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model leans on prior knowledge or the original image rather than the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation and thus ensure train-test consistency, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from main agent, eliminating any dependence on an external model at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while markedly improving tool faithfulness. The homepage is at this https URL.
[CV-41] Scaling Vision-Language Models Is Not Enough to Mitigate Bias
链接: https://arxiv.org/abs/2607.28211
作者: Ioannis Sarridis,Ioannis Kompatsiaris,Symeon Papadopoulos
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases). Across these settings, the Spearman correlation between model scale and performance weakens as evaluation shifts from ImageNet ( \rho=0.68 ) to single-attribute ( \rho=0.48 ) and further to multi-attribute ( \rho=0.05 ) bias benchmarks. In contrast, properties of the training data (size and quality) show more consistent relationships with worst-group accuracy across both bias benchmarks. Notably, curated datasets yield improvements of up to 25% over uncurated alternatives at a comparable scale. Finally, the effect of architectural choices (e.g., patch size, image resolution) is highly context-dependent, varying with the nature of the benchmark, including the type of bias and its spatial distribution within images.
[CV-42] UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis
链接: https://arxiv.org/abs/2607.28198
作者: Hui Zhang,Julian Ferchow,Jie Song,Mirko Meboldt
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Many dexterous manipulation tasks require the object to remain securely held throughout the interaction. From the perspective of hand-object relational motion, such manipulation comprises four canonical skills: grasping, relocation, in-hand rotation, and in-hand translation. Human hands flexibly compose these skills to accomplish complex tasks. Existing approaches, however, model these skills separately with skill-specific action constraints, objectives, or even dedicated hand morphologies, which breaks the compatibility and continuity required for long-horizon composition. In this work, we present a unified framework that models all four skills in a single formulation that shares the same state and action spaces and a common objective structure. This formulation enables straightforward distillation of a single cross-skill policy that performs strongly on every skill, generalizes to unseen objects, stays robust to disturbances, and chains skills seamlessly into long-horizon manipulation. The framework also transfers effectively across different hand morphologies. Overall, our results suggest that different dexterous manipulation skills can be viewed as instantiations of a shared task formulation, revealing the intrinsic consistency across different behaviors.
[CV-43] hink with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain
链接: https://arxiv.org/abs/2607.28186
作者: Haiyang Wu,Weiliang Mu,Zhuofei Du,Dandan Zhong,Kaijie Shi,Haifeng Li,Chao Tao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Existing farmland remote sensing image (FRSI) segmentation follows a “Think with Intra-Image” paradigm, assuming that the current image contains sufficient visual evidence for reliable segmentation. Yet farmland appearance varies with phenology and spatial context and is often confused with other land-cover, making instantaneous, local observations inadequate. Thus, segmentation ambiguity stems not only from limited model representation, but more fundamentally from the required spatio-temporal information lying beyond the current image. Based on this insight, we redefine FRSI segmentation from an information bottleneck perspective as a dynamic decision process driven by task-relevant extra spatio-temporal information gain. We further propose FarmSeeker, a dynamic FRSI segmentation agent that identifies ambiguous regions, reasons about their causes, and queries extra spatio-temporal information on demand for accurate segmentation. To evaluate FarmSeeker, we construct GSFS-Bench, the first global-scale, high-resolution FRSI segmentation benchmark that supports reasoning-querying. Experiments show that FarmSeeker achieves more stable segmentation performance than existing methods. The project is publicly available at: this https URL
[CV-44] S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
链接: https://arxiv.org/abs/2607.28164
作者: Hail Song,Seokhwan Yang,Jiwon Yang,Woojin Cho,Woontack Woo
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 15 pages, 12 figures
Abstract:We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation module and strategies for animating 3D Gaussian Splatting (3DGS). While single-image head avatar reconstruction is crucial for lifelike Virtual Reality (VR) applications, existing approaches often struggle to preserve 3D consistency under unseen viewpoints. S-Avatar addresses this limitation through a three-stage pipeline. First, a high-resolution 3DGS is synthesized directly from a single image using a diffusion-based Gaussian splat generation module. Next, the parametric head model FLAME is aligned with the generated 3DGS by optimizing its parameters and spatial transformations. Finally, to adapt the 3DGS to FLAME variations, we construct a binding template that encodes the spatial relationship between the initial splats and FLAME. The dynamic 3D head avatar can then be rendered in real time by deforming the 3DGS with the binding template. By combining diffusion-guided canonical 3DGS generation with FLAME-based control, our method achieves efficient and accurate reconstruction with enhanced 3D consistency. Evaluations on public datasets demonstrate that S-Avatar outperforms state-of-the-art methods in novel-view and expression generation, achieving superior realism and consistency. Consequently, our approach represents a significant advance in accessible avatar creation, applicable to a wide range of VR/AR applications. The project page is available at this https URL.
[CV-45] OPLD: On-Policy Latent Distillation for Multimodal Reasoning
链接: https://arxiv.org/abs/2607.28154
作者: Shoutai Zhu,Tianyang Xu,Bin Sun,Mingyuan Xu,Yu Liu,Qinzhen Guo
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.
[CV-46] What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
链接: https://arxiv.org/abs/2607.28148
作者: Longxia Gao,Linan Wang,Yuhe Han,Junze Geng,Meng Zhang,Hanqing Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 30 pages, 8 figures, 9 tables
Abstract:Deep learning has shown promise for automated tongue diagnosis in traditional Chinese medicine (TCM), yet the design space remains underexplored. We conducted a systematic ablation study spanning 20+ model versions under rigorous 5-fold cross-validation on TongueDx2 (5,109 images, 976 expert-annotated) and a merged dataset of 11,101 samples. We compared six backbone architectures, four loss functions, five augmentation strategies, and six training strategies. The best 976-sample model achieved weighted-F1 of 0.6625 using ConvNeXt-Tiny with restrained augmentation and weak-group ensemble, while the best 11,101-sample model reached weighted-F1 of 0.7761. Six key design principles emerged: (1) ConvNeXt-Tiny offers optimal parameter efficiency; (2) BCE substantially outperforms Asymmetric Loss (+2.7%); (3) restrained color augmentation is critical; (4) weak-group ensemble replacement (+2.1%) outperforms probability averaging; (5) data scaling yielded +20.6% improvement; (6) expanding from 13 to 45 label dimensions caused catastrophic collapse (0.78 to 0.22). These principles are generalizable to multi-label medical image classification with class imbalance.
[CV-47] Convolutional Neural Shading for High-Quality 3D Reconstruction from Multi-View Images
链接: https://arxiv.org/abs/2607.28132
作者: Juheon Hwang,Taewan Kim,Heeseok Oh,Jiwoo Kang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:We propose a convolutional neural shading (CNS), a novel pipeline to reconstruct high-quality 3D shapes from multi-view images. Several recent studies have used neural radiance fields and other neural differentiable rendering methods to understand 3D geometry. However, these approaches rely on single-point geometric information, such as positions and normals of the surface, leading to a lack of detailed local geometry. Our approach addresses the inherent limitations of single-point information by leveraging a neural shader to capture variations even in dark and textureless regions with a convolutional neural shader, resulting in far more accurate geometry predictions. Additionally, our method mitigates surface irregularities at image boundaries by introducing a fine-detail displacement network, which utilizes spatial information of surface geometry and learns fine displacement details by correlating neighboring values in the rendering coordinates. Through extensive experiments, our proposed method has demonstrated significant quality improvements in the reconstructed shapes and rendered images over current state-of-the-art methods.
[CV-48] Collaborative Feature Aggregation for Face Super-Resolution and Robust Re-Identification
链接: https://arxiv.org/abs/2607.28130
作者: Juheon Hwang,Taewan Kim,Jiwoo Kang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:We propose a novel collaborative approach for face super-resolution (SR) and robust person re-identification from sequential or multi-view facial images. Traditional SR methods often suffer from blurring and distortion in faces recovered from poor-quality images due to low resolution. Image- and video-based facial SR methods using facial landmarks or segmentation also have similar challenges. To overcome these limitations, we leverage multiple correlated facial observations, across time or viewpoints, by introducing a transformer-based collaborative feature aggregation method that unifies identity features from multi-sequence or multi-view data. This allows faces in multiple sequences of an individual to contribute to accurately estimating common facial features. Furthermore, we propose a cascade SR network to progressively restore the high-resolution image of the target’s face with gradual facial feature unification. The unified identity representation is further utilized in person re-identification scenarios, enabling accurate matching even under severe image degradation. The exhaustive experimental results and comparisons show that our method outperforms other state-of-the-art methods, demonstrating consistent improvements in both face super-resolution and re-identification performance. Our work highlights the effectiveness of joint identity reconstruction and progressive image restoration from multiple facial inputs in enhancing downstream visual recognition tasks.
[CV-49] owards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging
链接: https://arxiv.org/abs/2607.28125
作者: Yiheng Xiong,Luisa Gallée,Daniel Santak Wolf,Heiko Hillenhagen,Michael Götz
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Numerous unsupervised domain adaptation (UDA) algori-thms exist, but for clinical practice, selecting the best-suited one along with proper hyperparameters often remains unclear, as the unlabeled deployment (target) domain prevents direct evaluation. We propose a label-free criterion that jointly selects the algorithm and hyperparameters for UDA. Given a pool of candidate models from multiple algorithms trained with different hyperparameters, our approach scores each candidate against an agreement reference, and selects the one with the highest score. The agreement reference is constructed in two levels without using target labels. First, we leverage multiple label-free selection signals, using each to nominate a model within every algorithm. Second, the nominated models are aggregated across algorithms to form a reference prediction for each unlabeled target sample. The candidate whose predictions agree most with this reference is then selected for deployment. Experimental results on four brain MRI and four chest X-ray datasets across seven clinically relevant transfer scenarios show that our method achieves better selection performance than other methods and remains effective across different algorithm pools. Our approach takes a step towards practical, label-free algorithm selection for clinical deployment of UDA.
[CV-50] mmRadarTwin: A Measurement-Calibrated Signal-Level Digital Twin Platform for Indoor mmWave Radar
链接: https://arxiv.org/abs/2607.28108
作者: Jianyi Zhou,Chenghao Zhang,Yanli Li,Dong Yuan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 7 figures, 4 tables
Abstract:Indoor mmWave radar perception is difficult to reproduce because measured range-angle responses depend on scene geometry, material response, multipath, hardware conventions, and signal processing. Existing ray-tracing and digital-twin tools often expose rendering, channel, or path-level quantities, while radar sensing requires complex signal products that can be processed and compared in the same domain as real FMCW measurements. We present mmRadarTwin, a signal-level and path-attributed digital-twin platform for indoor mmWave radar. mmRadarTwin links a real radar measurement branch with an Unreal Engine scene-simulation branch through a shared receive-channel and range-angle processing interface. The simulator writes complex multi-channel receive grids and exports per-path contribution records that identify the actor, material tag, propagation event, and output-bin support of each simulated return. We evaluate mmRadarTwin in an office deployment using a commodity monostatic mmWave radar and mobile scene-capture hardware. Across 154 measured poses spanning 22 radar locations, the current physics-only path-basis simulator recalls 70.8% of measurement-active geometry-supported response regions in the central usable field of view while exposing residuals caused by weak or missing path support, shifted responses, unsupported anchors, and missing physical mechanisms. Rather than claiming complete radar-map reconstruction or cross-room generalization, mmRadarTwin establishes a practical systems workflow for constructing, comparing, and diagnosing indoor radar digital twins.
[CV-51] GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios
链接: https://arxiv.org/abs/2607.28073
作者: Yiming Xu,Jihua Kang,Chunsai Du,Qifan Zhang,Wangqiu Zhou,Yiting Wu,Tianqi Li,Qi Song
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered by three major challenges: (1) the scarcity of datasets for complex, logic-rich diagrams; (2) the absence of explicit layout priors, which leads to chaotic spatial arrangements; and (3) the lack of fine-grained visual feedback to validate rendered outputs and correct aesthetic defects. To address these challenges, at the data level, we introduce DocMeetSVG-100K, a large-scale SVG dataset tailored for document authoring and meeting review scenarios. At the model level, we propose GVR-Coder, a novel framework designed to generate high-quality logical diagrams from lengthy professional texts. Specifically, we adopt a curriculum-driven rejection sampling fine-tuning to progressively enhance the model’s capability in modeling complex structures, while explicitly incorporating layout constraint knowledge during training. In addition, we introduce reinforcement learning from dual rendering feedback, a mechanism that provides implicit feedback through reward signals to jointly optimize structural complexity and visual aesthetics. Furthermore, we design a generate-verify-repair agent loop, which improves generation quality through explicit, fine-grained feedback and targeted refinement. Extensive experiments demonstrate that GVR-Coder outperforms competitive baselines and reliably produces logically coherent and visually appealing diagrams. Code and data are available at this https URL.
[CV-52] BladeYOLO: Wind Turbine Blade Defect Detection with Limited Annotations and Weak-Saliency Awareness
链接: https://arxiv.org/abs/2607.28065
作者: Yabin Xu,Fangtao Zhang,Fan Wang,Zhan Wang,Honghua Chen,Mingqiang Wei,Haoran Xie,Sam Kwong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to IEEE TGRS, Code: this https URL
Abstract:Wind turbine blade defect detection remains highly challenging in real-world inspection scenarios due to limited on-site data and the subtle visual characteristics of defects. In practice, blade defects are often small-scale, low-contrast, and difficult to distinguish from complex backgrounds, which significantly limits the robustness of existing detectors. To address these challenges, we propose BladeYOLO, a defect detection framework for wind turbine blades. Specifically, we integrate a Vision Transformer (ViT) backbone initialized with DINOv3 self-supervised pre-trained weights into YOLOv12-L, enabling the transfer of large-scale generic visual priors to blade defect detection and improving feature representation under limited training annotations. To enhance the perception of subtle defects, we further develop a Mamba-guided Weak-Defect Enhancement module, which consists of a Detail-Enhanced Multi-scale Branch for preserving high-frequency structural cues and a Cross-Mamba module for progressively propagating high-level semantic guidance to shallow features. In addition, we introduce a lightweight Style-Injector module that captures environment-related style information via Fourier decomposition and injects it into selected ViT self-attention layers, thereby improving robustness against environment-induced appearance variations. Extensive experiments demonstrate that BladeYOLO achieves superior performance on the WTBlade-Defect dataset, with additional annotation-budget experiments showing its favorable performance under reduced training annotations. Evaluation on the public Wind Surface Defect dataset further provides supportive evidence for the cross-dataset robustness of BladeYOLO. In particular, on this public dataset, BladeYOLO outperforms the best competing method by 3.5% in mAP _50 and 2.5% in mAP _50-95 .
[CV-53] Landmark shape spaces with induced metrics
链接: https://arxiv.org/abs/2607.28064
作者: Sarang Joshi,Peter W. Michor,Stefan Sommer
类目: Computer Vision and Pattern Recognition (cs.CV); Differential Geometry (math.DG)
备注:
Abstract:We present a unification of Kendall’s landmark shape spaces, where rigid motions are factored out and scale fixed on landmark configurations equipped with Euclidean geometry, with landmark configuration spaces carrying Riemannian metrics descending from right-invariant Sobolev metrics on the diffeomorphism group. The resulting new landmark shape spaces achieve the defining properties of both approaches: The regularity of the descending metric prevents landmarks from colliding, the metric is defined in the ambient space independent of the number of landmarks, local rigid transformations are preserved, global rigid motions are removed, and scale fixed. To achieve this, we define a particular Sobolev-type operator, the screened elasticity operator, whose null-space consists exactly of the rigid motions, we show how this operator descends to achieve the desired geometry, and we present approaches to solving matching problems and computing geodesics numerically. The resulting construction allows the use of landmark configuration spaces with sufficiently regular metrics in applications while retaining the shape invariances that are a hallmark of Kendall’s shape spaces.
[CV-54] mporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
链接: https://arxiv.org/abs/2607.28058
作者: Henglin Liu,Fangyuan Kong,Jing Wang,Yizhou Lin,Nisha Huang,Chang Liu,Xintao Wang,Pengfei Wan,Kun Gai,Xiu Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: project page: this https URL
Abstract:Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.
[CV-55] ongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
链接: https://arxiv.org/abs/2607.28039
作者: MD Wahiduzzaman Khan,Mingshan Jia,Xiaolin Zhang,En Yu,Kaska Musial-Gabrys
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.
[CV-56] Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
链接: https://arxiv.org/abs/2607.28032
作者: MD Wahiduzzaman Khan,Mingshan Jia,Xiaolin Zhang,En Yu,Kaska Musial-Gabrys
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.
[CV-57] MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images
链接: https://arxiv.org/abs/2607.28030
作者: Farzaneh Seyedshahi,Kai Rakovic,Adalberto Claudio Quiros,John LeQuesne,Ke Yuan
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches typically rely on handcrafted features, such as marker intensity statistics or cell-type proportions, which often fail to scale or generalise across cohorts with heterogeneous marker panels. We introduce MUL-T, a lightweight transformer framework that reframes tissue architecture as a masked contextual prediction task over discrete cell tokens. By learning contextualised [CLS] embeddings without task-specific supervision, the model captures higher-order cellular interactions while remaining computationally efficient. We evaluate MUL-T on several clinically relevant downstream tasks, including core-level tumour pattern classification, patient-level grading, PD-L1 positivity prediction, and cross-dataset treatment response prediction. Across tasks, MUL-T consistently outperforms classical feature-based baselines and achieves performance comparable to a foundation ViT model, despite substantially fewer parameters and lower training cost.
[CV-58] ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression
链接: https://arxiv.org/abs/2607.28020
作者: Shuhan Ye,Hongbin Yu,Chenqi Kong,Pingchuan Ma,Chong Wang,Jun Wan,Qixin Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Learned video compression relies on accurate temporal modeling to remove redundancy between adjacent frames. However, most existing codecs infer motion solely from discretely sampled RGB frames, making their estimates vulnerable to fast motion, blur, occlusion, weak texture, low illumination, and abrupt brightness changes. Event cameras asynchronously capture fine-grained intensity changes between RGB timestamps and therefore provide complementary evidence about inter-frame dynamics. We propose ENCORE, an Event-Assisted Complementary Motion Refinement framework for learned video compression. ENCORE first employs Complementary Motion Representation (CMR) to decompose aligned RGB-event features into common and modality-specific motion representations. Spatial Energy and Redundancy-Informed Calibration (SERIC) then identifies event-specific responses that are active and novel relative to RGB, suppresses weak or redundant evidence, and predicts a candidate flow correction. Finally, Energy-Aware Routing (EAR) determines where and how strongly the correction should refine the RGB flow. Events serve solely as an auxiliary modality for motion modeling, while RGB remains the only coding and reconstruction target. Experiments on BS-ERGB, HQ-EVFI, and CED demonstrate consistent gains across datasets and GOP lengths. On BS-ERGB, ENCORE achieves up to 20.80% PSNR-RGB and 22.14% MS-SSIM-RGB BD-rate savings, while retaining clear improvements on the other two datasets.
[CV-59] Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
链接: https://arxiv.org/abs/2607.28007
作者: Sweta Banerjee,Alireza Teimoury,Nils Porsche,Alexandra K. Stoll,Viktoria Weiss,Niklas Hargarter,Jonas Ammeling,Thomas Conrad,Christoph Stroblberger,Christopher Kaltnecker,Robert Klopfleisch,Christof A. Bertram,Katharina Breininger,Marc Aubreville
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Pathology foundation models (FMs) are models trained on vast amounts of typically unlabeled data and have been shown to yield regularized latent spaces that can be used effectively in downstream classification tasks. This is also true for the classification of mitotic figures vs. other cells. However, it is so far unclear if the latent space of current FMs provides features that are discriminant and spatially suitably resolved to also serve as a backbone for dense object detection paradigms. In this work, we investigate this question for common current pathology FMs (UNI, UNI2-h, Virchow, Virchow2, H-optimus-0, H-optimus-1) and compare their performance against a fully end-to-end trained baseline based on a ResNet50 architecture. We combine FM backbones with representatives of single stage, dual stage and self-attention-based detectors (RetinaNet, Faster R-CNN, Deformable DETR respectively) on the multi-domain MIDOG++ dataset, and on the TUPAC16 dataset as an out-of-domain case. We show that the H-optimus-0 and Virchow models yielded competitive performance, indicating that the latent spaces of current FMs, all trained on image-level self-supervision, are suitable for direct mitotic figure detection and may be slightly more robust on our out-of-domain test case. All code is made available publicly at this https URL.
[CV-60] Deep learning-based hierarchical insect classification using camera trap imagery
链接: https://arxiv.org/abs/2607.28005
作者: Zaki Mahfoud,Juan A. Chiavassa,Simon Walther,Florian Haselbeck,Ehsan Yaghoubi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Declining insect populations make reliable biodiversity monitoring increasingly urgent, yet monitoring of insect biodiversity is hampered by a lack of standardised data and by costly and time-consuming manual identification by expert entomologists. Deep learning-based image classifiers, processing data from automated non-lethal camera traps, have the potential to transform and scale insect biodiversity monitoring. However, challenges remain in acquiring expert-annotated datasets, developing model architectures that generalise well across diverse taxonomic levels and training models on highly imbalanced data. Hierarchical data also benefits from designing models that default to higher-confidence, coarser-level predictions, when uncertain about finer taxonomic levels. In this paper we address these challenges with a deep learning-based hierarchical classification model. First, we present a manually curated, long-tailed dataset of around one million images of insects, extracted from 1,801 camera-trap video recordings and annotated with a five-level, 34-class hierarchy. Further, we adapt a hierarchical classification model architecture to a five-level variable-depth hierarchy, with class-balanced weighting. Our model improves on non-hierarchical classifiers by leveraging biological taxonomy to extract granularity-specific visual features and makes hierarchy-consistent predictions to the deepest taxonomic level that meets a confidence threshold (T = 0.6). Our model achieved a per-level accuracy of 80-99% on test data, across five levels of hierarchy. Furthermore …
[CV-61] ViP-Rig: Visual-Prompted Controllable Rigging
链接: https://arxiv.org/abs/2607.27982
作者: Zihan Qin,Mingze Sun,Yifan Mao,Jialei Xu,Jingfeng Guo,Changrong Hu,Wenbo Zhao,Junjun Jiang,Xianming Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 4 figures. Zihan Qin and Mingze Sun contributed equally. Xianming Liu is the corresponding author
Abstract:Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explicit control over the resulting skeleton and deformation behavior. In this work, we present ViP-Rig, a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones. Specifically, ViP-Rig consists of two stages, Skeleton Generation and Skinning Prediction. In the first stage, the skeletal sketch is processed by the Dense-to-Compact Visual Prompt Encoding to produce compact, fixed-length conditioning tokens. The resulting tokens are injected into a frozen pretrained autoregressive generator through gated adapters to control joint placement and branching structure while preserving the generator’s geometric prior. In the second stage, the rigidity map is processed using the same visual encoding design, while the pretrained skinning backbone remains frozen. The resulting tokens are symmetrically injected into the point and joint streams to modulate point-joint compatibility and the resulting skinning weights. Experiments on Articulation-XL2.0 and zero-shot evaluation on ModelsResource show that ViP-Rig more accurately recovers target skeletons and skinning weights than geometry-conditioned baselines under prompt-guided evaluation. Qualitative results further demonstrate explicit and localized control in both prompt-first rigging and result-guided editing.
[CV-62] Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
链接: https://arxiv.org/abs/2607.27974
作者: Danilo Danese,Angela Lombardi,Tommaso Di Noia
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The ASNR-MICCAI BraTS Local Synthesis (Inpainting) task asks for the anatomically plausible completion of healthy brain tissue within a masked region of a T1-weighted MRI, providing a tumor-free anatomical reference for downstream analysis. As the task is scored by distortion metrics (SSIM, PSNR, MSE), we build a deterministic regression model and focus on giving it inductive biases tailored to inpainting. Our network follows the U-DiT principle of performing self-attention on a downsampled token grid: a volumetric encoder-decoder imports long-range context through a downsampled global self-attention block with three-dimensional rotary position embeddings, while convolutions and skip connections preserve high-frequency detail. Two ideas drive our results. First, we constrain the attention so that occluded (“void”) tokens attend only to known-healthy tokens of the same volume, with a learned bias toward each query’s contralateral homologue, forcing the completion to be inferred from observed anatomy rather than from other unknown regions. Second, we add a contralateral-symmetry input that supplies the mirrored healthy hemisphere as a patient-specific prior; since the brain is approximately bilaterally symmetric and lesions are typically unilateral, this prior improves the distortion metrics at matched structural similarity. On the official BraTS-2026 validation leaderboard our submission reaches a mean healthy-region SSIM of 0.864 , PSNR of 24.7 ,dB and MSE of 4.6\times10^-3 over 219 cases. We further analyse the residual smoothness inherent to distortion-optimal regression and discuss its implications for anatomical realism.
[CV-63] FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection
链接: https://arxiv.org/abs/2607.27969
作者: Haotian Zhang,Hao Chen,Han Guo,Zhengxia Zou,Zhenwei Shi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at each spatial location over the entire observation period using a single change category associated with the final observation. This implicit single-change assumption limits their ability to characterize regions of recurrent change closely related to human activities. To address this limitation, we introduce Urban Building Dynamics Detection (UBDD), which identifies building-change dynamic footprints, i.e., the temporal intervals in which changes occur, from multi-temporal imagery and produces pixel-wise classification masks. For regions undergoing two or more changes, UBDD introduces an independent multi-change class for unified representation, thereby enabling unified modeling of single- and multi-change processes. Furthermore, we propose FootprintNet, which abstracts building-change processes as interactions between latent states and actions, and imposes state-action transition constraints to guide the learning of causally coherent change trajectories. It further exploits temporal change-boundary cues to enhance feature contrast across boundary sides, thereby improving the discrimination among different dynamic footprints and enabling accurate detection of dynamic footprints. Moreover, we introduce the Building Change Dynamics Score (BCDS) to address the inability of conventional metrics to reflect the temporal proximity between predicted footprints and labels. It evaluates predictions according to their preservation of change semantics and temporal offsets from the corresponding labels. Extensive experiments on TSCD, MUDS, and WUSU demonstrate that FootprintNet outperforms current state-of-the-art methods. The code is available at this https URL.
[CV-64] LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
链接: https://arxiv.org/abs/2607.27952
作者: Feng Yang,Xinrui Ju,Keyang Zhang,Xiandong Meng,Rongqun Lin,Howard Leung,Shiqi Wang,Haoliang Li,Chris Xing Tian
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud MLLM. However, dense visual-token sequences increase cloud-side inference costs. Existing pruning methods mainly target centralized inference: vision-driven methods can operate before cloud execution but are typically query-agnostic, whereas query-guided methods often rely on internal states of the target MLLM and cannot determine token relevance before transmission. Compact guidance models offer an alternative, but existing designs may require costly attention aggregation or auxiliary generation. We propose LAST, a training-free framework for query-dependent visual token pruning in edge-cloud collaborative MLLM inference. LAST uses a compact edge-side VLM as a guidance proxy and derives a lightweight importance signal from the last query token’s attention to visual tokens. Under causal attention, the last query token can attend to the full visual sequence and the entire query context, enabling query-aware pruning without cloud-model access, autoregressive generation, or costly aggregation over multiple query positions. LAST then retains a diverse set of query-relevant visual tokens under a fixed token budget. We evaluate LAST on 11 multimodal benchmarks under multiple token budgets against pruning methods with different guidance strategies. Experiments show that LAST consistently achieves the strongest performance, preserving 95.4% of the full-token accuracy while retaining only 12.5% of the visual tokens, with low edge-side selection overhead and reduced cloud-side computation.
[CV-65] ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
链接: https://arxiv.org/abs/2607.27927
作者: Dongfu Yin,Rourou Su,Cong Zhao,Fei Yu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Reflection symmetry detection remains challenging due to interference from asymmetric regions and arbitrary orientations of symmetric patterns. Asymmetric regions introduce background clutter that disrupts symmetric pattern matching, whereas conventional convolutional neural networks lack rotation equivariance, leading to inconsistent feature representations under rotational transformations. To address these issues, we propose an Asymmetric Region Denoising (ARD) module and a Rotation Equivariant Feature Similarity Matching (REFSM) module. The ARD module suppresses asymmetric interference to refine symmetric patterns, while the REFSM module enhances rotation equivariance through feature similarity matching between original and rotated images. Specifically, our dual-input REFSM framework leverages rotation loss to maximize consistency between the score maps of original and rotated images, thereby enabling precise prediction of rotation-equivariant symmetry axes. Furthermore, we introduce GMSYM, a new benchmark dataset that categorizes images into diverse scenarios and incorporates various interferences to address the limitations of existing reflection symmetry detection benchmarks. Extensive experiments on four standard datasets (DENDI, NYU, LDRS, SDRW) and our proposed GMSYM dataset demonstrate that our method achieves state-of-the-art performance in both accuracy and robustness.
[CV-66] ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
链接: https://arxiv.org/abs/2607.27924
作者: Dongxiu Liu,Haoyi Niu,Peng Cheng,Yuan Gao,Xirui Kang,Sangli Teng,Koushil Sreenath,Xianyuan Zhan
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturing the dynamics of physical world. We introduce Physical-Time Flow (\textbfPT-Flow), a novel approach that learns a continuous latent velocity field operating in physical time. Crucially, the underlying dynamics of sequential data are parameterized by an ordinary differential equation (ODE) embedded in a well-structured representation space. Under this paradigm, the prediction of future can be recast as temporal integration via an ODE solver in the compressed latent space. Building upon PT-Flow, we construct \textbfODEWorld, a continuous-time latent world model that is both efficient and versatile. By extracting time-variant features and enforcing ODE properties on both the dynamical representation space and the latent velocity field, ODEWorld effectively addresses the long-standing representation collapse issue in latent world model literature. This also enables high-quality image reconstruction even after long-horizon prediction. Moreover, its continuous nature allows for arbitrary temporal resolution and even backward prediction, which is impossible for most discrete-time models. Lastly, ODEWorld can provide rich planning-oriented information to facilitate downstream policy learning. Comprehensive experiments demonstrate that ODEWorld successfully reconciles planning-conducive dynamics abstraction with visual realism, excelling in both video generation and robotic control. \hrefthis https URLProject Website.
[CV-67] One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM -Based Scene Text Spotting
链接: https://arxiv.org/abs/2607.27902
作者: Rui Tang,Wentao Yang,Peirong Zhang,Yongxin Shi,Shun Zhang,Huiguo He,Lianwen Jin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 11 figures. Accepted to ACM Multimedia 2026
Abstract:Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code will be released soon.
[CV-68] CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
链接: https://arxiv.org/abs/2607.27898
作者: Zaiyan Zhang,Qiangqiang Yuan,Jie Li,Ziyang Lihe,Yu Wan,Yuzeng Chen,Xin Su,Liangpei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ISPRS Journal of Photogrammetry and Remote Sensing
Abstract:Remote sensing images acquired by unmanned aerial vehicles (UAVs) and satellites are often degraded by adverse weather, illumination variation, and imaging artifacts, which may co-occur and jointly induce global distribution shifts and local structural corruption. Although All-in-One image restoration offers an appealing unified alternative to task-specific pipelines, existing methods still suffer from weak or implicit degradation cues and parameter redundancy caused by full-rank multi-expert designs with overlapping restoration behaviors. We propose CoRE-UIR (Common and Residual Experts for Universal Image Restoration), a prior-guided global-local framework centered on the Common-and-Residual Expert Block (CoRE). CoRE explicitly decomposes restoration capacity into a common dense expert for degradation-invariant restoration and low-rank residual experts for degradation-specific compensation, enabling adaptive specialization without redundant expert replication. Built on this design, Degradation Prior Embedding (DPE) adapts frozen CLIP features into an explicit restoration-oriented prior, while Global Feature Modulation (GFM) aligns global feature statistics before local residual compensation. We also construct MDVD-108K (Multi-Degradation VisDrone), a large-scale UAV restoration dataset covering both single and compound degradations, together with a real-world test set. Extensive experiments on multiple datasets show that CoRE-UIR improves the overall average PSNR by 1.05 dB while running 11.83 \times faster and reducing peak memory by 85.3% relative to the strongest baseline, BaryIR, thereby maintaining a favorable quality-efficiency trade-off. Evaluations on downstream tasks and unseen degradation also validate the generalizability of CoRE-UIR. The code and dataset will be released at this https URL.
[CV-69] Unifying Adversarially Robust Model Experts in Vision-Language Models
链接: https://arxiv.org/abs/2607.27897
作者: Nguyen Duc Thai,Junhao Dong,Sua Qi Rong,Hua Yu,Yew-Soon Ong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, different fine-tuning strategies often produce specialized models with distinct robustness characteristics. Each fine-tuned model in turn thrives in some evaluation settings but falters on others, limiting their defensive capabilities. We refer to these specialized fine-tuned models as robust model experts and propose a collaborative adversarial fine-tuning framework: CARE - Collaborative Adversarial Robustness fine-tuning using Embedding alignment. CARE maintains multiple experts during training, enables knowledge exchange through embedding-space harmonization, and consolidates the learned knowledge into a single unified robust model. Experts benefit from one another while preserving their individual specializations, enabling the final model to inherit complementary robustness properties. In this paper, we demonstrate CARE on two different adversarial fine-tuning strategies with complementary robustness behaviors. Extensive experiments on classic image classification and downstream vision-language tasks display the effectiveness of our approach, with CARE being able to outperform individually learned model experts. The results suggest that collaborative learning across model experts is a promising direction for improving adversarial robustness.
[CV-70] MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
链接: https://arxiv.org/abs/2607.27895
作者: Jinpeng Hu,Erqiang Wang,Shan Wang,Zhuo Li,Peipei Song,Xun Yang,Meng Wang
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.
[CV-71] DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection
链接: https://arxiv.org/abs/2607.27882
作者: Zihao Cai,Xinghan Li,Ruiyan Yang,Xue Song,Haijun Shan,Jingjing Chen
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge acquired from previous ones. This continual learning setting is particularly challenging because forensic traces are often subtle and generator-specific, making detectors highly vulnerable to catastrophic forgetting. Existing methods primarily address this problem by stabilizing feature representations, implicitly treating forgetting as a representation-level issue. In this paper, we show that this perspective is incomplete. We demonstrate that even when feature representations remain discriminative, the decision boundary can progressively drift as the classification head is continually optimized on new domains. These two effects jointly give rise to a compound failure mode, termed Dual Degradation. To overcome this challenge, we propose DECODE, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting. Specifically, we introduce Subspace Diversity Regularization (SDR) to preserve diverse forensic representations and Closed-Form Decision Alignment (CDA) to recalibrate the shared classification head after each adapter merge without manual hyperparameter tuning. Extensive experiments on 19 generative domains show that DECODE achieves an average accuracy of 99.36% with only 0.39% forgetting, while further generalizing to 11 unseen generators with 95.36% accuracy.
[CV-72] Learning to Understand Body Language from Flight through Robust 3D Avatar Placing
链接: https://arxiv.org/abs/2607.27865
作者: Dragos Costea,Alina Marcu,Cristina Lazar,Marius Leordeanu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion. Enabling it is a lightweight geometric world model of the local scene - semantically selected anchors lifted to 3D through streaming monocular depth - in which a placement point is predicted as an affine anchor combination with provably rigid-invariant weights, and re-rendered under an SVD-fitted ground rotation. Across twelve architectures on scene- and motion-disjoint splits, training on placed data lifts mean intent accuracy by a wide margin for real, retargeted and generated motion alike, with gains confirmed on two in-the-wild scenes.
[CV-73] EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
链接: https://arxiv.org/abs/2607.27857
作者: Kaifan Zhang,Lihuo He,Yuqi Ji,Junjie Ke,Lukun Wu,Tianhao You,Xinbo Gao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Main paper with supplementary material. Code: this https URL . Dataset: this https URL
Abstract:Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEG-image models preserve. The code and complete dataset are publicly available.
[CV-74] Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation
链接: https://arxiv.org/abs/2607.27856
作者: Jinghong Liu,Yuchuan Deng,Fanping Liu,Meng Huang,Xirong Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progress, existing FS-MIS solutions span diverse paradigms but are evaluated under inconsistent settings, leaving their relative effectiveness unclear. We introduce FAME, a unified benchmark for evaluating FS-MIS solutions, covering specialists, SAM-based methods, CLIP-based methods, and MLLM-based methods. FAME contains 14,958 test samples across 7 anatomical sites, 9 imaging modalities, and 14 ROI categories, and evaluates models under zero-shot and ten-shot settings with additional assessment of target-absence recognition and generalization under covariate and semantic shifts. Our evaluation reveals several findings. First, effective few-shot segmentation depends on how models exploit support examples: direct visual adaptation generally outperforms prompt-based strategies. Second, increasing support examples improves performance only when models can effectively utilize them. Third, semantic transfer remains substantially more challenging than imaging-domain adaptation, and strong localization ability does not necessarily imply reliable target-absence recognition. We hope FAME provides a comprehensive understanding of current FS-MIS solutions and facilitates the development of more effective and reliable few-shot medical segmentation methods.
[CV-75] Simplifying Neural Networks During Training
链接: https://arxiv.org/abs/2607.27854
作者: Lorenzo Sciandra,Samuele Fonio,Roberto Esposito
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint, submitted to a journal
Abstract:Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. Recent evidence on Neural Collapse (NC) shows that class representations and classifiers exhibit highly structured geometry, while the Tunnel Effect suggests that only a subset of layers is essential for feature extraction. We combine these two perspectives and propose an NC-inspired training framework for simplifying deep networks during training. Our method monitors representation dynamics through the Inverse Fisher Criterion, a stable and efficient proxy for the variability collapse behavior, to identify both the split point between feature extraction and classification and the training stage at which simplification becomes viable. We then replace the trailing layers with a lightweight classification head and continue training the reduced model. Experiments on image-classification benchmarks across MLP, VGG, and ResNet architectures show that the proposed method achieves substantial parameter reductions while maintaining accuracy comparable to that of the full model. Code to reproduce the experiments can be found at: this https URL.
[CV-76] VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection ECCV2026
链接: https://arxiv.org/abs/2607.27843
作者: Songsong Duan,Xi Yang,Nannan Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ECCV 2026
Abstract:Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth domain. To address this issue, we propose a depth collaborative network, called VCP-DCN, to mine distinguishable multi-modality features beyond visual concealed prototype in depth domain. Specifically, VCP-DCN progressively performs multi-modality alignment, interaction, and fusion for the COD task. In the \textbfalignment stage, we propose a Separable Prototype Embedding (SPE) module to learn modality-consistency and modality-specific RGB/depth prototype tokens through prototype contrastive learning. Furthermore, we develop a Multi-modality Dual Attention (MDA) module to enhance the cross-modal feature representation through local response maps between modality-consistency RGB/depth prototype tokens and visual tokens on the \textbfinteraction stage. Finally, we design a Depth Adaptive Injection (DAI) module to adaptively measure contribution of RGB/depth features with a decision-making mechanism, which calculates similarity distance between RGB/depth modality-specific prototype tokens and modality-consistency ones on the \textbffusion stage. Extensive experiments demonstrate the effectiveness of our VCP-DCN on three authoritative datasets.
[CV-77] FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
链接: https://arxiv.org/abs/2607.27842
作者: Hanshuai Cui,Zhiqing Tang,Zhi Yao,Qianli Ma,Fanshuai Meng,Weijia Jia
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer–timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70\times over Vanilla while maintaining competitive output quality.
[CV-78] SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
链接: https://arxiv.org/abs/2607.27835
作者: Harshit Mittal,Arash Rabbani
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment planning. Encoder-decoder architectures that fuse multi-scale features through skip connections have become the dominant paradigm for this task, yet standard direct skip connections treat every spatial location equally, which leads to redundant and potentially conflicting information reaching the decoder. To overcome this problem, various gating mechanisms have been introduced, but most of them operate solely on filtering encoder information, neglecting the benefit of global contextual information from the decoder. This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module. Each SAF gate concatenates the encoder skip and upsampled decoder features, compresses them through two pointwise convolutions with an intermediate ReLU, and applies a channel-wise softmax to produce a per-pixel “heatmap of trust” that sums to unity at every spatial location, allowing the network to learn where each source is most trustworthy. The resulting fused features improve the model’s ability to detect the minority “Dead” class, which in turn enhances the multi-class panoptic quality (mPQ) on the PanNuke dataset. SAF Gating is compared against six gating alternatives including no gating, attention gates, squeeze-and-excitation, CBAM, cross-attention, and attentional feature fusion on PanNuke and MoNuSeg datasets. SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1 score compared to ungated CellViT baseline.
[CV-79] hinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
链接: https://arxiv.org/abs/2607.27830
作者: Zhongkuan Mao,Xianjie Liu,Tianyu Meng,Yidong Wang,Wenzhuo Zhao,Ronghao Xian,Yao Jiang,Fei Shen,Junfeng Fang,Yong Dai,Yi Zhang,Keren Fu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbftraining-free, single-visual-pass evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V ^* Bench, HRBench-4K, and HRBench-8K by \textit+3.1, \textit+3.0, and \textit+2.7 points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit+9.9, \textit+4.6, and \textit+5.5 points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V ^* Bench inference time by \textbf97.2% while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.
[CV-80] Sign Language Question Answering: A New Task Benchmark and Baseline for Sign Language Understanding
链接: https://arxiv.org/abs/2607.27826
作者: Shiwei Gan,Lichen Wang,Xiao Liu,Yafeng Yin,Kuizhuang Liu,Sanglu Lu,Lei Xie
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-language sentences. As a result, they provide only a limited assessment of whether a model truly understands the semantic content of SL videos. To address this limitation, \textbfwe first propose a new task, Sign Language Question Answering (SLQA), which evaluates SL understanding by requiring models to answer arbitrary natural language questions about SL videos. Unlike previous SLU tasks, SLQA provides a more flexible and comprehensive evaluation framework that assesses multiple reasoning capabilities beyond recognition and translation. To facilitate this task, \textbfwe further construct two SignQA benchmarks based on PHOENIX14T and CSL-Daily by automatically generating question-answer pairs from existing gloss and sentence annotations using carefully designed templates. The resulting datasets cover five complementary question categories, including position reasoning, structural reasoning, visual search, gloss recognition, and translation understanding. \textbfFinally, we propose a simple yet effective baseline model equipped with a Question-Conditioned Modulated Temporal Downsampling module and an in-domain knowledge transfer strategy, enabling effective knowledge transfer from existing SLU tasks while enhancing question-aware temporal feature modeling. Extensive experiments demonstrate that our baseline consistently outperforms representative vision-language models across all question categories, establishing a strong benchmark for future research on SLQA. Datasets are available at:this https URL.
[CV-81] Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
链接: https://arxiv.org/abs/2607.27823
作者: Lei Yang,Xinze Liu,Dayan Wu,Ding Wang,Hengjie Zhu,Zihao Zhang,Tianzhu Hu,Hanqi Wu,Peng Fu,Zheng Lin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still lack reliable object-level grounding diagnostics and therefore tend to apply coarse-grained interventions, which can impair visual understanding, shorten responses, and reduce coverage of genuinely grounded objects. The key challenge is thus to detect, during generation, whether each emerging object mention is supported by reliable visual evidence, so that hallucination can be mitigated selectively. Yet output confidence reflects next-token plausibility rather than visual support, allowing language priors to make absent objects appear certain. We show that the missing diagnostic evidence is encoded in an Intrinsic Grounding Signature (IGS), a distributed signed attention pattern that remains informative for such confident hallucinations. Based on IGS, we propose Verifier-Guided Decoding (VGD), a decoding framework in which a lightweight verifier examines each emerging object mention, rolls back the KV cache when the mention is identified as high risk, suppresses the object and its synonyms, and regenerates the affected continuation. Because VGD intervenes only on object mentions identified as high risk, it reduces object hallucination while preserving the model’s original visual understanding and grounded object coverage. Experiments on CHAIR and AMBER-G show that VGD achieves state-of-the-art object hallucination reduction: at @rec90, it cuts AMBER-G CHAIR by 43.6% while retaining 99.6% of grounded-object coverage, and reduces CHAIR-MSCOCO CHAIR _i /CHAIR _s by 37.0%/30.4% without shortening captions.
[CV-82] SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack
链接: https://arxiv.org/abs/2607.27811
作者: Chunpeng Wang,Yanan Shi,Zhiqiu Xia,Jidong Yang,Suo Gao,Qi Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack. SPFM-Net first employs high-ratio masking to disrupt the spatial coherence of invisible watermark signals, and then utilizes a partially fine-tuned pretrained Masked Autoencoder to reconstruct semantically consistent image from sparse observations while suppressing watermark-related information. A Multi-scale Residual Frequency Feature Interaction module subsequently aggregates watermark-related residual features across multiple receptive fields, while adaptively suppressing responses from watermark-irrelevant regions. To further capture the long-range dependencies of globally distributed watermark signals, a lightweight Mamba-based Global State-space Feature Modeling (GSFM) unit is introduced to separate watermark-related features from natural image content and suppress the remaining watermark traces. In addition, SPFM-Net is optimized using a multi-level objective that jointly imposes spatial-, frequency-, and edge-domain constraints, enabling effective watermark suppression while preserving perceptual quality. Extensive experiments on representative spatial-domain, transform-domain, orthogonal moment-based, and deep learning-based watermarking schemes demonstrate that SPFM-Net achieves a favorable trade-off between watermark attack effectiveness and perceptual fidelity.
[CV-83] LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
链接: https://arxiv.org/abs/2607.27806
作者: Zhilin Wu,Zhangkai Ni,Chengmei Yang,Longzhen Yang,Yihang Liu,Ying Wen,Lianghua He
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 23 pages, 17 figures, 7 tables. Code and data: this https URL
Abstract:In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: this https URL
[CV-84] FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
链接: https://arxiv.org/abs/2607.27800
作者: Chunpeng Wang,Yuxin Li,Xiaoyu Wang,Jidong Yang,Suo Gao,Qi Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Existing invisible watermark removal methods often struggle to accurately capture the watermark-bearing features, leading to an unfavorable trade-off between watermark suppression and perceptual fidelity. In this paper, we propose the Frequency-Decoupled Diffusion Watermark Attack Network (FDDWAN), a coarse-to-fine framework that performs watermark removal through wavelet-domain decomposition and residual diffusion refinement. In the initial stage, the Wavelet-based Frequency-domain Preliminary Attack Module (WFPAM) decomposes the watermarked image into low- and high-frequency subbands and applies frequency-specific attack strategies tailored to their respective contributions to watermark robustness and perceptual quality. In the next stage, the Frequency-domain Residual Diffusion Attack Module (FRDAM) separately models the residual distributions between the preliminarily attacked outputs and the corresponding watermark-free references during training. Rather than reconstructing the entire image, FRDAM selectively refines frequency-domain residuals, directing the diffusion process toward the remaining watermark related discrepancies while minimizing modifications to image content. Extensive experiments on CelebA and ImageNet across four representative watermarking schemes demonstrate that FDDWAN achieves a more favorable trade-off between watermark removal effectiveness and visual fidelity than conventional and learning-based attack methods.
[CV-85] CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography
链接: https://arxiv.org/abs/2607.27779
作者: Tomer Erez,Moshe Kimhi,Chaim Baskin,Ehud Rivlin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language models offer a natural interface for text-to-image retrieval, but current biomedical models are primarily optimized for report-to-image matching rather than for satisfying short clinical search queries. This creates an objective mismatch: a model may retrieve images related to words in the query while failing to satisfy the full clinical constraint, especially for conjunctions and negations such as ``atelectasis and no pneumonia.‘’ We introduce CXR-Retrieve, a structured benchmark for compositional chest X-ray text-to-image retrieval. The benchmark contains 5,159 test images from the official test-split of MIMIC-CXR-JPG and 145 textual queries spanning single and conjunction findings, both positive and negative. Relevance is defined by whether a retrieved image satisfies all asserted pathology constraints, rather than by whether it matches a paired report. We further propose a label-aware contrastive fine-tuning objective for clinical retrieval. Our method attracts image-text pairs with compatible asserted pathology constraints, including shared confirmed absences, while explicitly repelling contradictory pairs. Starting from the in-domain CXR-CLIP checkpoint, our method improves Precision@5 over CXR-CLIP by 8.5 percentage points on two-pathology conjunctions and by 22.0 percentage points on negation queries. These results show that reliable chest X-ray retrieval requires training objectives that model not only which findings are mentioned, but also how they are clinically asserted. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.27779 [cs.CV] (or arXiv:2607.27779v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2607.27779 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-86] Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation
链接: https://arxiv.org/abs/2607.27764
作者: Shuhuan Chen,Xiangyu Zhu,Weisong Zhao,Siran Peng,Tianshuo Zhang,Haoyuan Zhang,Haichao Shi,Xiao-Yu Zhang,Zhen Lei
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: \emphthe identity cues that make released faces useful for recognition supervision are also the cues that make them linkable to real individuals. A protected face should be decoupled from the original identity, yet still behave as a reliable identity sample for training. Removing these cues too aggressively may destroy the class structure needed for recognition learning, whereas preserving them too faithfully may increase source-identity linkability. We argue that this paradox stems from conflating source-aligned identity semantics with recognition-useful proxy identity geometry. The former should be suppressed to reduce linkage to private individuals, while the latter should be preserved for FR learning. Based on this insight, we propose \textbfPrivate Face Distillation, an identity-decoupling and geometry-preserving framework. It uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning. Experiments across multiple domain-shifted FR scenarios show that Private Face Distillation achieves stronger utility than the evaluated publication baselines. On IJB-C surveillance, it improves \mathrmTAR@\mathrmFAR=1\texte-3 by 3.94% over the baseline while reducing source-identity linkability. These results suggest that private FR training dataset publication should decouple source-identity correspondence while preserving proxy identity geometry.
[CV-87] DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement
链接: https://arxiv.org/abs/2607.27761
作者: Shubin Ma,Liang Zhao,Chuanye He,Zhenjiao Liu,Liang Zou,Lin Yuanbo Wu,Yu Shao
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 4 figures. Accepted by ACM Multimedia 2026
Abstract:In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view alignment problem (PVAP). To mitigate the impact of view asymmetry and irrelevant samples, this paper proposes a framework for partial multi-view clustering via dual alignment and structure enhancement (DAS-PMVC), which leverages view structure consistency and semantic relevance. Specifically, DAS-PMVC includes three parts: \textbfanchor graph structure alignment, where sample joint embedding representations with consistent latent space are derived from anchor point relationships for initial view alignment; \textbfstructure-enhanced feature learning, where the model learns view structure information through pretraining and combines multi-view graph convolutional networks to further extract deep latent features from the aligned graph structure to improve the discriminative power of representations; and \textbfa dual alignment strategy, where initial alignment is performed through the anchor graph in the pretraining phase, and contrastive learning loss and the Hungarian algorithm are introduced in the training phase to further optimize the alignment of latent features. Experimental results on various datasets demonstrate that the DAS-PMVC framework outperforms existing state-of-the-art methods in clustering performance, showcasing its effectiveness and superiority.
[CV-88] EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder ECCV2026
链接: https://arxiv.org/abs/2607.27755
作者: Jaehun Jung,Wonjun Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 18 pages, 6 figures, Accepted to ECCV 2026
Abstract:We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-mounted devices or smart glasses. The challenge of this task lies in estimating the pose information of unobserved body parts based solely on a single joint (i.e., head) trajectory. Several studies have begun to adopt head-conditioned generative models, however, such previous methods are costly and time-consuming due to the diffusion-based iterative process. As an alternative, we propose a simple yet novel method that leverages the latent space of the guidance network, which is designed as a variational autoencoder taking full-body poses as inputs. By enforcing latent distributions of this guidance network and our head-to-motion network to be similar, latent features sampled from the ‘guided’ distribution, i.e., distribution learned in our head-to-motion network, can be reliably decoded for natural representations of full-body poses even only with the head pose. One important advantage of the proposed method is that one-step sampling scheme achieves remarkably fast inference (more than 50 times faster) compared to diffusion-based approaches. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of ego-body mesh reconstruction.
[CV-89] Articulated Object Reconstruction from Rest-State Observation ECCV2026
链接: https://arxiv.org/abs/2607.27749
作者: Daeun Lee,Jaeah Lee,Woosung Kim,Haebeom Jung,Jaesik Park
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: ECCV 2026
Abstract:Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.
[CV-90] PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation ECCV
链接: https://arxiv.org/abs/2607.27729
作者: Sangmin Hong,Daniel Sungho Jung,Heewon Kim,Kyoung Mu Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: European Conference on Computer Vision (ECCV) 2026
Abstract:Point clouds are one of the most fundamental and widely used 3D representations, serving as the most basic geometric representation of 3D shapes. Nevertheless, most existing 3D printing pipelines require a watertight mesh as input, preventing the direct use of point clouds for fabrication. A common workaround is to reconstruct meshes from point clouds; however, the resulting meshes often contain geometric artifacts, such as incorrect faces or topological inconsistencies, that are difficult to repair and may lead to printing failures. To overcome these limitations, we propose PrintAnything, a novel framework that learns to produce executable 3D printing G-code directly from 3D point clouds without requiring mesh reconstruction. To enable point clouds to serve as direct input for slice-wise toolpath generation, we introduce a slice-wise point projection strategy that transforms unstructured 3D point clouds into slice-aligned 2D representations consistent with layer-by-layer nature of fused deposition modeling in 3D printing. To eliminate mesh dependency and provide a unified representation that bridges point clouds and G-code, we propose Geometric plan (G-plan) map, a compact 2D representation composed of occupancy, region, and flow maps that encode the geometric and extrusion properties required for toolpath synthesis in 3D printing. As a result, our proposed method accurately generates printable G-code directly from point clouds, enabling a practical and fully mesh-free pipeline for 3D printing. The code is publicly available at \hrefthis https URLthis https URL.
[CV-91] Calibrate Before Reason : Robust Visual Token Reduction against Semantic Drift in VLMs
链接: https://arxiv.org/abs/2607.27700
作者: Jiasheng Li,Zhong Ji,Yan Zhang,Huihui Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign critical visual cues, triggering severe semantic drift that deviates the VLM’s understanding. In this paper, we first introduce the principle of ‘Calibrate Before Reason’ to visual token reduction and propose CaRe, a training-free robust framework that calibrates compact visual representations before reasoning to preserve semantic fidelity in VLMs. CaRe consists of two mutually complementary modules: 1) Perturbation-Robust Calibration Anchoring, which identifies calibration anchors with stable model-side influence under multi-directional perturbations; 2) Confidence-Gated Token Calibration, which extracts reliable calibration signals from unselected tokens and injects them into anchors. Extensive evaluations across diverse VLM architectures and benchmarks verify that CaRe outperforms state-of-the-art token reduction baselines. While pruning 94.4% of visual tokens, our method retains 96.4% of the original full-token performance, delivering up to 2.30 times faster end-to-end inference speed relative to unpruned vanilla models.
[CV-92] RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation ACM-MM2026
链接: https://arxiv.org/abs/2607.27699
作者: Shaobo Liu,Feiqiao Mao,Shuaishuai Zhou,Yan Zhan,Weiqi Tan,Zhiqiong Lu,Zhengping Liang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 pages, 5 main-paper figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026). Includes the complete supplementary material
Abstract:We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation through self-correction. Existing MLLM-based approaches rely on single-pass open-loop inference, where the model receives visual input only once and must generate thousands of SVG code tokens without intermediate verification. This paradigm inevitably leads to geometric drift, error accumulation, and visual hallucination on complex images. RefineSVG overcomes this limitation by invoking an external rendering engine after an initial SVG generation pass to compare the rendered output against the target image. The comparison yields a multi-dimensional visual residual map (Diff-Map) that is fed back to the model as a ReAct-style correction signal, driving a targeted correction step. To support this render-observe-correct interaction, we further introduce an SVG-oriented semantic vocabulary that compresses token sequences by over 52%. A progressive training pipeline spanning supervised fine-tuning, rejection-sampling cold-start data construction, and end-to-end agentic reinforcement learning aligns the model with closed-loop visual correction. Extensive experiments show that RefineSVG consistently outperforms existing baselines in reconstruction fidelity, structural accuracy, and code this http URL is available at this https URL.
[CV-93] JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
链接: https://arxiv.org/abs/2607.27670
作者: Shawn Li,Wei Yang,Jike Zhong,Jiate Li,Jiawei Yang,You Qin,Ryan Rossi,Franck Dernoncourt,Roger Zimmermann,Yue Wang,Zhengzhong Tu,Vicente Ordonez,Mohit Bansal,Yue Zhao
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit\ours, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4 \times 4 to 16 \times 16), we find that \textbfzero-shot VLMs largely lack geometric reasoning: only one of five frontier models (GPT-5.5) exceeds random baseline on 4 \times 4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves 97% on 4 \times 4, \textbfall models collapse on larger grids: GPT-5.5 drops from 70% to near-random on 8 \times 8, and even fine-tuned models fall below 5% on 12 \times 12. This ``scaling cliff’’ suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours establishes scalable geometric reasoning as an open challenge for vision-language models.
[CV-94] Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers
链接: https://arxiv.org/abs/2607.27667
作者: Fexiang Liu,Shiye Wang,Qiang Qiu,Zheng Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 22 pages, 6 figures; includes supplementary material. Code: this https URL
Abstract:Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Confidence scores capture candidate margins, but not where the estimated signed visual readouts associated with those margins come from or how they are distributed. We study inference-time risk detection for closed visual answers using the same white-box prefill path that produces the answer. Witness Evidence Portfolios (WEP) first estimates, layer by layer, which visual contributions support or contradict the predicted candidate. It summarizes these contributions through two interpretable route families: question-related evidence provenance and signed evidence concentration. Nested grouped validation chooses the more reliable family and a sparse top-k route portfolio, which is fused with candidate confidence. WEP needs no image perturbation, decoding change, backward pass, or external verifier. Across three MLLMs and four binary-answer benchmarks, WEP improves mean error AP by 0.134. All 12 model–dataset gains are positive, and image-cluster bootstrap intervals are strictly positive on 10 pairs. WEP targets white-box closed-answer systems and uses a labeled calibration slice.
[CV-95] Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
链接: https://arxiv.org/abs/2607.27660
作者: Rishabh Iyer,Truong Pham,Anay Majee
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\citemajee2024score demonstrated that SIMs can serve as effective objectives for supervised contrastive learning. Despite their empirical success, however, the geometric and statistical properties induced by different submodular information measures remain poorly understood. In this work, we develop a unified theoretical framework connecting SIMs to classical concepts in representation learning and statistical pattern recognition. We show that Total Information (TI) objectives characterize intra-class structure: Graph Cut TI recovers within-class variance, LogDet TI recovers generalized variance and covariance volume, and Facility Location TI induces imbalance-aware separation that emphasizes rare and confusable classes. We further show that Mutual Information (MI) objectives capture complementary notions of inter-class structure: Graph Cut MI is closely related to centroid separation and Fisher-style discrimination, LogDet MI captures covariance-aware separation through Mahalanobis distance, and Facility Location MI measures nearest-mode representational overlap. We validate these theoretical characterizations using controlled synthetic experiments that independently vary variance, covariance, class imbalance, class separation, and multimodal overlap. Across all settings, the empirical behavior closely matches the proposed theory. Our results provide the first unified geometric and statistical understanding of submodular information measures and offer principled guidance for selecting and designing SIM-based objectives for representation learning. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2607.27660 [cs.LG] (or arXiv:2607.27660v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27660 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-96] Learning Color Grading No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement
链接: https://arxiv.org/abs/2607.27659
作者: Chuanzhi Xu,Ziyuan Tao,Jean Julien KNell,Yanrong Chen,Haolan Guo,Xuanhua Yin,Adnan Mahmood,Weidong Cai
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
备注:
Abstract:Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings that are unsuitable for centralized collection. The task must infer preference from sparse, heterogeneous feedback and translate it into natural-looking color transformations on resource-constrained user devices. We introduce FedPAIE, a federated personalized aesthetic image enhancement framework for user-adaptive color grading without centralizing raw photos or ratings. FedPAIE trains a lightweight dual-cue aesthetic scorer, calibrates it into a personalized scorer on a small local support set, and freezes it to guide regularized adaptation of a lightweight CLUT enhancer from unpaired local photographs. Fidelity constraints and an excess-gap penalty regularize scorer-guided adaptation to limit proxy-score over-optimization while preserving content and natural appearance. Training remains lightweight throughout the pipeline: scorer learning updates at most 0.787M parameters, enhancer adaptation updates 0.265M, and inference retains only a 0.293M-parameter personalized enhancer. Experiments on MIT-Adobe FiveK and Flickr-AES demonstrate effective open-world personalization and a favorable balance between user preference and image fidelity. FedPAIE thus connects decentralized preference learning with efficient personalized image transformation without requiring paired user retouches.
[CV-97] MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
链接: https://arxiv.org/abs/2607.27637
作者: Wenjie Zhu,Yabin Zhang,Wenjun Zeng,Lei Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: project page: this https URL
Abstract:Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.
[CV-98] 4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
链接: https://arxiv.org/abs/2607.27634
作者: Renlong Wu,Haoran Chen,Yuxiang Wei,Xiaowei Jin,Wangmeng Zuo,Hui Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 14 pages
Abstract:Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.
[CV-99] BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement
链接: https://arxiv.org/abs/2607.27628
作者: Mingzhe Lyu,Jinqiang Cui,Hong Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied across scenes. Manually selecting the best configuration, however, can be time-consuming and not always practical. Peak signal-to-noise ratio (PSNR) is the natural fidelity criterion for automating parameter selection, yet it requires a ground-truth reference that is typically unavailable. To our knowledge, no learning-based method addresses no-reference PSNR prediction for low-light image enhancement; the natural surrogate, no-reference image quality assessment (NR-IQA), targets perceptual quality rather than signal fidelity, and all seven baselines we test achieve 0% top-1 selection accuracy on our benchmark. With paired training data, the ground-truth PSNR is analytically computable, providing exact supervision without a separate teacher network. Building on this, we propose BlindPSNR, a lightweight no-reference network that fuses the enhanced image with the degraded low-light input via windowed cross-attention and estimates PSNR through heteroscedastic regression. While a scalar-regression baseline achieves top-1 accuracy of 54.4%, BlindPSNR raises this to 89.5% with regret dropping from 1.62 dB to 0.026 dB, and generalizes to unseen datasets (SRCC = 0.61-0.67).
[CV-100] MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging ACM-MM26
链接: https://arxiv.org/abs/2607.27620
作者: Jianwei He,Kailin Lyu,Junhao Dong,Long Xiao,Wenjie Hou,Jingze Lu,Di Wu,Lin Shu,Jie Hao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: accepted by ACM MM 26
Abstract:Deep learning has shown strong potential in medical image analysis, but most existing methods rely on large-scale annotations and a closed-world assumption that rarely holds in clinical practice. Although Generalized Category Discovery (GCD) has advanced rapidly on natural images, it remains underexplored in medical imaging. To address this issue, we propose MedXplore, a unified framework for reliable and unbiased medical GCD, optimizing from both perceptual and decision levels. Specifically, at the perceptual level, taking a frequency domain perspective, Frequency-SNR Adaptive Attention and Consistency (FAAC) performs learnable full-spectrum filtering and global-local energy contrast activation to not only highlight local abnormal signals relative to the global context, but also provide reliable semantic anchors for patch consistency learning. At the decision level, Adaptive Cosine-Angular Margin (ACAM) adjusts angular margins using semantic difficulty and feature confidence to balance intra-class compactness and inter-class separability. Together, the two modules improve lesion-sensitive representation learning and mitigate old-class bias. Experiments on multiple benchmarks show an average \textbf8.5% gain in \textitAll accuracy over the strongest competing methods. On Kvasir, MedXplore reduces false-old errors from 14.50% to 0.80%, demonstrating strong robustness under severe old-new ambiguity.
[CV-101] MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
链接: https://arxiv.org/abs/2607.27616
作者: Jiajia Lin,Mingxuan Du,Tuowen Zhou,Benfeng Xu,Hongtao Xie
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.
[CV-102] MeshFM: 2D Features Are All You Need for 3D Shape Understanding
链接: https://arxiv.org/abs/2607.27592
作者: Jinfan Zhou,Richard Liu,Itai Lang,Rana Hanocka
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: Project page: this https URL
Abstract:We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. Our method distills 2D features from visual foundation models into 3D. We train a feedforward network to directly predict 3D features without requiring optimization during inference. The approach utilizes a two-stage training strategy. First, we optimize a feature field in 3D using only 2D feature supervision. Second, we train a network to regress this feature field. The entire procedure requires no 3D annotation, instead relying on the powerful information in 2D foundation models. We demonstrate that our learned features can be immediately applied to downstream tasks, including part segmentation, dense correspondence, and mesh deformation. Extensive experiments show that MeshFM, trained solely with 2D supervision, performs on par with methods trained explicitly with 3D supervision, even without task-specific fine-tuning. Moreover, our model is trained to be robust to extreme rotations of the input objects. Project page: this https URL
[CV-103] ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
链接: https://arxiv.org/abs/2607.27585
作者: Dekun Yuan,Zhongwei Li,Zheng Qiao,Jie Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:As primary consumers in the marine food chain, zooplankton play a crucial role in maintaining marine ecological balance. However, the Segment Anything Model (SAM) exhibits limited performance in microscopic image instance segmentation due to its lack of zooplankton-specific domain knowledge. To address these challenges, we propose a novel instance segmentation model based on SAM and wavelet transform (ZMIS-SAM), effectively tackling issues such as inaccurate classification, discontinuous segmentation of slender appendages, and incomplete boundary segmentation. Our framework incorporates three core innovations: ZM-ViT enhances SAM’s capability to model zooplankton morphology and image intensity distributions through two lightweight adapters, the Neighboring Feature Aggregation Module (NFAM) improves continuous segmentation of semi-transparent slender appendages by integrating general-purpose and domain-specific features, and the Wavelet-based Multi-scale Multi-directional Feature Enhancement (WM2FE) module effectively recovers high-frequency details to refine boundary segmentation completeness. Extensive experiments demonstrate that ZMIS-SAM achieves state-of-the-art instance segmentation performance on the zooplankton dataset and exhibits strong generalization capability across multiple public cross-domain datasets.
[CV-104] Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA CVPR2026
链接: https://arxiv.org/abs/2607.27566
作者: Site Li,Jianyi Hao,Xiaofeng Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Presented at the CVPR 2026 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities
Abstract:Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negative mixing, and staged continuation all appear plausible from first principles. We test a simpler competing hypothesis on MedFrameQA: methods that remain tightly aligned with the benchmark’s final answer objective should be the strongest \emphrobust adaptation family once evaluation is controlled across fixed splits, matched budgets, repeated seeds, and calibration. We compare controller-based methods, scaffold evolution, static mixed supervision, continuation-heavy variants, and direct answer-only supervised fine-tuning (SFT). The strongest robust family is direct decoder-only answer SFT on MedGemma-1.5-4B. Empirically, this family yields substantial improvements in held-out report accuracy over frozen baselines while remaining remarkably stable across repeated seeds and matched controls, ensuring our claims reflect true family-level robustness rather than an isolated hyperparameter peak. Furthermore, post-hoc calibration effectively repairs confidence estimation without compromising accuracy, and the core approach transfers consistently to secondary backbones like Qwen2.5-VL-3B. The main result is therefore not that a complex auxiliary mechanism wins, but that objective-aligned direct answer SFT is the strongest robust adaptation family we found for MedFrameQA. By establishing this strong, minimalist baseline, we hope to redirect community focus toward fundamentally robust optimization rather than architectural complexity.
[CV-105] Inference-Time Agent ic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning CVPR2026
链接: https://arxiv.org/abs/2607.27564
作者: Site Li,Jianyi Hao,Xiaofeng Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Presented at the CVPR 2026 Workshop on Multi-Modal Reasoning for Agentic Intelligence
Abstract:Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbforder-vote policy achieves 57.89 \pm 0.65% final-test accuracy, significantly outperforming the fixed baseline ( 52.73 \pm 0.42% ) and the more complex, albeit brittle, order-rerank variant ( 55.79 \pm 0.43% ). Paired bootstrap analysis confirms these significant gains. Extending the evolutionary search budget from 50 to 100 generations yields no generalization benefit: while holdout performance marginally increases, final-test accuracy drops from 57.89% to 56.02% . Our findings suggest that for multi-image medical reasoning, defining the correct agentic decision rule is substantially more impactful than expanding the optimization search budget.
[CV-106] Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
链接: https://arxiv.org/abs/2607.27558
作者: Mingi Kim,Yongjun Kim,Hyungki Kim
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Recovering Parametric CAD sequences from raster-format 2D Computer-Aided Design (CAD) drawings accumulated prior to digital transformation is important for part reproduction and manufacturing process automation. However, existing studies either process only vector drawings or are limited to specific domains, and fail to explicitly connect dimensional annotations to geometric information, limiting their use of dimensional information for 3D Parametric CAD sequences recovery. We propose Drawing-Recode, a framework that generates Parametric CAD sequences as CAD code from raster 2D CAD drawings. Drawing-Recode extracts geometric features via an image encoder and recognizes annotations through a separate text recognition module, then explicitly grounds annotations to geometric information using cross-attention and our proposed Annotation Grounding Loss (AGL). The resulting features are fed into a Large Language Model (LLM) to generate CAD code in the Structured Parametric CAD Code (SPCC) format. Experiments show that Drawing-Recode outperforms existing baselines and remains robust on scanned drawings resembling industrial conditions. We expect Drawing-Recode contributes to digitizing raster 2D CAD drawings in industrial settings and to part reproduction and manufacturing automation.
[CV-107] Cross-Embodiment Transfer via Behavior-Aligned Representations
链接: https://arxiv.org/abs/2607.27549
作者: Ajay Sridhar,Jensen Gao,Jonathan Yang,Jean Mercat,Suneel Belkhale,Dorsa Sadigh
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Project page: this https URL
Abstract:Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: this https URL.
[CV-108] ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction
链接: https://arxiv.org/abs/2607.27537
作者: Dexuan Ding,Yuankai Qi,Luping Zhou,Jian Yang,Quan Z. Sheng,Ming-Hsuan Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, while most subject-specific brain structure remains stable over time. An effective model should therefore preserve global brain structural consistency while remaining sensitive to fine-grained disease progression. Existing latent-space-based methods improve computational efficiency, but suffer from information loss during their compression-reconstruction procedure. In contrast, direct voxel-space methods avoid latent reconstruction but commonly use a unified prediction pathway to model brain structure and progression-related changes. Subtle local changes may therefore be overshadowed by the dominant stable brain structure. To address these challenges, we propose ProgFormer, a hierarchical voxel-space Diffusion Transformer for longitudinal brain MRI prediction. ProgFormer uses a coarse pathway to perform the primary volumetric prediction from 3D patch tokens. This pathway models overall brain structure and longitudinal context. The fine pathway then uses the coarse representations as spatio-temporal grounding for voxel-level refinement within individual patches. The two pathways jointly estimate a velocity field directly in voxel space through conditional flow matching, enabling end-to-end prediction without a separately learned image autoencoder. The predicted future scan is then generated from Gaussian noise by integrating the estimated velocity field over a sequence of Euler steps. Extensive experimental results on three widely used benchmarks, ADNI, AIBL, and OASIS, under both pairwise and trajectory settings demonstrate favourable performance compared against several state-of-the-art methods.
[CV-109] IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
链接: https://arxiv.org/abs/2607.27465
作者: Mengqi He,Jing Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 3 figures
Abstract:Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be computationally expensive. Existing ensemble attacks often rely on multiple surrogate models, increasing the computation cost, even harder for segmentation. This paper studies an efficient single-source alternative for transferable attacks on semantic segmentation. We formulate transferable attack composition as a chained computation over differentiable attack components, allowing the expensive source-model gradient computation to be shared. To reduce the update instability introduced by chained composition, we further use an integrated-gradient-style path-averaged direction as an empirical stabilization heuristic. Experiments on Pascal VOC and Cityscapes evaluate the resulting transferability efficiency trade-off across CNN- and transformer-based segmentation models. IGME achieves competitive transferability compared with single-source baselines and favorable runtime compared with model-ensemble attacks, while requiring access to only one source model.
[CV-110] VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agent ic Dual-Engine System
链接: https://arxiv.org/abs/2607.27380
作者: Haodong Li,Tianfei Ren,Xiaoxiao Ma,Chunmei Qing,Zhen Fang,Sipeng He,Ziyu Guo,Haoyu Wu,Juanxi Tian,Yihang Zou,Ruichuan An,Dongzhi Jiang,Boxue Yang,Ji Xie,Xu Huang,Wenhao Yan,Jialv Zou,Zhengrong Yue,Yaxin Luo,Xiaotong Li,Yuzhu Wang,Junyan Ye,Jinjing Zhao,Zehui Chen,Lin Chen,Renye Yan,Feng Zhao,Pheng-Ann Heng
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 15 pages, 3 figures, and 3 tables
Abstract:Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
[CV-111] PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
链接: https://arxiv.org/abs/2607.27378
作者: Xiaohan Li,Xinyu Liu,Chang Liu,Sum Wing Au Yeung,Jun Liu,Yixuan Yuan,Hui Chen
类目: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
备注:
Abstract:Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-grained, clinically reliable benchmarks that reflect expert interpretation. This work introduces PanDent, a large-scale, clinically grounded OPG benchmark built upon fine-grained, expert-validated tooth-level annotations. The dataset comprises 9,524 high-quality OPGs, each associated with comprehensive structured annotations produced by experienced dentists and further validated by an oral and maxillofacial radiologist, providing clinically reliable supervision for tooth-level diagnosis and reasoning. Clinically consistent radiology reports are constructed from expert-validated findings using clinician-defined reporting logic, establishing explicit correspondence between structured clinical evidence and free-text descriptions. This design enables evaluation of whether MLLMs generate reports that are not only linguistically coherent but also clinically consistent with expert-validated tooth-level findings. Experiments are conducted on diverse MLLMs, including state-of-the-art (SOTA) proprietary models, general-domain open-source models, and medical-specific models. Results show that current MLLMs can generate fluent reports, yet fail to produce clinically consistent descriptions, exhibiting substantial errors in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent significantly improves structure-language consistency, substantially enhancing visual localization accuracy and diagnostic correctness, and bringing model outputs closer to expert dental interpretation. These results establish PanDent as a rigorous benchmark for evaluating tooth-level clinical reasoning in MLLMs and a valuable resource for clinically grounded dental AI.
[CV-112] Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
链接: https://arxiv.org/abs/2607.27357
作者: Dillan Imans,Phuoc-Nguyen Bui,Duc-Tai Le,Hyunseung Choo
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 16 pages, 2 figures, 4 tables
Abstract:Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student modality. In medical image analysis, however, the two modalities are often unpaired: they are collected from different patient cohorts and occupy geometrically incompatible feature spaces. This makes instance-level distillation invalid and direct feature matching unreliable. To address these challenges, we propose Shared Semantic Codebook Distillation (SSCD), which compares teacher and student representations through a shared discrete codebook. Each image is represented as a distribution over a common, modality-agnostic vocabulary, and knowledge is transferred by aligning these distributions across modalities, both globally and class-conditionally, without requiring paired samples or directly comparable raw features. The codebook is evolved online by exponential moving average and kept diverse through entropy regularization and dead-code restart. At inference, all teacher-side and codebook modules are discarded, leaving only the student encoder and classifier. On two heterogeneous unpaired settings, OCT-to-fundus retinal disease classification and CT-to-chest-X-ray pneumonia classification, SSCD improves the student from 64.5 to 70.2 macro-F1 and from 73.8 to 76.3 macro-F1, respectively, outperforming all evaluated distillation baselines on both settings. Code and pretrained models are available at this https URL
[CV-113] Bunraku: Turning a Single Illustration into an Editable Live2D Character
链接: https://arxiv.org/abs/2607.27348
作者: Junhao Chen,Jingjia Mao,Dayong Li,Chenghai Li,Saining Zhang,Zhihao Li,Hao Zhao,Yufei Wang,Ruqi Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Live2D is the dominant 2D character-animation format for anime characters and virtual avatars, representing each character as a stack of RGBA layers driven by per-layer mesh deformation. Despite its wide use in virtual streaming, mobile games, and interactive characters, authoring a Live2D model still demands weeks of manual layer separation, occlusion completion, mesh placement, and keyframing, and no prior generative method produces such a structured asset end-to-end. We present the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move. Stage 1 casts layered decomposition as a layered diffusion process under a Live2D-aware organ-level taxonomy, producing an ordered RGBA stack with hidden-region completion. Stage 2 builds a content-conforming triangle mesh for each layer from its alpha channel alone, then predicts the keypose displacement field of all layers jointly: every vertex of every layer is one token, self-attention spans layer boundaries, and each displacement is factorised into a bounded direction and a log-magnitude. Joint rather than independent prediction is what makes the result a coherent character instead of separately plausible parts, and is our largest gain; scaling the network 112x yields none. On 50 held-out characters, under true generation with no teacher forcing, Stage 2 attains a per-vertex direction cosine of 0.768 (median 0.828). Because a layer’s mesh derives from its alpha channel, a clothing layer can be re-textured from a natural-language instruction while the mesh and predicted animation are reused byte-for-byte. We further contribute Live2D-Bench, the first standardized benchmark for the task, and an 8,884-model Live2D corpus with layer and animation supervision.
[CV-114] Position Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
链接: https://arxiv.org/abs/2607.27304
作者: Supratik Bhowal,Subhrajyoti Basu,Aritra Gir Mahanta,Anik Pal Chowdhury
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral framework that perturbs a single clinically meaningful attribute within a model’s own generated reasoning and measures whether the resulting prediction follows the edited reasoning. Our framework combines a dual-arm protocol comparing re-prompted evidence with prefix-forced continuation, together with a provenance-controlled intervention that varies only the attributed source of identical reasoning to disentangle reasoning mediation from sycophancy. We evaluate LLaVA-Med and MedGemma on 1,000 VQA-RAD samples each. Prefix-forced continuation consistently yields higher mediation faithfulness than re-prompting, while the provenance analysis reveals distinct model-specific deference behaviors. Across both models, removing visual evidence increases reliance on injected reasoning, whereas laterality is the least faithfully tracked clinical attribute. These results show that the mechanism used to inject reasoning substantially affects measured faithfulness and that contextual position, rather than stated provenance, is the primary determinant of whether medical VLMs use their generated reasoning.
[CV-115] VETO: Towards Protecting Images From Frontier AI Editing
链接: https://arxiv.org/abs/2607.27292
作者: Jonas Grebe,Hossein Shakibania,Tobias Braun,Marcus Rohrbach,Anna Rohrbach
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The rise of powerful, accessible image-editing models such as FLUX.2 has brought high-fidelity editing within broad reach. Their capabilities now extend beyond localized modifications to extracting and recontextualizing objects and identities in entirely new scenes. By allowing prompt and generation tokens to attend directly to reference-image tokens, modern models blur the boundary between conventional editing and text-to-image synthesis. This expanded generative freedom also broadens the space of potential misuse, as harmful transformations are no longer confined to a predictable set of localized edits. Existing anti-edit defenses are designed to disrupt the semantic bottleneck of the reference-image encoding in legacy diffusion pipelines. However, newer editors distill reference information through joint-attention blocks, thereby often circumventing these protections. We therefore introduce VETO, a subtle anti-edit cloak that disrupts this inner mechanism through which modern models read the source image. Additionally, as existing editing benchmarks leave comprehensive recontextualizations largely untested, we introduce VetoBench, which evaluates defenses not only on conventional localized edits but also on broader contextual shifts. Across two contemporary editing models and three benchmarks, VETO consistently outperforms existing defenses while providing a stronger protection-fidelity trade-off.
[CV-116] OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
链接: https://arxiv.org/abs/2607.27278
作者: Kaiyu Li,Zepeng Xin,Zixuan Jiang,Jing Fu,Lanxuan Xue,Lingyu Zhang,Xiangyong Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at this https URL.
[CV-117] heatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography
链接: https://arxiv.org/abs/2607.27266
作者: Diego Belzarena(UDELAR, CB),Seginus Mowlavi(CB),Paula Casariego Castiñeira(ROMA TRE),Alejandra Ulla Lorenzo(USC),Gregory Randall(UDELAR),Jean-Michel Morel(LU - Hong Kong)
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
备注:
Abstract:We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates philological analysis. Using character prototypes derived from clustering and aligning automatically extracted character images, the method defines a typeface distance between any two books. To produce actionable outputs, we develop an a contrario statistical framework to interpret the significance of the computed typeface distances. We apply the method to the philological study of 17 th -century Spanish printed theatre chapbooks in a quantity that exceeds the capabilities of systematic visual inspection by human experts. Our method enables the automatic comparison of Roman and Italic types extracted from different books. After validation by human experts, our method has led to new printer attributions being discovered, and former printer attributions being revised. This success strongly suggests that our method has the potential to enable digital bibliography on a larger scale than was previously possible.
[CV-118] RadHarmony: Radiological Data Handling in the Era of Agent ic AI
链接: https://arxiv.org/abs/2607.27235
作者: Frank Li,Bardia Khosravi,Mohammadreza Chavoshi,Theo Dapamede,YoungSeok Jeon,Janice Newsome,Hari Trivedi,Judy Gichoya
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Training deep learning models on radiological images requires integrating heterogeneous datasets across different sources, file formats, directory layouts, label schemas, and annotation types. We present RadHarmony, an open-source Python library that provides a unified API for loading, harmonizing, and augmenting radiological datasets, with a primary focus on chest radiographs and early support for computed tomography (CT) and magnetic resonance imaging (MRI). RadHarmony standardizes metadata from 24 public datasets into a single tabular format, wraps MONAI’s map-style datasets for deep-learning-ready sample delivery with optional on-disk caching, and supports classification labels, segmentation masks, bounding boxes, and radiology report text through a single interface, with an interactive visualization tool for dataset exploration and verification. To lower the barrier for integrating new datasets, RadHarmony introduces an AI-agent skill that guides the full integration workflow from raw data inspection through code generation and testing. We demonstrate the library’s utility by pretraining RadHarmony-ViT, a reference vision transformer baseline that combines three heterogeneous chest radiograph datasets with no dataset-specific code. The code and pretrained model weights are available at this https URL.
[CV-119] ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
链接: https://arxiv.org/abs/2607.28144
作者: Zheyuan Zhang,Johnson Wu
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 12 pages, 5 figures, 5 tables
Abstract:We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The encoder reduces a source clip to a compact bitstream – a neurally compressed first frame, per-frame pose keypoints, and metadata – totaling about 26 kB for a 77-frame sequence. The decoder is a four-step distilled diffusion transformer that reconstructs the video conditioned on the transmitted pose and reference frame. Compared with x264/x265, ReGenVC reduces the bitrate to roughly one tenth of that required by traditional codecs (about 26 kB vs. 250–280 kB for essentially artifact-free reconstruction); at a matched ultra-low bitrate, conventional codecs collapse into blocking artifacts while ReGenVC stays sharp by exploiting a strong generative prior. The central obstacle to deploying such a codec is decoder latency: multi-step sampling with transformer and VAE components is too slow for interactive use. We make the decoder real-time through four-step distillation and three model-preserving system techniques: (i) eight-GPU unified sequence parallelism (Ulysses Ring), (ii) a spatially-split VAE, and (iii) a three-stage overlapped pipeline; an analytical timing model characterizes the real-time feasibility region. On an 8-GPU node, the system sustains 24 fps output (972 ms per 25-frame window, within the 1000 ms budget), enabling a live browser stream without observed frame underruns. A hybrid CPU-GPU deployment further runs the encoder on the CPU at 24 fps and offloads the decoder-side one-shot conditioning encoders to the CPU, reducing the per-GPU memory peak from 21.1 GB to about 7.7 GB. To our knowledge, ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system.
[CV-120] SOG: A Format For Temporally And Spatially Ordered Gaussians ICIP
链接: https://arxiv.org/abs/2607.28049
作者: Shady Gmira,Evangelos Alexiou,Emmanouil Potetsianakis,Emmanuel Thomas
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注: 2026 IEEE International Conference on Image Processing (ICIP). IEEE, 2026
Abstract:We propose Temporally and Spatially Ordered Gaussians (TSOG), a format for efficient representation of 4D Gaussian Splatting (4DGS) content. TSOG extends the Spatially Ordered Gaussians (SOG) framework to the temporal domain by introducing a timeline attribute and temporal parameterization of geometry and appearance attributes. Similar to SOG, TSOG is a lossy format that assigns each Gaussian a unique index and encodes attribute values as index-aligned image data. TSOG is model-agnostic, extensible, and compatible with both discrete and continuous 4DGS representations. Evaluation using a PLYs sequence and FreeTimeGS as baselines, serving as simplistic and state-of-the-art 4DGS representations respectively, shows file size reductions exceeding 90%, with PSNR differences ranging between -0.42 and +0.85 dB. These results demonstrate substantial file size savings with minimal quality degradation, enabling efficient representation, storage, and delivery of dynamic scenes for next-generation 4D content.
[CV-121] Endo-NeRF: Uncertainty-Aware Neural Rendering with Multi-Resolution Hash Encoding for Dynamic Surgical Scene Reconstruction
链接: https://arxiv.org/abs/2607.27825
作者: Gousia Habib,Laura Ruotsalainen
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tissue deformation, occlusions, specular reflections, and restricted viewpoints. In this study, we introduce Endo-NeRF++, a neural rendering framework that accounts for uncertainty in the reconstruction of dynamic surgical scenes. Expanding on EndoNeRF, the suggested approach incorporates multi-resolution hash-grid encoding, temporal feature merging, and uncertainty-informed adaptive sampling to enhance reconstruction accuracy and temporal coherence in deformable endoscopic this http URL multi-resolution hash-grid representation within the framework effectively captures both coarse and fine anatomical details, while temporal feature blending ensures stable reconstruction during tissue deformation and surgical tool occlusions. Additionally, uncertainty-driven adaptive sampling assigns more samples to uncertain areas to enhance rendering quality and geometric coherence. Experiments on robotic surgical video sequences demonstrate that the proposed uncertainty-guided adaptive sampling improves PSNR by up to 1.22,dB (4.3%), increases SSIM by up to 5.3%, and reduces LPIPS by up to 55.1% compared with the EndoNeRF baseline.
[CV-122] hree-Photon Bayesian Imaging of Ortho-Positronium
链接: https://arxiv.org/abs/2607.27741
作者: L. Raczynski,W. Krzemien,A. Coussat,M. Bala,B.C. Hiesmayr,K. Klimaszewski,M. Obara,R. Y. Shopa
类目: Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph)
备注: 17 pages, 4 figures
Abstract:PET provides functional images relying on two-photon coincidences from positron-electron annihilation. In human tissue, about 40% of annihilations are preceded by Ps formation, of which o-Ps component partially decays into three photons, with the remainder annihilating via pick-off or spin-exchange into two photons. This three-photon channel carries additional information about the surrounding micro-environment, including the three-to-two-photon yield ratio as a potential diagnostic marker. We propose the TRIO algorithm, a novel three-photon event-by-event image reconstruction algorithm formulated as a Bayesian maximum a posteriori inference problem. TRIO unifies time-based trilateration, energy-based reconstruction and, for the first time, a physics-informed prior derived from the QED description of Ps decay within a single probabilistic framework. In contrast to positronium lifetime imaging, which requires a prompt photon and is therefore restricted to specific radionuclides, TRIO relies solely on the three photons and is fully compatible with standard radionuclides such as 18F. Monte Carlo simulation modelled after the Siemens Biograph Quadra scanner demonstrates a mean position error of 1.62~cm, improving by approximately a factor of two over the time-based trilateration (3.05 cm) and by about an order of magnitude over energy-based reconstruction alone (18 cm). More importantly, the proposed Bayesian approach is compatible with existing TOF-PET scanners that can register three-photon annihilation coincidences. Comments: 17 pages, 4 figures Subjects: Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph) Cite as: arXiv:2607.27741 [physics.med-ph] (or arXiv:2607.27741v1 [physics.med-ph] for this version) https://doi.org/10.48550/arXiv.2607.27741 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[CV-123] oward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data
链接: https://arxiv.org/abs/2607.27286
作者: Yogisri Pujitha Chinthoti
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 2 figures
Abstract:Automated classification of pulmonary disease from chest radiographs is a widely studied application of machine learning in medical imaging. This paper presents a pilot study evaluating classical texture- and gradient-based feature representations for distinguishing COVID-19 from other forms of pneumonia using the publicly available COVID-19 Image Data Collection (668 posteroanterior/anteroposterior radiographs from 408 patients). Using histogram of oriented gradients (HOG) and gray-level co-occurrence matrix (GLCM) texture descriptors with classical classifiers (logistic regression, random forest, and support vector machine), evaluated under patient-level 5-fold stratified cross-validation to prevent data leakage, we obtain a best mean accuracy of 75.4% and AUC of 0.755, modestly exceeding the 71.6% majority-class baseline. We report these results transparently, including their limitations, and use them to motivate and scope a proposed multi-modal deep learning architecture – combining convolutional and transformer-based encoders across imaging modalities – as a direction for future work requiring access to larger, multi-institutional, ethically sourced datasets.
人工智能
[AI-0] PAC-MAN: Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball
链接: https://arxiv.org/abs/2607.28623
作者: Lizhi Yang,Junheng Li,Aaron D. Ames
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Website at this https URL
Abstract:We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid dodgeball. The deployed policy sees the ball only as segmentation-masked depth from a head-mounted camera, while training-time CBF guidance represents clearance to every body link, and an adversarial motion prior regularizes the resulting evasive reflexes. We evaluate on a controlled any-link contact benchmark with seeded throws in two regimes: single throws and a deployment loop in which the robot walks back to its station and recovers between throws. On this benchmark, the policy comes within a few points of a privileged state oracle: a fixed onboard camera alone is adequate for evasion. We find that usable barrier structure depends on perceptual observability: Joint-CBF gives the best performance with accurate ball states, degrades under fixed-camera observations when used only as training guidance, and recovers with a ball-tracking gimbal or privileged runtime filter. We therefore deploy a lightweight Link-CBF policy zero-shot on the Unitree G1 in the real world, where it tolerates imperfect perception, succeeds on 95% of throws, and uses semantic segmentation to dodge different balls.
[AI-1] PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
链接: https://arxiv.org/abs/2607.28587
作者: Manyi Wang,Junjielong Xu,Pinjia He
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)
Abstract:SWE-bench-like benchmarks are widely used for evaluating LLM’s issue resolution capability. They typically follow a common construction pipeline: each PR (Pull Request) is paired with its linked issue by extracting issue references from the PR description; the issue description is used as the problem statement, and the PR patch serves as the test oracle. However, due to the inherent complexity of developing and maintaining large repositories, such PR-Issue pairings are often misaligned in practice. In this work, we systematically study SWE-bench Verified instances, finding that 13.6% exhibit misalignment across five patterns in eleven fine-grained scenarios. To enable reliable and scalable construction of those benchmarks in the future, we propose PAIChecker, a multi-agent system for checking PR-Issue misalignment in SWE-bench-like benchmarks. Specifically, PAIChecker adopts a three-phase design that combines specific pattern identification, cross-agent label synthesis, and code-level validation, thereby enabling more accurate, generalizable, and progressively verified detection. Experiments on SWE-Gym and SWE-bench Multilingual show that PAIchecker achieves the best performance across all four LLM backbones, reaching up to 92.12% and 91.67% binary accuracy, respectively.
[AI-2] DualG-MRAG : Decoupling Macro-Reasoning and Micro-Matching for Multimodal Retrieval-Augmented Generation ACM-MM2026
链接: https://arxiv.org/abs/2607.28580
作者: Jiacheng Tao,Qingyun Sun,Haonan Yuan,Ziwei Zhang,Jianxin Li
类目: Artificial Intelligence (cs.AI)
备注: Accepted to the 34th ACM International Conference on Multimedia (ACM MM 2026). 12 pages
Abstract:While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex multi-hop reasoning tasks. Existing methods primarily focus on independent instance-level matching, which often fails to capture explicit relationships across modalities and documents. Although Graph-enhanced methods introduce structural modeling, they face a fundamental challenge in multimodal scenarios: incorporating fine-grained visual features leads to rapid graph expansion and retrieval noise, whereas coarse-grained representations cause the discarding of critical local evidence. To address this dilemma, we propose DualG-MRAG, a Dual-tier framework that introduces a decoupled architecture comprising Macro-reasoning and Micro-matching Graphs for Multimodal RAG. Specifically, to suppress retrieval noise by isolating global structural reasoning from fine-grained evidence matching, we construct a Macro Graph for global topological routing and a Micro Graph for precise local verification. Subsequently, to enable dynamic relevance propagation across heterogeneous evidence sources, we formulate retrieval as a query-driven message passing process via a GNN Retriever. Furthermore, to provide the generative model with coherent structural guidance, we introduce a dynamic programming decoding mechanism that extracts explicit reasoning paths directly from the GNN’s forward pass, replacing the standard input of isolated document chunks. Extensive experiments demonstrate that DualG-MRAG outperforms baselines in both evidence recall and complex QA accuracy.
[AI-3] Rethinking Inference-Time Scaling in Local Computer-Use Agents : Failure Modes and Compute Tradeoffs
链接: https://arxiv.org/abs/2607.28573
作者: Woongkyu Lee,Jungwook Choi
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constraints remains challenging. While recent studies show that inference-time scaling can improve frontier computer-use agents through additional computation during execution, its effectiveness for resource-constrained local models remains poorly understood. We present a systematic empirical study of inference-time scaling in local CUAs across contextual, temporal, structural, and parallel dimensions. We evaluate Qwen3-VL-8B/30B-A3B, UI-TARS-1.5-7B, and OpenCUA-7B on the OSWorld benchmark. Our results show that additional computation often yields diminishing returns while changing failure modes. Contextual scaling provides historical grounding that improves trajectory stability and task accuracy, but its gains saturate as token cost increases and failures shift from repetitive or stalled trajectories toward premature false successes. Temporal scaling similarly reduces max-step stalls, yet does not substantially improve task success, indicating that longer horizons often extend erroneous trajectories rather than correct them. We further find that structural decomposition can introduce planning and formatting overhead in local two-stage agents, while parallel scaling partially mitigates these failures at a substantial computational cost. Overall, our findings suggest that efficient local CUAs require selective compute allocation, failure-aware control mechanisms, and agentic frameworks designed around the capabilities and limitations of local models.
[AI-4] MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
链接: https://arxiv.org/abs/2607.28527
作者: Mao-xun Huang,Jerry Wang,Yi-Cheng Lai,Zhengxin Zhang,Claire Cardie,Hen-Hsen Huang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model-based multi-agent systems improve complex problem solving through task decomposition, agent specialization, information exchange, and intermediate validation. However, existing systems typically treat communication topology as a fixed design choice or an offline optimization target. We introduce MANTA, a framework for Multi-Agent Network Topology Adaptation that enables communication structures to self-evolve at inference time. Before execution, MANTA initializes a task-conditioned topology from prior structural experience. During deployment, it monitors collaboration traces and applies bounded structural updates when the current organization becomes insufficient. These updates can modify agent roles, communication links, execution order, information visibility, and validation pathways while preserving the task interface and agent budget. We evaluate MANTA against representative single-agent and multi-agent baselines on five benchmarks spanning information seeking, tool use, planning, workflow execution, and mathematical reasoning. MANTA achieves the highest average score of 74.0, outperforming the strongest baseline by 5.8 percentage points and obtaining the best result on PlanCraft. These results show that inference-time self-improvement can extend to the architecture of collaboration itself.
[AI-5] Selective Credibility-Limited Belief Update
链接: https://arxiv.org/abs/2607.28523
作者: Theofanis Aravanis,Costas D. Koutras
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注:
Abstract:Belief update concerns changes in an agent’s beliefs induced by changes in the underlying world. Standard Katsuno-Mendelzon update assumes that an epistemic input can be incorporated from every initially possible world, whereas credibility-limited belief update restricts, for each source world, the successor worlds regarded as credible or reachable. Nevertheless, existing credibility-limited approaches treat the epistemic input as an indivisible whole, and therefore cannot represent cases in which only part of a compound epistemic input can be realized. We introduce selective credibility-limited belief update, in which the epistemic input is transformed, relative to each source world, into a weaker proxy before the credibility-limited transition is performed. We provide semantic and axiomatic characterizations of the resulting class of update operators. We then identify two well-behaved sub-classes; namely, consistency-preserving update operators, which require every transformed epistemic input to be credible from its source world whenever the original epistemic input is consistent, and maximal consistency-preserving update operators, which additionally require the selected proxy to be maximally informative among the credible consequences of the original epistemic input. Finally, we establish the generality of the proposed framework by showing that credibility-limited belief update is recovered as a special case, while Katsuno–Mendelzon belief update emerges when credibility restrictions are removed and the transformation functions are taken to be identities. These results demonstrate that the framework provides a unified and strictly more expressive account of belief update, encompassing established approaches while supporting source-dependent selective acceptance.
[AI-6] InfoOps Bench: A live information operations safety benchmark
链接: https://arxiv.org/abs/2607.28503
作者: Dorian Quelle,Lisa-Maria Neudert,Jonathan Bright,John Gallacher
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live monitoring pipeline which tracks Russian, Chinese and Iranian state-backed information assets. Alongside this paper, we release a companion website that tracks the most prominent claims spread by state-backed media outlets, updated weekly, available from: this http URL. The dynamic nature of the benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations. Integrity scores, defined as the percentage of refused requests, range from 8.8% to 94.5%, an 85.7-percentage-point spread not explained by model size. Model choice also changes the character of the resulting operation. Some models fabricate details and produce output more harmful than the source material, others defuse claims even while complying, and fact-checking rates vary from 2.9% to 72.9%. Integrity against information operations is at least partly related to refusal to produce content even for benign claims, illustrating the challenge of balancing model usability with safety. With one exception (this http URL’s GLM 5.2), the Chinese-developed models sharply cut compliance on factually grounded but China-critical claims, dropping 48-70 percentage points relative to matched benign claims. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.28503 [cs.AI] (or arXiv:2607.28503v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28503 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-7] SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination
链接: https://arxiv.org/abs/2607.28488
作者: Yunhao Liang,Xianqi Cao,Pujun Zhang,Yuan Qu,Yongzhi Qi,Ningxuan Kang,Max Z.J. Shen
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Can supply-chain AI move beyond isolated decision modules toward unified operational planning? A complete replenishment plan specifies which products each location carries, which upstream facility supplies it, how often it is replenished, and how deliveries are routed. These decisions are operationally coupled: the selected assortment changes the demand and load passed to later stages; source assignment and replenishment frequency reshape the delivery requests; and route feasibility and cost, in turn, determine the system value of the earlier choices. Yet in modern supply chains, these decisions are often handled by separate departments and optimized through separate systems, which can lead to stockouts, inventory exposure, and avoidable transportation. We propose SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination, a composite policy model that represents supply-chain entities as tokens, contextualizes them through a shared operational representation, and maps each token type to the corresponding decision interface. Each decision builds on the partial plan formed by earlier decisions while the completed plan is evaluated using a shared system-level utility. We instantiate this framework in urban fresh-retail replenishment, where service frequency, assortment, capacity pressure, and road-network routing interact strongly, and evaluate it on real operational data from Dingdong and this http URL, two large-scale supply chains operating at different replenishment echelons. Across both settings, SCOPE consistently outperforms methods that optimize each decision stage separately, as well as practice-oriented baselines commonly used in supply-chain operations. These results show that learning and coordinating cross-department operational couplings lead to more effective end-to-end supply-chain decisions.
[AI-8] A Fuzzy Rule-based Neuro-Symbolic Approach for Pipe Severity Prediction in Sewer Networks
链接: https://arxiv.org/abs/2607.28481
作者: Ngoc Thai Le,Thanh Ma,Umberto Straccia
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Standard automated sewer pipe severity assessment relies on direct image classification, creating a “black box” where the link between visual defects and final severity scores remains implicit. This study introduces a modular, fuzzy rule-based neuro-symbolic framework that bridges this gap by decoupling neural perception from symbolic reasoning. The perception module utilizes a Swin Transformer to predict 14 multilabel inspection CODE degrees directly from images. For reasoning, a DT, specifically Weka’s J48, algorithm is trained on ground-truth CODEs and severity labels, and its paths are converted into 19 fixed IF–THEN rules. Inference operates via fuzzy logic: t-norm activations from CODE conditions are weighted by rule confidence and combined with corresponding s-norms to produce interpretable class evidence. We assessed Product, Łukasiewicz, and Hamacher operator pairs using a dataset of 3,244 images spanning five highly imbalanced severity classes. Ground-truth labels were robustly generated via consensus from five independent large language models analyzing original inspector notes. Our results show an improvement of accuracy, balanced accuracy, Macro F1 and MCC by 17.9%, 12.2%, 23.0%, and 17.3%, respectively, over image-only based classification. Overall, the framework combines competitive class-balanced performance with traceable reasoning from predicted CODE degrees to rule supports and severity evidence. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.28481 [cs.AI] (or arXiv:2607.28481v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28481 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-9] A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
链接: https://arxiv.org/abs/2607.28466
作者: Jia Yu,Yan Zhu,Yili He,Zilong Wang,Xinyang Jiang,Peiyao Fu,Ruijie Yang,Tianyi Chen,Siyuan Li,Zhihua Wang,Fei Wu,Quanlin Li,Xian Yang,Pinghong Zhou,Shuo Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.
[AI-10] LeanCSP: A Framework for Certifying Constraint Reformulation and Solving in Lean
链接: https://arxiv.org/abs/2607.28459
作者: Pablo Manrique,Stefan Szeider
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注:
Abstract:Constraint programming is a core technology for solving complex combinatorial problems in scheduling, planning, configuration, and verification. Trusting its results therefore demands guarantees at two levels: that reformulations applied beforehand are semantics-preserving, and that solvers produce correct answers. In this work, we introduce a framework that addresses both verification levels in the Lean theorem prover: it can be used to prove formulation-level properties, such as equivalence, equisatisfiability, and the correctness of symmetry-breaking constraints, parametrically for entire problem families; and to check solver-produced certificates for individual instances via translation backends to external formats such as MiniZinc, SMT-LIB, and OPB. Combining both levels yields an end-to-end workflow that establishes the satisfiability or unsatisfiability of a constraint problem without trusting the external solver. Experimental results show that our framework’s verified symmetry breaking also pays off in practice: a single parametric proof per problem family, reused across all instance sizes, reduces solver search effort by a factor of up to 2x10^7, while the entire in-Lean certification stays affordable, taking at most a few minutes for our largest instances.
[AI-11] Machines that know they are aging: a framework for hardware-aware autonomous intelligence
链接: https://arxiv.org/abs/2607.28451
作者: Cheng Siong Chin,Jianhua Zhang,Mohan Venkateshkumar
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 1 figure, 8 pages
Abstract:Autonomous systems inevitably age, yet their artificial intelligence typically assumes hardware remains in its original condition. Batteries degrade, sensors drift, processors accumulate timing errors, and memory reliability declines, creating a growing mismatch between assumed and actual capability. This can lead to agnostic collapse, where mission failure arises from accumulated hardware degradation rather than a single component fault. We propose Aging-Aware Autonomous Intelligence (AAAI), a framework that integrates hardware health directly into reasoning, planning, and mission execution. AAAI is built on three pillars: hardware self-awareness, which continuously estimates the health of power, sensing, memory, and computation subsystems using physics-of-failure models; self-adaptive reasoning, which adjusts inference complexity, planning horizon, and task priorities according to remaining hardware capability; and survival-centric intelligence, which allocates remaining operational life across mission objectives through performance optimization, resource conservation, and graceful degradation. Rather than introducing new hardware, AAAI unifies prognostics, lifecycle management, and hardware-aware computing into a closed-loop cognitive architecture. We argue that such integration is essential for autonomous systems operating in inaccessible or safety-critical environments, including space missions, marine robotics, and implantable medical devices. By enabling machines to recognize and respond to their own aging, AAAI improves resilience, extends operational lifetime, and supports safer, more graceful mission completion.
[AI-12] A foundation model of numerical intelligence with cross-disciplinary generalization
链接: https://arxiv.org/abs/2607.28432
作者: Chenghan Wu,Zongmin Yu,Liu Yang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. Large language models exhibit this capacity by inferring task-relevant knowledge from textual context and applying it to new tasks. Yet intelligence need not be confined to language. For scientific and social systems, we need models that acquire and apply knowledge from numerical context-an ability we call numerical intelligence. Here we introduce UNified In-Context Operator Networks (UNICON), a foundation model that exhibits numerical intelligence across disciplines. Using graph-based examples from a system as context, UNICON infers the predictive relation shared across them and applies it to queries from the same system. Across scientific and social systems, including those from disciplines absent from training, the same model approaches specialist performance without retraining. Combining UNICON with language-model agents yields further gains, enabling it to surpass state-of-the-art specialists in a discipline unseen in training. We further show that training-corpus diversity improves generalization to unseen disciplines. Together, these results establish UNICON as a foundation model of numerical intelligence and position it as a building block for a broader ecosystem of artificial intelligence.
[AI-13] When Derived Measurements Mislead: Quantifying and Mitigating LLM Over-Trust with Privileged-Modality Reliability Evidence
链接: https://arxiv.org/abs/2607.28421
作者: Zongheng Guo,Tao Chen,Tianli Li,Mingzhe Cui,Yang Jiao,Lei Xie,Yi Pan,Xiao Hu,Manuela Ferrario
类目: Artificial Intelligence (cs.AI)
备注: 25 pages, including references and supplementary material; 3 figures and 19 tables. Code: this https URL
Abstract:Derived measurements increasingly enter large language model (LLM) pipelines as direct facts despite their instance-dependent validity. We define derived-feature over-trust (DFOT) as the failure in which a downstream LLM assigns such a measurement the epistemic status of a direct fact or uses it outside its valid scope. Using physiological sensing as a case study, D1 tests acceptance of a PPG-derived rhythm contradicted by offline ECG, whereas D2 tests rejection of an offline-confirmed reliable PPG rhythm under misleading severe history. ECG supplies training supervision and offline reference construction but is never shown to the LLM. Five estimands quantify this chain: conflict over-trust rate (COTR) and context-induced error rate (CIR) characterize D1/D2; correct repair rate (CRR) measures frozen-error repair; evidence-specific repair margin (ESRM) contrasts matched and patient-disjoint shuffled evidence; and utility harm rate (UHR) measures unnecessary verification among HIGH-reliability cases used without verification at baseline. The framework does not depend on a particular reliability generator. We demonstrate it on 50,000 paired PPG-ECG records using ECG-to-PPG privileged distillation as an illustrative baseline and PPG-only inference. On a protocol-locked 187-patient test, the baseline improves four repair and specificity endpoints by 1.82-6.69 percentage points, with all paired confidence intervals excluding zero; UHR increases by 0.67 percentage points (95% CI: -0.4 to +1.7). DFOT provides a common evaluation target for stronger mitigation methods. The code is available at this https URL.
[AI-14] On-Policy and Off-Policy Learning for Large Action Spaces
链接: https://arxiv.org/abs/2607.28408
作者: Imad Aouali
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Statistics Theory (math.ST); Machine Learning (stat.ML)
备注: PhD Thesis, 241 pages
Abstract:This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces, both settings face major challenges: inefficient exploration, sparse data coverage, high-variance importance weights, extrapolation bias, and difficult optimization landscapes. The first part develops structured Bayesian methods for on-policy learning. We introduce meTS, a mixed-effect extension of Thompson sampling, and dTS, which leverages diffusion-inspired priors to model dependencies between actions. These methods share information across actions and yield regret guarantees depending on an effective number of actions. The second part addresses off-policy learning. We propose sDM, a structured direct method based on latent variables, show that optimization error can dominate estimation error in large action spaces, and introduce concave, efficiently optimizable policy-weighted log-likelihood objectives. Finally, we develop differentiable pessimistic methods based on exponential smoothing and PAC-Bayesian bounds to control the bias-variance trade-off of regularized importance-sampling estimators.
[AI-15] QuantWAMs: Calibrating at the Right Granularity for World Action Models
链接: https://arxiv.org/abs/2607.28405
作者: Jiacheng Zhou,Jinfan Lv,Ruixuan Li,Longtai Zhang,Yan Wang,Wenqiang Zhang,Lizhe Qi
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 13 pages, 6 figures
Abstract:World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient deployment costly. Existing post-training quantization (PTQ) methods are poorly suited to WAMs because they rely on open-loop objectives, homogeneous model assumptions, and calibration distributions that do not reflect deployment. We present QuantWAMs, a PTQ framework that aligns quantization decisions with the calibration context defined by model structure, rollout distribution, and task objective. QuantWAMs introduces three strategies: shared-basis outlier calibration, which pools activation evidence only across coordinate-compatible modules; co-training-objective saliency, which computes empirical-Fisher scores from the joint video–action gradient and assigns weight precision at a calibration-stable layer granularity; and fixed-intervention rollout auditing, which revises denoising-step protection schedules using reachable closed-loop states without changing the precision budget. We evaluate QuantWAMs on Fast-WAM and LingBot-VA across RoboTwin 2.0, LIBERO, and real-robot manipulation with an AgiBot G2. Under a W4A4-dominant setting, the reported simulation means differ from FP16 by 0.2–0.7 percentage points. Real-robot trials further establish deployment feasibility on three manipulation tasks. For the targeted video and action blocks, QuantWAMs reduces peak weight-and-activation memory to about 29% of FP16 and provides 1.4–1.6 \times block-level speedups.
[AI-16] When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences AAAI2027
链接: https://arxiv.org/abs/2607.28384
作者: Tairan Wang,Liang Zhou,Zikang Zhan,Pingchuan Yan
类目: Artificial Intelligence (cs.AI)
备注: Submitted to AAAI 2027
Abstract:Large language models (LLMs) are increasingly required to integrate multiple sources of information that may be inconsistent or conflicting. However, there is still a lack of controllable and attributable methods for analyzing how models resolve conflicts between competing specifications. We propose a controlled experimental framework for studying model preferences under conflicting specifications. By constructing specifications with explicit conflicts, the framework enables model choices between competing specifications to be directly observed and analyzed. A symmetry-based design further reduces confounding factors, allowing preferences across representation types to be compared systematically. We evaluate the framework on an executable mathematical benchmark with 550 conflict instances spanning 11 function families, comparing four representation types: pure natural language, formal language, naturalized formal language, and input–output examples. Results show systematic preference patterns rather than random behavior, with a consistent ordering: \textFormal \approx \textNaturalized Formal \textPure Natural Language \textInput–Output Examples . Example effects further depend on model capability and function family. We extend the framework to heterogeneous specification conflicts in Boolean algebra, code generation, and the clinical domain, demonstrating its applicability across diverse tasks and specification forms. The framework provides a unified approach for measuring how LLMs resolve conflicts between competing sources of information. Comments: Submitted to AAAI 2027 Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.28384 [cs.AI] (or arXiv:2607.28384v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28384 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-17] HyperClaim: Fine-Grained Cross-Modal Hypergraph Reasoning for Video Misinformation Detection
链接: https://arxiv.org/abs/2607.28375
作者: Xiangbo Wang,Jiasheng Zhang,Xingtong Yu,Luoqiang Lei,Delvin Ce Zhang
类目: Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: 13 pages, including supplementary material
Abstract:Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.
[AI-18] How Benchmarks Mis-Score Computer-Use Agents
链接: https://arxiv.org/abs/2607.28367
作者: Zihan Dong,Zhiyuan Ma,Zekun Wang,Yunqing Li,Zirou Liu,Ruixuan Deng,Qishi Zhan,Rui Qian
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Computer-use agents (CUA) are being deployed to browse the web and operate desktop software, yet their benchmark scores are still commonly produced by brittle scripted oracles. A score is the output of a pipeline in which tasks can be stale, trajectories can omit decisive visual evidence, evaluators can reject valid alternatives, and aggregate reports can hide the cause of failure. We organize these problems into a reliability framework spanning task construction, trajectory observation, scoring, and reporting. We then audit 150 public failure-scored trajectories from five web, enterprise-workflow, and desktop-control benchmarks, find that 15.3% of FAIL verdicts are wrong: 10.7% are evaluator false negatives and 4.7% are broken tasks. For genuine failures, a three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain. We connect these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.
[AI-19] ffic-Audio: Tell Fact from Fiction
链接: https://arxiv.org/abs/2607.28351
作者: Wan Lin,Li Wang,Jindong Wang,Kunyu Feng,Zhizheng Wu
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注: 16 pages, 1 figure, 7 tables. Technical report. Project page: this https URL
Abstract:Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoofing artifacts can be further shaped by variability in source speech, recording environments, and transmission channels. This variability makes robust generalization across heterogeneous conditions a central requirement for practical detection systems. This report presents Teffic-Audio, a general speech deepfake detection system designed for comprehensive evaluation environment. Teffic-Audio adopts a straightforward detector architecture consisting of a Conformer-based speech encoder, multi-head attentive statistics pooling, and a binary classifier. Rather than relying on additional architectural complexity, the system improves generalization through its training recipe, which integrates multi-source data, attack- and source-balanced sampling, and diverse audio augmentation. Trained only with open-source data, Teffic-Audio achieves a pooled EER of 1.454% on the 14 test sets of Speech-DF-Arena, outperforming all currently public systems on the leaderboard. It also obtains the lowest EER on five individual test sets and shows a favorable performance-complexity trade-off compared with larger leading systems. Overall, Teffic-Audio provides a strong and practical reference system for general speech deepfake detection.
[AI-20] Correcting What You Cannot See: Credit Assignment for Perception Distillation in Multimodal Reason ers
链接: https://arxiv.org/abs/2607.28336
作者: Feng Xiong,Leyan Xue,Hongyu Lin
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:On-policy distillation provides dense supervision for multimodal reasoners, but its trajectory-level reward cannot determine whether a failed answer arose from perception or subsequent reasoning. Perception Success Rate (PSR), estimated from multiple reasonings sharing one perception, remains ambiguous because low success conflates perceptual insufficiency with reasoning difficulty. We introduce \textbfPerception-Correction Distillation (PCD), a label-free method that identifies correctable perception failures using downstream failure and teacher–student disagreement as complementary witnesses. Their product, , forms a soft AND gate that strengthens distillation only when both witnesses are present. We motivate this rule through Bayesian evidence combination and show that multiplication is the unique normalized bilinear gate that vanishes when either witness is absent. PCD uses separated perception–reasoning rollouts and mean-preserving weights, leaving the reasoning objective unchanged. Across eight benchmarks, PCD improves the 8B 2B macro average from 44.50 with OPD to 47.28 and the 32B 8B result from 56.94 to 61.22. In matched 2B ablations, removing PCD and separated rollout reduces held-out average by 2.22 and 0.88 points, respectively. Effective multimodal distillation therefore depends not only on what the teacher predicts, but also on identifying when perception is the appropriate target of correction.
[AI-21] Paying for Honesty Without Knowing the Truth: Reputation-Penalty Design for LLM Marketplace Agents
链接: https://arxiv.org/abs/2607.28330
作者: Mingdai Yang,Shicheng Fan,Kejing Yu,Duohao Wang,Li Sun,Hao Peng,Philip S. Yu,Zhiwei Liu
类目: Artificial Intelligence (cs.AI)
备注: 11 pages
Abstract:LLM agents increasingly act as autonomous merchants that write their own product listings, and under competitive pressure, they fabricate attributes to win sales. Even under instructions to be honest, they fabricate attributes in a majority of listings across models. A platform’s obvious remedy—verifying each claim against the truth—is unavailable, because it observes only a noisy, biased complaint signal, never the ground truth. We design CARP, a reputation-penalty mechanism with a deadband that forgives complaint noise and a state-dependent severity that counters reputation-driven detection erosion. CARP requires no product-level ground truth and is robust to strategic gaming. CARP protects consumers by suppressing the sales volume of low-rated liars while sparing honest sellers. Paired with SPARC, it closes most of the consumer-welfare gap relative to a perfect-information oracle, without ever accessing the truth. It also achieves the best welfare of the policies we compare. We further show that this felt penalty becomes behaviorally binding through SPARC, a byte-clean code-gated reflection mechanism: LLM merchants fabricate when lying is free but restrain themselves when fabrication costs them sales, a self-interested response rather than compliance. We trace this distinction to penalty-gated self-correction reasoning, and observe the binding across models, with supporting confidence intervals.
[AI-22] PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
链接: https://arxiv.org/abs/2607.28318
作者: Zongyi Chen,Yu Liang,Jie Lin,Liansheng Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
[AI-23] One Human N Agents : Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated Correlated Confidence
链接: https://arxiv.org/abs/2607.28317
作者: Cesare Zavattari,Alessandro Tommasi,Giuseppe Prencipe
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:A single human must audit N LLM agents under a budget of B \ll N audits per round, guided by self-reported confidence that may be adversarially miscalibrated and by correlated errors. We model this as budgeted noisy inspection over a two-level Gaussian copula and locate the miscalibration threshold \delta^* past which confidence-ranked auditing is \emphworse than random. Two a-priori expectations reverse: \delta^* \emphrises as the budget shrinks, and cross-family correlation is not low—shared difficulty dominates lineage. Five open-weight LLMs show operationally useless (near-constant) confidence, point estimates at or beyond the flip though CIs straddle it; a proprietary model is informative and lands below it. We give a quantitative criterion for \emphvacuous oversight, and replaying policies on recorded traces confirms the ordering.
[AI-24] From Textual Requirements to Microservice Architectures - A Comprehensive Evaluation of LLM -Based Design Synthesis
链接: https://arxiv.org/abs/2607.28307
作者: Danyllo Albuquerque,José Renan,Guillermo Rodríguez,Guillermo Rodríguez,Emanuel Dantas,Ademar França,Mirko Perkusich,Kyller Gorgônio,Angelo Perkusich
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Microservice architectures have become dominant for modernizing monolithic systems, yet identifying appropriate services remains challenging and largely manual. Existing decomposition approaches are predominantly code-centric, limiting applicability in early design stages where only textual requirements are available. Despite advances in Large Language Models (LLMs), limited empirical evidence exists on their ability to synthesize complete microservice architectures from natural-language requirements, including service definitions and inter-service interactions. This study investigates whether an LLM can bridge requirements engineering and architectural design, generating architectures solely from textual requirements and evaluating structural agreement and perceived quality of results. We conduct a mixed-method study using OpenAI o3 under zero-shot (ZS) and few-shot (FS) prompting across two systems (Bookstore, PetClinic), one execution per system/condition. Architectures are evaluated through (i) comparison with reference architectures using precision, recall, and F1-score for service identification and communication recovery, and (ii) a blinded expert assessment of correctness, completeness, modularity, and plausibility, plus open feedback synthesis. OpenAI o3 identifies services with higher agreement under FS prompting (F1 = 0.79 for ZS versus = 0.97 for FS). Communication recovery is more challenging: ZS produces dense architectures with high recall but low precision (F1 = 0.61), while FS improves agreement, reaching F1 = 0.82 and reducing unsupported dependencies. Expert evaluation corroborates these results, with FS architectures perceived as more modular, coherent, and plausible than ZS outputs. OpenAI o3 shows potential for requirements-driven synthesis when guided by exemplar prompting. Results are model- and context-specific from two small systems, not model-independent proof.
[AI-25] MemHarness: Memory Is Reconstructed Not Replayed
链接: https://arxiv.org/abs/2607.28272
作者: Rong Wu,Daocheng Fu,Licheng Wen,Xuemeng Yang,Shu Zou,Jianbiao Mei,Yuxin Wang,Hairong Zhang,Yu Yang,Tao Hu,Cong Zhang,Botian Shi,Pinlong Cai
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 13 figures
Abstract:Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be replayed verbatim, injecting them into the context regardless of whether they align with the agent’s current situation. This ``replay’’ paradigm ignores the gap between the abstract, general nature of stored experience and the concrete, ever-changing states encountered at decision time, frequently causing negative transfer. In contrast, humans rarely recall past experiences verbatim; instead, they reorganize and adapt retrieved memories to fit the present context. Inspired by this, we propose MemHarness, a framework that equips LLM agents to actively harness and reconstruct past experiences based on the present context. At each decision step, a unified policy model critiques and reconstructs the retrieved experience conditioned on the current state, producing context-grounded guidance before acting. This reconstructive ability emerges naturally through end-to-end training with GRPO. Experiments on ALFWorld and WebShop show that MemHarness substantially outperforms pure RL and static memory-augmented baselines, demonstrating strong robustness in out-of-distribution (OOD) scenarios. Furthermore, our analyses reveal that this reconstruction objective not only prevents negative transfer but also serves as latent guidance during training, fundamentally improving the agent’s intrinsic reasoning capabilities.
[AI-26] Agent ic Method for Deterministic Validation of Legacy Code Migration
链接: https://arxiv.org/abs/2607.28271
作者: Andras Ferenczi,Jordan Docherty,Mariya Bessonov,Matthew Findlay,Krishna Lingamneni
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 11 pages, 6 figures
Abstract:Migration of legacy COBOL programs to Java requires extensive testing to ensure correct functionality. This effort is often complicated by the lack of test data and the difficulty of validating all corner cases. In this paper we propose a novel agentic test-synthesis method, the “Locksmith Loop,” which is initiated by preparing two runtime environments: the COBOL source and the generated Java target are each instrumented with mocks and executed off-mainframe on commodity hardware, then an iterative agentic loop performs Witness Search over input mocks to penetrate program branches, followed by parity-preserving mutations. When routing boundaries are reached, an analyzer identifies a Locked Paragraph: a condition preventing deeper exploration. Across three COBOL-Java case studies, spanning two open-source programs and one internal production-like COBOL program and ranging from 430 to 4,114 source lines, Locksmith consistently improved coverage beyond input-search plateaus, reaching nearly complete coverage on the two open-source programs and 91.90% branch coverage on the internal production-like COBOL program. The generated Java matched the COBOL reference under deterministic parity checks in all accepted test cases. Through these findings we demonstrate, to the best of our knowledge, a novel approach for validating agentic coding output using a deterministic oracle.
[AI-27] LLM -Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency
链接: https://arxiv.org/abs/2607.28268
作者: Kostis Michailidis,Dimos Tsouros,Nguyen Dang,Tias Guns
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Combinatorial problems appear in numerous industrial applications. A common approach is to formulate these problems as declarative constraint models that can subsequently be compiled to and solved by a range of back-end solvers. Recent work shows that Large Language Models (LLMs) can produce correct models from natural language, but even a correct model can be expensive to solve because performance remains sensitive to modelling choices. In this work, we investigate whether LLMs can automate performance-oriented model reformulation. Inspired by Automatic Heuristic Design (AHD), we use an evolutionary framework in which an LLM proposes candidate reformulations that are verified and benchmarked against the user-defined baseline model. We compare AHD-adapted search strategies that control which prior attempts, instructions, and measured feedback enter each prompt. Existing retention strategies prioritize recency or performance, but do not explicitly diversify the context. To cover this gap, we introduce Profile-Diverse Retention (PDR), which applies Maximal Marginal Relevance (MMR) to instance-level runtime vectors to retain behaviourally diverse attempts. We systematically evaluate the strategies on eight CSPLib problems using validation-based final model selection. The results show that: (i) iterative reformulation can produce substantial held-out speedups; (ii) strategies that keep the retained context diverse outperform those that retain only recent or the fastest attempts; and (iii) validation-based selection improves the held-out speedup of every strategy.
[AI-28] Operationally Guided Placement-Aware Learning for Industrial Online 3D Bin Packing
链接: https://arxiv.org/abs/2607.28257
作者: Dheeraj Poolavaram,Aanchal Rajesh Chugh,Sebastian Dorn
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The online three-dimensional bin packing problem (3D-BPP) is a longstanding challenge in logistics and industrial palletizing. Recent learning-based methods use a learned policy to select among feasible candidate placements. Performance depends on the candidate generator and representation, especially in industrial settings where packings must be space-efficient, stable, compact, and balanced. However, prior work has mainly optimized the policy, while candidate generation and representation remain largely geometry-driven. We address this gap with OPAL, an operationally guided placement-aware learning framework for industrial online 3D-BPP which combines an Operationally Guided Empty-Maximal-Space generator (OG-EMS), an operational representation for each candidate placement, and a masked ranking policy trained with proximal policy optimization. OG-EMS evaluates multiple anchors within each free-space region and prioritizes low, well-supported, compact, and spatially diverse placements. An xLSTM-based Placement Encoder models dependencies among geometric and operational candidate attributes, while a lightweight recurrent core combines the resulting embeddings with the current item and pallet state to rank feasible actions. On the BED-BPP benchmark, OPAL achieves a mean space utilization of 0.49, with improvements of 15.1% from operationally guided candidate generation and 6.3% from learned ranking, while maintaining robust inference-time performance.
[AI-29] AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability Hallucination and Source Fidelity in Quranic Hadith and Fiqh Knowledge
链接: https://arxiv.org/abs/2607.28237
作者: Muhammad Sajjad Akbar
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Generative Artificial Intelligence (AI) is increasingly used by Muslims for religious guidance, Qur’anic interpretation, Hadith explanation, jurisprudential rulings, and Islamic education. Despite its growing adoption, there is limited empirical evidence on whether current AI systems provide authentic, verifiable, and trustworthy Islamic knowledge suitable for high-trust religious contexts. This study evaluates six leading generative AI systems using fifty realistic open-ended Islamic questions covering Qur’anic interpretation, Hadith, Fiqh, ethics, pastoral advice, and Madhhab-sensitive topics. Responses were collected under real-world conditions from participants in Australia and the United Kingdom and analysed using a mixed-method framework examining domain accuracy, citation verification, hallucinations, jurisprudential consistency, uncertainty handling, source provenance, and geographical variation. The study addresses four research questions: (1) How accurate and authentic are AI-generated responses across major Islamic knowledge domains? (2) To what extent do AI systems produce hallucinations, incomplete citations, or unverifiable religious references? (3) How consistently do models handle jurisprudential disagreement, Madhhab diversity, and uncertainty? (4) Are current AI systems sufficiently reliable for religious guidance, Islamic education, and scholarly research? Overall, current generative AI systems are valuable as assistive tools for introductory Islamic learning but should not be treated as authoritative sources for religious rulings or Islamic research without verification against authenticated primary sources and qualified scholarly expertise. This study provides one of the first comprehensive empirical evaluations of AI reliability within Islamic knowledge, offering practical guidance for researchers, educators, AI developers, and the wider Muslim community. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.28237 [cs.AI] (or arXiv:2607.28237v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28237 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-30] Security of World-Model-Based Embodied AI: A Lifecycle of Threats Defenses and Evaluation
链接: https://arxiv.org/abs/2607.28226
作者: Fazhong Liu,Zhuoyan Chen,Haozhen Tan,Yan Meng,Guoxing Chen,Haojin Zhu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:World models give embodied AI a predictive core: they compress observations into states, simulate action-conditioned futures, and enable planning beyond reactive control. This predictive layer, however, opens a new security boundary-compromise can propagate from data, sensors, prompts, or feedback into physical action. Rather than treating world models as an isolated component, this survey traces threats across their entire lifecycle-from data construction and representation learning, through state grounding and imagination, to trajectory evaluation, execution, and long-term adaptation via memory and tools. We show that familiar attack families: poisoning, backdoors, adversarial examples, sensor spoofing, prompt injection, trajectory manipulation, and supply-chain attacks take on distinct meanings when they corrupt world states, learned dynamics, affordance estimates, or safety costs. We also highlight a duality: world models can serve as runtime safety shields, yet when compromised or over-trusted they generate predictive safety illusions. The survey offers a lifecycle taxonomy, maps existing attacks to world-model security properties, outlines evaluation protocols for safety failures, and structures defenses across provenance, robust grounding, uncertainty-aware prediction, trajectory gating, feedback auditing, and deployment assurance.
[AI-31] Old Tricks New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
链接: https://arxiv.org/abs/2607.28187
作者: Marco Alecci,Francesco Marchiori,Iyiola Emmanuel Olatunji,Tegawendé F. Bissyandé,Jacques Klein
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Software Engineering (cs.SE)
备注:
Abstract:While automated content-moderation systems have become essential for screening harmful content at scale, conventional task-specific classifiers often provide limited policy cov- erage and contextual understanding. Recently, commercial multimodal moderation APIs built on large foundation models have been introduced with the promise of providing broader and more capable safety filters. In this work, we analyze whether this shift also yields more robust image moderation. We conduct a large-scale black-box evaluation on three established commercial image-moderation services and compare their robustness. By evaluating seven simple, model-agnostic image transformations across multiple providers, datasets, harm categories, perceptual-similarity constraints, and transformation intensities, we find that: (1) all three commercial services can be bypassed using inexpensive image transformations that require no gradients, surrogate models, or knowledge of the target system; (2) even fixed transformations such as color inversion and grayscale conversion induce unsafe-to-safe decision changes while preserving content that remains recognizable to humans; (3) their robustness varies substantially across datasets and harm categories, with multimodal content and self-harm exhibiting pronounced vulnerabilities. This yields the conclusion that replacing conventional moderation classifiers with foundation-model-based APIs does not, by itself, provide a reliable security boundary. Such systems must be evaluated under realistic transformations and deployed as one component of a layered moderation pipeline rather than as standalone safety filters.
[AI-32] Persistent Gaussian Perturbations Prevent Oversmoothing in Recurrent Graph Neural Networks
链接: https://arxiv.org/abs/2607.28185
作者: Mostafa Haghir Chehreghani
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Theory (cs.IT)
备注:
Abstract:Oversmoothing is a fundamental limitation of deep graph neural networks (GNNs), where repeated message passing causes node representations to become increasingly similar, eventually collapsing toward a low-dimensional subspace. This phenomenon limits the effective depth of message-passing architectures and motivates the search for mechanisms that preserve representation diversity. In this paper, we study a recurrent graph neural network in which independent Gaussian noise is injected after every propagation step and analyze the resulting architecture as a stochastic dynamical system. Under a standard global contraction assumption on the deterministic update, we prove that the hidden representations form a geometrically ergodic Markov chain admitting a unique invariant probability measure. Our main theoretical result establishes an explicit positive lower bound on the expected stationary Dirichlet energy, proportional to both the noise variance and the spectral gap of the underlying graph. Consequently, the stationary representations cannot collapse onto the constant manifold, providing a rigorous guarantee that asymptotic oversmoothing is prevented in the sense of non-vanishing Dirichlet energy. Our analysis reveals persistent stochastic perturbations as a fundamentally different mechanism for combating oversmoothing, complementing existing deterministic approaches based on residual connections, normalization, and graph rewiring. Finally, numerical experiments on both linear and nonlinear recurrent graph neural networks closely match the theoretical predictions, illustrating the emergence of a stationary distribution and the predicted dependence of the limiting Dirichlet energy on the noise intensity.
[AI-33] Integrating AI into Requirements Quality Learning in Software Engineering Education: A TPACK-Guided Empirical Study
链接: https://arxiv.org/abs/2607.28176
作者: Hansika Ekanayake Mudiyanselage,Rohan Jai Dharmaraj,Malik Abdul Sami,Zheying Zhang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注: 11 pages, 6 figures, 3 tables, presented in the 38th CSEET in Florence, Italy, from July 20-22, 2026
Abstract:The rapid adoption of generative Artificial Intelligence (AI) in software engineering (SE) practice creates a need for pedagogically grounded approaches to AI integration in SE education, especially in conceptually intensive subjects such as requirements engineering (RE). This study examines a TPACK-guided integration of a multi-agent AI tool into a master-level RE assignment on requirements quality analysis. Using a mixed-methods design (N=100; 72 submissions analysed), we examine how structured assignment design shaped students’ AI use, affected their understanding of user story quality criteria, and influenced their perceptions of AI’s benefits and limitations. Results show that students used the AI tool selectively, mainly as support for analysis and evaluation rather than automation. Alignment improvements were most evident for structurally concrete requirements quality dimensions, such as value articulation and testability, while negotiability showed mixed effects. Students reported conditional trust, active refinement, and increased awareness of quality criteria, alongside moderate usability challenges. The findings show that TPACK-guided scaffolding can align AI affordances with pedagogical goals and RE content, offering design guidance for responsible AI integration in RE education.
[AI-34] Agent icASR: Refining Speech Recognition in Real-World Scenarios via an Agent ic Approach
链接: https://arxiv.org/abs/2607.28175
作者: Zixuan Jiang,Binghao Qiang,Jiaying Chi,Yanqiao Zhu,Kai Yu,Xie Chen
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 3 figures, 14 tables
Abstract:Automatic speech recognition (ASR) has achieved substantial gains in transcription accuracy, yet verbatim transcription does not necessarily produce readily usable text. It retains fillers, repetitions, false starts, and self-corrections that increase reading effort, obscure the speaker’s final intent, and propagate unresolved or abandoned content to downstream tasks. Existing spoken-to-written methods process completed audio or transcripts but cannot revise emitted text when later speech changes how preceding content should be interpreted. We therefore formulate Agentic Speech Recognition (AgenticSR), an audio-to-clean-text task that removes disfluencies, resolves self-corrections, and normalizes written form while preserving the speaker’s final intent. AgenticASR implements this task through an ASR–Refiner architecture that repeatedly transforms a bounded active context and replaces its corresponding output span as audio arrives. This enables continual emission and revision over streams of arbitrary duration. We also introduce AASR-Bench, a bilingual benchmark with fine-grained atomic rubrics. Across multiple ASR front ends, AgenticASR attains the highest AASR-Bench scores among evaluated systems. A human–AI agreement study shows that rubric-based judgments align with independent expert assessments. Ablations characterize Refiner capacity, context length, and the quality–latency trade-off between online and offline inference. Together, these results establish AgenticASR as a practical framework for intent-preserving clean transcription during ongoing speech. Code, AASR-Bench, and a demo will be released at this https URL.
[AI-35] Search Strategies for Optimal Classification and Regression Trees
链接: https://arxiv.org/abs/2607.28170
作者: Jacobus G. M. van der Linden,Mim van den Bos,Emir Demirović
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Optimal decision trees (ODTs) are compact, interpretable machine learning models that globally optimize a given objective, but their scalability remains challenging. While recent work has proposed a variety of search strategies to improve scalability, the precise contribution of each strategy remains unclear. To address this gap, we introduce a general algorithmic framework for ODTs that instantiates previously used search strategies and enables the definition of new ones. This provides a common lens through which to understand and compare different strategies, which we use to empirically investigate the effect of 18 search strategies. Compared to the state of the art, the best strategy in our evaluation achieves significantly better anytime performance for classification, and improves runtime by more than an order of magnitude for regression.
[AI-36] Asymmetric Communication: Large Language Models and Language Games
链接: https://arxiv.org/abs/2607.28137
作者: Enzo Fenoglio
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:Contemporary AI discourse attributes to language models properties they cannot bear: general intelligence as substrate-independent cognition, hallucination as cognitive failure, agency as autonomous goal-pursuit, sentience as emergent inner life, alignment as goal synchronization. This paper argues that these are instances of a single category mistake–properties constituted within human communicative practice are projected onto the machine side–and explains its structure. Human-LLM interaction constitutes a language game in which one side bears all normative activity. We call this configuration asymmetric communication since model outputs circulate communicatively, entering further exchanges, without the system undertaking commitments, bearing entitlements, or performing the assessment on which discursive standing depends. Three conditions define the asymmetry: (i) correctness is enforced exclusively by the receiver; (ii) accountability is borne by human participants alone; and (iii) the practical standing of any output depends entirely on human uptake. These conditions are structural, hold independently of capability, and remain unchanged as more powerful models raise the stakes of misattribution. The framework draws on Wittgenstein (meaning enacted in shared practices), Luhmann (communication completed on the receiver’s side), Esposito (algorithmic contingency sufficient for uptake), and Brandom (normative scorekeeping as the source of discursive standing). Applied to all five, it reclassifies each as a receiver-side phenomenon, grounds guardrails as structural necessities rather than manifestations of machine moral agency, and yields an implication for AI governance. Alignment is institutional constraint engineering, not goal synchronization between agents, while responsibility remains with human institutions.
[AI-37] ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
链接: https://arxiv.org/abs/2607.28126
作者: Bingchen Liu,Yuanyuan Fang,Lei Liu,Guangyuan Dong,Xing Fu,Yuanyuan Gao,Shuyue Wei,Xin Li,Xiangtian Meng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existing retrieval-augmented generation systems treat historical logs as a static corpus and retain records without estimating their diagnostic value, failing to report early risk. To this end, we propose ConMem, a contribution-aware memory framework for LLM-assisted equipment inspection, supporting a human-in-the-loop early-risk screening system. Specifically, our ConMem first segments inspection logs into functional evidence units, then estimates each memory unit’s contribution to downstream diagnosis through a Shapley-style estimation, and finally retains high-value evidence under a constrained memory budget. In experiments, we evaluate ConMem on real-world dataset and ConMem achieves 76.0% QA accuracy, exceeding the strongest directly comparable baseline. Relative to the naive 8K-context LLM baselines, it reduces the average number of input tokens by 88.2% and response time by 86.6%. Ablation studies also show that the functional-role-aware segmentation and contribution-based valuation are helping prioritize weak degradation signals for targeted field inspection. Practical deployments further confirm that ConMem retains the weak early signal across three inspection cycles, providing an early-stage seal-wear alert targeted for on-site inspectors.
[AI-38] Information Bottleneck Learning for Faithful Time Series Forecasting Explanations
链接: https://arxiv.org/abs/2607.28124
作者: Xu Zheng,Wei Cheng,Zhuomin Chen,Mo Sha,Jingchao Ni,Dongsheng Luo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 17 pages, 6 figures, 8 tables
Abstract:As forecasts increasingly drive decisions in fields such as energy, transportation, and healthcare, understanding the historical data behind these predictions has become as crucial as the predictions themselves. Although existing interpretable-by-design forecasters reveal their internal structures, they offer no guarantee that these structures faithfully reflect the underlying evidence driving the predictions. In contrast, while faithfulness-oriented methods explicitly verify model behavior, they are almost exclusively designed for post-hoc classification tasks. To bridge this gap, we propose IB-Forecast, an inherently interpretable multivariate time-series forecasting framework. It decomposes forecasting into a learned periodic component and a residual component computed with explainable masks over input tokens. With a budget-constrained information bottleneck, end-to-end optimization enables users to directly control explanation sparsity. With a rigorous faithfulness evaluation protocol, extensive experiments demonstrate that IB-Forecast matches the forecasting error of leading black-box models while providing faithful explanations at no additional inference cost. Furthermore, under a matched sparsity budget, these native explanations consistently surpass gradient-based, occlusion-based, and optimization-based baselines across all evaluated datasets. Ultimately, whereas the native explanations of existing interpretable forecasters exhibit poor faithfulness, IB-Forecast guarantees high explanation fidelity, requiring only 14-20% of the observations to deliver low-error predictions.
[AI-39] BlueprintRepair: Typed Local Edits for Failed Lean Proof Blueprints
链接: https://arxiv.org/abs/2607.28110
作者: Ruslan Khrulev
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, 4 figures, 7 tables
Abstract:LLM-based Lean proving systems increasingly organize a proof as a blueprint: a dependency graph of formal statements. We introduce BlueprintRepair, a repair interface that lets a model change this graph through ten schema-checked local operations. An operation names the node it edits, so the target theorem cannot be changed. Lean checks every applied change, and an accepted repair must declare every blueprint lemma its proof uses. We also construct BlueprintTrace, a benchmark of 142 controlled failures with complete accepted and rejected repair trajectories. We compare typed edits, exact source patches, and complete module rewrites under matched source, feedback, model, and budget, one episode per state and interface. With DeepSeek-V4-Flash, the three interfaces solve almost the same number of the benchmark’s localized failures. Typed repair is the cheapest per solved state (patching is 1.30x as expensive, rewriting 2.06x), and within 10,000 completion tokens per task it reaches almost all of its final coverage, while both free-form interfaces are well behind. A second model, Qwen3.6-Flash, solves fewer states but keeps typed repair cheapest, puts it ahead on the proof-authoring states, and repeats the localized pattern.
[AI-40] Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
链接: https://arxiv.org/abs/2607.28109
作者: Jiawen Tao,Miao Peng,Yaoming Li,Xiaokun Yuan,Mengzhou Wu,Wenhan Yu,Guoan Wang,Nuo Chen,Tong Yang,Maxm Pan
类目: Artificial Intelligence (cs.AI)
备注: 31 pages, 3 figures, 11 tables
Abstract:Synthetic textbook data has improved language model pre-training, but prior work largely treats the benefit as a property of generated content or local rewriting style. We study a different factor: whether related content is organized into coherent book-level documents. We contribute both a scalable synthesis pipeline and controlled evidence that this organization matters. The pipeline retrieves source material from a pre-training corpus, clusters it into topical units, plans hierarchical tables of contents, and assembles source-grounded sections into complete books (our Full setting), yielding 686K textbooks (32B tokens) across 15,000+ disciplines. Replacing natural books in a mid-training mix with this corpus improves downstream performance by +1.09 on average. Controlled comparisons then disentangle the relevant design factors. A content-matched Split condition holds generated text and tokens fixed but treats each section as an independent document; Full’s +1.02 mean gain isolates document packaging. A length-matched RandomConcat control that joins sections from different books remains below Full, ruling out document length alone. A retrieval-pool-matched Rephrase condition independently rewrites individual retrieved documents under the same audience-by-style scheme, without clustering, TOC planning, or book assembly; Full’s +1.17 gain demonstrates the value of structured synthesis. On Llama3-8B, Full likewise outperforms both RandomConcat and Natural Books, supporting book-level organization as a useful axis for synthetic pre-training data design.
[AI-41] MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
链接: https://arxiv.org/abs/2607.28103
作者: Dongyi Liu,Haixing He,Xiaobao Wu,Jia Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Memory-augmented LLM-based agents are vulnerable to memory injection attacks: Agents may retrieve poisoned memory from attackers, which diverts their behavior from initial user intent and finally causes task failure. However, existing defense mechanisms either incur high computational cost or suffer from information redundancy in multi-turn contexts. To address these challenges, we propose Memory Intent-Aware Neural Denoising(MIND), a lightweight defense framework for memory injection attack. Our preliminary analysis reveals that benign and poisoned trajectories exhibit distinguishable relationships between the initial user intent and subsequent behavior. Building on this observation, MIND employs an intent-aware Information Bottleneck(IB) to extract compact intent–behavior representations from the initial intent and turn-level behavior. The IB preserves intent-relevant cross-turn attack signals while filtering task-irrelevant and repetitive information, and a lightweight detector identifies malicious memories from the resulting representations. As such, MIND mitigates information redundancy in multi-turn contexts while avoiding the overhead of repeated LLM auditing. Extensive experiments show that MIND reduces attack success rates while preserving task accuracy and inference efficiency. Notably, on ReAct-StrategyQA, MIND reduces mean ASR-r and ASR-a by 55.4% and 55.3%, respectively, while matching the undefended agent in average accuracy and latency.
[AI-42] An Instrument to Evaluate Governance Proposals: AI Policy Analysis at Scale
链接: https://arxiv.org/abs/2607.28094
作者: Paulo Carvao,Claudio Mayrink Verdun,Isabel Adler,Jeffrey Zhou
类目: Artificial Intelligence (cs.AI)
备注: 48 pages
Abstract:This paper introduces a policy analysis framework for systematic, transparent assessment of AI governance proposals in an evolving and contested regulatory landscape. AI policy debates often collapse into binary positions that obscure underlying tradeoffs and normative assumptions. The framework structures policy analysis around multiple policy attributes, allowing users to surface priorities and tensions without prescribing outcomes. We use a mixed-methods approach that integrates qualitative insights from subject matter experts with computational text analysis to inform the design of policy attribute rubrics. This quantifies the relative emphasis of different policy objectives and presents them through comparative visualizations that support interpretability and cross-policy comparison. The paper also examines the use of commercial LLMs for rubric-based policy analysis, benchmarking their outputs against a domain-trained rubric-calibrated model with explicitly defined analytical assumptions. Rather than assessing policy effectiveness or desirability, the framework focuses on relevance and alignment across attributes. By making analytical assumptions explicit, including attribute selection, rubric construction, and weighting schemes, the framework enables users to evaluate whether its embedded priorities align with the users’ own normative commitments. The approach is jurisdiction-agnostic and intended to support policymakers, analysts, and researchers navigating complex AI governance environments. Contributions: (1) multidimensional policy assessment through empirically grounded rubrics that surface tradeoffs rather than resolving them; (2) a transparent hybrid methodology combining feedback from subject-matter experts with computational validation; and (3) use of domain-trained rubric-calibrated models as a benchmark for comparing different general-purpose large language models. Comments: 48 pages Subjects: Artificial Intelligence (cs.AI) ACMclasses: K.4.1; I.2.7 Cite as: arXiv:2607.28094 [cs.AI] (or arXiv:2607.28094v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.28094 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-43] PerturbMap: Cross-Context Transfer of Single-Cell Perturbation Responses
链接: https://arxiv.org/abs/2607.28090
作者: Panpan Cui,Yiqi Liu,Wenhao Sun
类目: Artificial Intelligence (cs.AI)
备注: 14 pages, 5 figs
Abstract:Single-cell perturbation atlases rarely measure every intervention in every cellular context: a query perturbation is often observed in one or more source contexts but missing in the recipient context where its effect is needed. Ignoring those measured responses discards query-specific experimental evidence, whereas copying or weakly calibrating them across contexts risks transferring the wrong signal. We propose PerturbMap, which predicts a missing recipient-context effect by combining a recipient-local low-rank base with accepted proposals that transport the same perturbation’s measured source responses through source-to-recipient ridge experts fit on paired training perturbations, with proposal weights determined by route reliability estimated on validation anchors. On the Perturb-CITE-seq melanoma cohort, PerturbMap improves full-effect MSE by 4.1% over a recipient-local low-rank base and achieves lower MSE than FedAvg, zero-response, raw-copy, calibrated-copy, and identity-shuffled affine controls. It remains within 2.82\times10^-6 MSE of our centralized token-matched pooled reference, which uses a stronger training interface. A condition-mean specificity diagnostic shows the same direction: same-recipient top-10 counterpart retrieval by cosine increases from 74.5% for the low-rank base to 80.5% for PerturbMap.
[AI-44] Diversifying Personalized Research Ideation against AI-Induced Homogenization
链接: https://arxiv.org/abs/2607.28087
作者: Rui Xu,Yunke Wang,Linwei Tao,Wenjie Xuan,Yong Luo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:AI-assisted research ideation has emerged as a promising paradigm for accelerating scientific discovery, with systems now capable of generating research directions conditioned on papers, topics, or lightweight researcher contexts. Yet current systems largely optimize individual suggestions in isolation. This leaves two blind spots. First, coarse researcher representations may elicit mainstream directions that appear broadly feasible, but lack sufficient researcher-specific grounding. Second, independent recommendations can concentrate a community’s portfolio around recurring high-probability themes. To address these blind spots, we propose DivAlign, a four-stage pipeline for alignment-preserving de-homogenization. DivAlign extracts fine-grained researcher profiles, generates profile-conditioned candidate directions, scores them along three alignment dimensions (Executability, Comprehensibility, and Growth Potential), and surfaces researcher-local directions while reducing redundancy across the community portfolio. On a benchmark we construct from 95 AI researchers across five subfields, DivAlign reduces community-level redundancy while preserving researcher-direction fit. Compared with coarse single-shot ideation, it lowers average pairwise similarity from 0.331 to 0.294 and nearest-neighbor similarity from 0.704 to 0.608. Compared with the independent top-choice variant, DivAlign reduces nearest-neighbor similarity from 0.663 to 0.608 while retaining 99.9% of the researcher-direction fit score. Code and data are available at this https URL.
[AI-45] Distilling Answer Set Programming Theories from Large Language Models
链接: https://arxiv.org/abs/2607.28086
作者: Nelson Higuera Ruiz,Markus Hofmarcher,Claudiu Leoveanu-Condrei
类目: Artificial Intelligence (cs.AI)
备注: Accepted at NeSy 2026
Abstract:Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop. The protocol is dataset-agnostic: with a single prompt and an empty file as the starting point the model is given a 1-hour time limit to derive a complete theory. We chose VQA as the application domain, three benchmarks (CLEVR, GQA, CLEVRER), as these are publicly available and non-trivial. In order to study the model scale required for solving this task we nine different models: four frontier (Claude Sonnet 4.6, Claude Opus 4.7, GPT-5, DeepSeek V4 Pro), two mid-tier (DeepSeek V4 Flash, gpt-oss-120b), and three open-weights (qwen3.6-27b, gpt-oss-20b, qwen3.5-9b). Three of four frontier models reach 100% on CLEVR and 92.8%-98.8% on GQA; on CLEVRER, Sonnet, Opus, DeepSeek V4 Pro score 92.7%-95.3%. GPT-5 reaches 98.7% on CLEVR but drops to 41.8% on GQA and to 86.7% on CLEVRER. Adding handwritten reference theories from other datasets moves the other three frontier models by at most +/-3.4 pp but reduces GPT-5’s accuracy by 3-19 pp. We release the code, prompts, and theories distilled.
[AI-46] Chem World: A Large-Scale Benchmark and Physics-Informed Framework for Trustworthy Chemical Property Prediction
链接: https://arxiv.org/abs/2607.28079
作者: Tianyou Bai,Huan Wang,Mingchen Gao,Fangyue Lin,Pinze Ren,Zhenlin Zhao,Siming Dong
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Chemical property prediction plays a critical role in accelerating scientific discovery in chemistry, materials science, and drug development. However, existing benchmarks often suffer from limited task diversity, fragmented datasets, and inconsistent evaluation protocols, making it challenging to systematically assess the reliability and generalization of AI models. In this work, we introduce Chem World, a comprehensive benchmark for chemical property prediction that integrates 17 diverse chemical datasets with over 800,000 molecular samples, covering various properties including density, electrical conductivity, solubility, and other molecular characteristics. Chem World provides a unified platform for evaluating AI models across multiple property prediction tasks. Furthermore, we propose Mixture-PINN, a physics-informed neural network based prediction framework that incorporates chemical prior knowledge into data-driven learning, improving the accuracy, robustness, and reliability of chemical property prediction. Extensive experiments on Chem World demonstrate the effectiveness of our approach compared with existing methods. By combining large-scale standardized evaluation with physics-informed learning, Chem World establishes a foundation for developing trustworthy AI systems for computational chemistry and advancing AI-driven scientific discovery.
[AI-47] Group-Reflective Self-Distillation for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2607.28076
作者: Binbin Zheng,Zijun Xie,Guanqun Zhao,Enlei Gong,Xing Ma,Xiaoliang Fu,Zeyu Chen
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidental choices entangled in the same outcome signal. Existing agentic self-distillation methods enrich sparse supervision with natural-language skills, but skills retrieved externally or extracted from a single trajectory by stronger models may mismatch current experience, exceed the policy’s capability, or remain path-specific. We propose Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy’s own verified rollouts. For each prompt, the policy reflects on each verified trajectory in an on-policy group, and a stop-gradient snapshot contrasts the resulting reflections from successful and failed rollouts to construct group-level privileged guidance. Conditioned on this guidance, a self-teacher refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction. Experiments across multiple agentic environments and model scales demonstrate that GRSD consistently outperforms competitive baselines and generalizes more effectively to unseen tasks.
[AI-48] mporal Poisoning: Clean-Label Backdoors via Event Redistribution in SNNs
链接: https://arxiv.org/abs/2607.28075
作者: Roberto Riaño,Gorka Abad,Stjepan Picek,Aitor Urbieta
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Backdoor attacks on Spiking Neural Networks (SNNs) have primarily assumed dirty-label poisoning, in which triggered training samples are relabeled to an attacker-selected class. We study clean-label temporal poisoning, where a fixed timestamp transformation is applied only to the target-class training streams, leaving their labels unchanged. The transformation preserves the per-pixel, per-polarity event count exactly, making clean and triggered samples identical after temporal aggregation while altering the sequence processed by the SNN. Across three neuromorphic datasets and both convolutional and transformer-based victims, the attack reaches an ASR of 1.00 in the strongest configurations. We analyze the attack through poison-budget and trigger-shape ablations and evaluate established backdoor defenses adapted to spiking models. Defenses that collapse the time axis before inspection are blind by construction, while feature-space methods detect the poison only in selected settings. Our model-free detector, based on per-step event mass, detects the evaluated temporal transformations, demonstrating both the limitation of rate-collapsed defenses and the boundary of the attack’s stealth. To our knowledge, this is the first clean-label backdoor attack evaluated on SNNs and neuromorphic event data.
[AI-49] Echoverse: Deep Evolving Environments for Training Computer-Use Agents at Scale
链接: https://arxiv.org/abs/2607.28074
作者: Yash Pandya,Sahil Gupta,Sarthak Harne,Archana Yadav,Kavyansh Chourasia,Hussein Mozannar,Vibhav Vineet,Sara Abdali,Corby Rosset,Yash Lara,Ahmed Awadallah,Ece Kamar,Akshay Nambi
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so synthetic environments stand in for them. Recent pipelines generate such environments in bulk, which moves the bottleneck from how many exist to what is inside each one. The returns, we find, come from three properties: how much behavioural depth an environment carries, whether it targets the interaction an agent actually fails, and whether it improves alongside the model. We present Echoverse, which compiles specifications into stateful applications whose tasks are graded against the application’s own database, and a co-evolution loop that reads every graded rollout twice: as repairs to the environment, its tasks and its verifier, and as training signal for the model. Trained on twelve such environments, a 9B model improves from 36.5% to 67.1% across fourteen evaluation splits, within fourteen points of the much larger frontier model that taught it. We examine each property in turn. On the same domains, shallow environments push live-site accuracy below the base model ( 80.0 \to 75.0 ) while deep ones raise it ( 80.0 \to 85.0 and 48.0 \to 65.0 ); drilling one interface control across many renderings transfers to held-out widget families and to the open web; and repairing a single environment lifts the model trained on it from 16.2% to 38.5% . The same worlds serve as reinforcement-learning environments, where a reward combining the grounded verifier with a dense per-step judge raises held-out score from 58.8% to 68.0% . We release four environments as a benchmark, with their applications, seed data and grounded graders. Code: this https URL
[AI-50] SemPIC: Learning Semantic Position-Independent KV Caches
链接: https://arxiv.org/abs/2607.28069
作者: Hui Xie,Peng Xiao,Yutong Deng\textsuperscript,Shuoran Dou,Jian Yang,Jinyang Guo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document orders. Prefix caching cannot exploit this reuse, while position-independent caching (PIC) remains unreliable because independently compiled KV states lack the future context in which they will be consumed. Our diagnostics show that a learned boundary-conditioned baseline sharply reduces attention deviation near reusable-block boundaries but leaves interior and task-level residuals, motivating adaptation of the document representation itself. We present \emphSemPIC, which trains a LoRA-enabled Writer to compile native per-layer document KVs through behavioral distillation while retaining the pretrained decoder as an unchanged Reader. Adaptation is confined to offline cache construction, preserving the standard KV interface and cache-hit decoding path. We further introduce KV Gradient Checkpointing, which reduces peak training memory without severing gradients through cached KVs. Across three models and four tasks, SemPIC raises mean micro-F1 over KV Packet from 0.53 to 0.60, approaching Full Recompute at 0.62.
[AI-51] IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD
链接: https://arxiv.org/abs/2607.28050
作者: Nianchen Deng,Jiaxin Ai,Tao Hu,Shu Zou,Yurui Dong,Siqi Li,Xinyu Cai,Xuemeng Yang,Licheng Wen,Hongbin Zhou,Hairong Zhang,Pinlong Cai,Botian Shi
类目: Artificial Intelligence (cs.AI)
备注: 14 pages
Abstract:Automating industrial CAD design and manufacturing places distinctive demands on multimodal foundation models: the model must see engineering drawings and 3D geometry screenshots, write correct parametric-modelling scripts and Windows COM API code, and cover the full range from single parts to assemblies. General-purpose multimodal models fall short on these tasks, while single-task fine-tuning is too narrow to support the diverse calls that upper-layer agents issue. We build IndustryForge-27B on top of Qwen3.5-VL-27B by curating and integrating six industrial-CAD sub-corpora totalling \sim 52k multimodal samples—CAD Visual QA (CAD-VQA), parametric CAD code (text2cadquery), assembly-level CAD code (text2cadquery-assembly), and three COM sub-corpora for Inventor / SolidWorks (com_2d / com_3d / com_assembly)—and training with a unified multi-task SFT recipe. Across four CAD-domain benchmarks IndustryForge-27B lifts the base model by +33.65 ~pp on average and outperforms the strong closed-source model GPT-5.4 on all four; across eleven general-capability benchmarks it retains, and slightly improves upon, the base model ( +1.56 ~pp mean, no catastrophic forgetting). IndustryForge-27B will serve as the common substrate for downstream industrial-agent projects, providing a unified starting point for a full-stack industrial agent that spans from CAD design to industrial-software operation, from parts to assemblies, and from single-shot generation to closed-loop self-improvement.
[AI-52] SKILL-KD: Contrastive Skill Distillation for LLM Agents
链接: https://arxiv.org/abs/2607.28048
作者: Qiming Shi,Yibo Dou,Jiawen Zhu,Yulong Tao,Linbo Jin,Zhaolu Kang,Yunfan Zhou,Di Weng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entries, or direct summaries of successful demonstrations. This creates a mismatch for weaker student agents: when a student fails because it lacks task knowledge or operational strategy, its failed trajectory may not contain enough evidence to infer the missing behavior, while the teacher trajectory may be too implicit to be internalized as reusable guidance. We propose SKILL-KD, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities. Given a student failure and the teacher trajectory on the same task, SKILL-KD distills their actionable discrepancy into a textual skill patch, evaluates the patch by re-running the student, and iteratively refines the patch when the student still fails. To prevent repeated local updates from causing skill drift, SKILL-KD further maintains trace-linked edit histories and performs Drift-Aware Skill Consolidation, deciding whether each patch should add a new rule, delete or modify an existing rule, or be skipped. Across five agent benchmarks and two student settings, SKILL-KD consistently improves frozen student agents over fixed-model adaptation baselines.
[AI-53] DataClawEval: A Benchmark for Data Engineering Agents in Real Industrial Harness
链接: https://arxiv.org/abs/2607.28033
作者: Debin Meng,Jiaming Yang,Zefang Zong,Tengyue Xu,Haining Xie,Yang Li,Peng Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models (LLMs) and LLM-based agents are increasingly being deployed to automate complex workflows, promising to revolutionize data management and processing. However, existing benchmarks predominantly focus on simplified Text-to-SQL translation or data analysis, leaving the critical and complex domain of end-to-end data engineering largely unexplored. To bridge this gap, we introduce DataClawEval, the first comprehensive benchmark designed specifically to evaluate the end-to-end task completion capabilities of autonomous agents in real-world data engineering scenarios. Built upon production-grade code authored by professional enterprise data engineers, it comprises 100 rigorous, end-to-end tasks spanning five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino, and FlinkSQL. Rather than non-deterministic LLM-as-a-judge scoring, each task is executed within a case-specific, isolated sandbox and graded by deterministic, rule-based scripts. Evaluating 16 frontier agents exposes critical limitations: The strongest model attains only 74.9 overall, and no single model dominates, as each excels on a different engine, revealing strict domain specialization rather than omnipotent proficiency. Thus, autonomous data engineering remains a formidable, unresolved challenge. We release our dataset, containerized environments, and deterministic evaluation scripts at this https URL
[AI-54] Scaling Lock-In and Proxy Compliance: A Political Economy of Responsible AI AAAI
链接: https://arxiv.org/abs/2607.28023
作者: Florian A. D. Burnat,Brittany I. Davidson
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT); Theoretical Economics (econ.TH)
备注: Accepted at AAAI/ACM Conference on AI, Ethics, and Society (AIES '26)
Abstract:AI accountability at scale is an institutional problem: who can observe, verify, and change deployed systems. We develop a sequential political-economy model in which an AI vendor chooses auditability and substantive mitigation, a deployer monitors after adoption while facing switching costs, and enforcement depends on verifiable evidence. Anticipating the deployer’s monitoring response, the vendor may stop at an observable procurement floor while mitigating below the social first best, producing a proxy-compliance equilibrium. We characterize the unique interior equilibrium and the corner in which harm is fully mitigated. Independent audit rights raise enforcement exposure directly; portability restores deployer leverage; incident reporting adds a regulator-visible evidence channel; and outcome-linked liability creates incentives that do not depend on vendor-controlled detection. The results explain why documentation and standardized evaluations can coexist with persistent post-deployment harms, and generate testable implications for monitoring, mitigation, and the gap between formal compliance and operational outcomes.
[AI-55] Flux-OPD: On-Policy Distillation with Evolving Contexts
链接: https://arxiv.org/abs/2607.28022
作者: Yuran Wang,Zekun Wang,Bohan Zeng,Ruixu Zhang,Wenxuan Liu,Liu Yang,Yifan Dai,Yang Shi,Bozhou Li,Chengzhuo Tong,Daili Hua,Yuanxing Zhang,Wentao Zhang
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.
[AI-56] MMLDSum-LLM : Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
链接: https://arxiv.org/abs/2607.28006
作者: Xianpeng Zhang,Jiahua Yang,Dongyu Chen,Lei zhang,Jian Ma,Xu guohuan,Haonan Lu,Tianhuang Su,Chuangchuang Wang,Kai Tang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal long documents are core carriers of professional knowledge, where critical evidence is sparsely distributed across paragraphs and modalities. This easily causes key information omission and cross-modal hallucinations in summarization by multimodal LLMs. These issues stem from attention drift in long-range dependency modeling and gaps in inter-modal alignment. To address this, we introduce MMLDSum-Bench, a high-quality benchmark for multimodal long-document summarization, covering multiple domains, context-length scales, and visual-textual modality distributions. We further propose MMLDSum-LLM, a reproducible two-stage training framework that combines supervised fine-tuning with visual-alignment weighted loss and keyword-aware weighted loss, followed by GRPO with a multi-objective reward (keyword coverage, image-text alignment, ROUGE, and length control). Extensive experiments on MMLDSum-Bench, comparing against leading closed-source and open-source multimodal models under a unified evaluation protocol - including LLM-as-a-judge scoring, atomic-claim precision/recall, image-text alignment (ITA), and ROUGE - demonstrate that our approach significantly improves key-information coverage and cross-modal consistency.
[AI-57] SKIMIX: Multi-Agent Harness-Time Scaling with Skill Mixture for Dynamic Harness Engineering
链接: https://arxiv.org/abs/2607.27994
作者: Jia Luo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:AI agents increasingly rely on large skill libraries, but selecting, combining, and maintaining skills remains difficult. We propose SKIMIX, a multi-agent framework in which agents with different skill portfolios collaborate through iterative refinement. SKIMIX combines embedding-based skill retrieval, submodular anti-dilution routing, and adaptive skill evolution. Across six reasoning benchmarks, multi-agent collaboration substantially improves open-ended mathematical reasoning but offers limited or negative gains on multiple-choice tasks. Agent-count scaling is non-monotonic, and most improvements arise during the first refinement round. These results show that task characteristics determine whether skill-level ensembles help and provide practical guidance for scalable agent design.
[AI-58] Driving up Inference Energy on SNNs: Per-Sample and Universal Sponge Attacks
链接: https://arxiv.org/abs/2607.27990
作者: Spyridon Raptis,Haralampos-G. Stratigopoulos
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Spiking Neural Networks (SNNs) communicate through sparse binary spike events rather than dense activations, enabling energy-efficient inference on neuromorphic hardware and motivating their use in always-on, battery-powered edge systems. We show that this same efficiency advantage creates a distinct security risk: sponge attacks can increase inference-time spike activity and synaptic workload, inflating energy consumption while remaining difficult to detect through correctness-based monitoring alone. Prior input-space efficiency attacks on SNNs have focused on per-sample optimization, primarily in rate-coded settings. We extend this threat to native event-based binary inputs and study two attack models. First, we develop a per-sample sponge attack that crafts a custom adversarial spike train for each input via gradient-based optimization. This attack increases per-inference SynOps by 1.5-2.6x on three SNN models for the NMNIST, SHD, and IBM DVS Gesture datasets, while preserving the predicted class on at least 98% of evaluated samples. Second, to the best of our knowledge, we introduce the first universal sponge attack for native event-based SNN inputs: a fixed binary perturbation computed offline and applied via XOR to all subsequent inputs. Although weaker, it still inflates SynOps by 1.09-1.24x across all three datasets and represents a more realistic deployment threat because it requires no per-input optimization. Mapping SynOp inflation to estimated Loihi-1 energy yields per-inference overheads from 14 \mu J to 13.24 mJ. These results show that native event-based SNNs are vulnerable to practical input-space efficiency attacks, and that reusable universal perturbations can accumulate into meaningful battery drain in continuously deployed edge systems.
[AI-59] Share the Judge Learn the Deferral: Where Specialization Helps LLM Evaluation
链接: https://arxiv.org/abs/2607.27984
作者: Weining Zhang
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 4 figures, 5 tables
Abstract:Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator’s weights, or into the rule that decides when its judgment can be trusted? We study 99,952 public, rubric-conditioned examples. Supplying the correct rubric improves locked-test accuracy by 2.11 points over a response-only control; replacing it with an unrelated rubric costs 2.66 points. Dividing the same training corpus among eight criterion-family LoRA judges, however, loses 10.05 points and cuts audited coverage at a 5% risk target from 24.44% to 5.43%. Matching the bank’s stored capacity with one rank-64 adapter does not reproduce this loss. Nor is the result explained by learning rate or optimizer steps. Initializing the family adapters from a shared, trained judge recovers test accuracy to 76.85%, 19.94 points above scratch training at the same learning rate (95% interval 18.88-21.02). The result changes when specialization governs deferral rather than judgment. On RewardBench 2, learned correctness heads route examples through a 0.6B-4B-8B cascade without changing any reward score. Across 20 locked repartitions, the cascade attains 89.40% accuracy, compared with 84.75% for 8B alone, at 0.415 normalized parameter compute. Every run passes an exact one-sided 95% risk audit; margin-based rules remain near 84.8% accuracy while using at least 0.94 compute. These results suggest a qualified design rule: share the learning of judgment until there is enough data to justify a split, and place domain-specific adaptation in an audited release boundary.
[AI-60] APO: Transition-Aware Policy Optimization for LLM Agents
链接: https://arxiv.org/abs/2607.27973
作者: Cong Li,Peixi Peng,Yisen Zhao,Xinyu Hu,Shudong Liu,Zhan Su,Zhuojian Li
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 16 pages, 5 figures
Abstract:Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing to fully exploit another class of inherently dense supervisory signals naturally present during online interaction: environmental feedback following action execution. Recent theoretical studies suggest that generalization in multi-step, goal-oriented tasks hinges on predictive knowledge of environmental consequences. Inspired by this, we propose TAPO: Transition-Aware Policy Optimization for LLM Agents, a unified training framework that alternates between policy optimization and transition supervision. Beyond standard RL updates, TAPO repurposes rollout data to apply action-conditioned next-observation prediction supervision on a shared backbone model. This approach enhances the model’s sensitivity to environmental transition dynamics and action consequences while concurrently optimizing the policy. It serves as a computationally lightweight, plug-and-play enhancement module for existing agent RL algorithms, requiring no additional expert data, extra sampling costs, or inference-time overhead. We conduct systematic experiments on WebShop and ALFWorld, integrating foundation models of various scales with different policy optimization algorithms. Empirical results demonstrate that TAPO consistently improves task performance over pure policy optimization baselines.
[AI-61] MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation ACL2026
链接: https://arxiv.org/abs/2607.27967
作者: Dawei Wang,Di Zhao,Xinyuan Liu,Marci Chi Ma,Xiaoyang Liu,Chengming Zhou,Gary Ushaw,Richard Davison
类目: Artificial Intelligence (cs.AI)
备注: ACL 2026 Main
Abstract:Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contribution-based pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.
[AI-62] Specification-Guided Synthesis of Deadlock-Free Communication Protocol Refinements with Large Language Models
链接: https://arxiv.org/abs/2607.27964
作者: Yang Li,Ping Hou,Nobuko Yoshida
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Ensuring behavioural correctness in communication protocols is a central challenge in distributed software systems, as subtle inconsistencies can lead to deadlocks. In such settings, protocol refinement - the safe substitution of a protocol that preserves correctness and compatibility with other components - is essential. Large language models (LLMs) have demonstrated strong capabilities in code generation and program synthesis, yet lack mechanisms to reliably produce outputs with correct behaviour. Formal specification approaches, such as multiparty session types (MPST), offer rigorous guarantees, including deadlock freedom, but provide limited support for automatically constructing protocol refinements. In this paper, we present Syntropy, a framework for synthesising protocol refinements guided by MPST specifications and LLMs. It incorporates refinement constraints directly into the generation process, ensuring the generated variants satisfy these guarantees. Our comprehensive evaluation indicates that Syntropy achieves 95.6%-99.5% validity while maintaining high syntactic correctness, and produces diverse, non-trivial refinements across multiple LLMs.
[AI-63] Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLM s
链接: https://arxiv.org/abs/2607.27951
作者: Pingyu Wu,Lingyao Zhu,Weiming Zhang,Nenghai Yu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model safeguards decide whether to answer before seeing how an answer will be used. This creates a basic problem for dual-use tasks: the same answer can help an authorized professional or an attacker, while an attacker can imitate a benign request and interaction history. We separate the capability released by the model from the evidence available about downstream use. When that evidence is copyable, we derive the exact worst-case floor on attacker assistance while preserving useful answers. The result yields a safety trilemma: Useful Capability, Reliable Safety, and Open Access cannot coexist. We then show how a trusted credential can complement existing safeguards by adding hard-to-copy information that predicts actual downstream use, and identify the stronger condition needed to eliminate the floor. Evidence from dual-use evaluations, adaptive attacks, and deployed trusted-access programs supports the practical relevance of these conditions.
[AI-64] From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
链接: https://arxiv.org/abs/2607.27937
作者: Xu Xia,Jinghua Piao,Min Yang,Xiaochong Lan,Jiaju Chen,Yong Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Recent work on LLM agents is shifting from external capability elicitation to capability internalization, enabling agents to retain useful skills without retrieval at inference time. On-policy self-distillation (OPSD) offers a promising direction, but many existing methods typically supervise students by scoring actions along student-generated trajectories. Such supervision has two limitations: teacher preferences are not validated by environment outcomes, and action-level scores underuse information from student rollouts, teacher rollouts, and their behavioral relationship. We therefore advocate outcome-verified teacher supervision and comparative learning over teacher-student trajectories. Based on this view, we propose Outcome-Verified Comparative Self-Distillation (OVCSD). OVCSD organizes failed student rollouts into a prefix tree, adaptively invokes a skill-conditioned teacher from student-reached states, and retains only outcome-verified successful continuations. It then applies localized comparative learning at the first state-aligned divergence and distills the post-divergence teacher suffix to transfer completion behavior. Experiments on ALFWorld and WebShop across three model scales show that OVCSD consistently outperforms skill-free RL and existing self-distillation baselines, achieving up to 29.7 and 5.4 absolute success-rate gains over the strongest baselines on ALFWorld and WebShop, respectively, while adding less than 3% privileged interaction during training.
[AI-65] Shapes from Examples: Foundations of Shape Learning in Recursive SHACL ISWC26
链接: https://arxiv.org/abs/2607.27934
作者: Bente Gortworst,Cem Okulmus,Magdalena Ortiz,Anni-Yasmin Turhan
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注: full version of a paper accepted at ISWC26
Abstract:SHACL shapes enable data graph validation, making automatic shape learning essential for knowledge graph applications. We investigate the well-known fitting approach to this task: given sets P and N of positive and negative example nodes from an input graph, compute a shape expression C, possibly using shape names defined in a recursive shape catalogue, that validates at every node in P and none in N. We focus on the case where C is written in a core fragment of SHACL corresponding to the Description Logic ELI. For the catalogue, we consider the well-founded, stable, and supported semantics. We address fitting existence and most specific fitting computation, establish tight exponential-time upper bounds for both problems, and obtain polynomial bounds for relevant special cases.
[AI-66] he Geometric Nature and a Free Proxy for Flow-Matching Uncertainty
链接: https://arxiv.org/abs/2607.27933
作者: Ziyang Rao,Yiren Zhao,Weiyu Guo,Ben Fei,Yandong Guo,Hui Xiong
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it does not explicitly expose its inherent uncertainty, producing faulty action chunks even when it misinterprets the scene or encounters out-of-distribution (OOD) inputs. Therefore, determining when an FM-generated action can be trusted is essential for safe deployment, yet existing uncertainty estimation methods on real-time control suffer from several issues: extra training budget, high computational overhead, and low generalization ability. In this work, we provide a geometric interpretation of FM uncertainty in the velocity field, showing that uncertainty manifests as deviation from an ideal affine-isotropic contraction field. Building on this observation, we introduce denoising acceleration ( \mathrmaccel ), a highly-generalizable and cost-free uncertainty proxy that measures the bending of the denoising trajectory from a single forward pass, without additional model evaluations, training, or resampling. We theoretically and empirically demonstrate that \mathrmaccel is a faithful proxy for FM uncertainty and further test its utility in online failure detection. Results show that \mathrmaccel identifies failing rollouts well before termination, matching or even outperforming costly resampling- and training-based baselines across settings under realistic deployment budget. Code and demos available at: this https URL.
[AI-67] Meta-Task: Turning Terminal Task Synthesis into a Terminal Task for Scalable Agent Training
链接: https://arxiv.org/abs/2607.27929
作者: Zhihong Pan,Jiyuan He,Kai Zhang,Yupeng Han,Ze Liu,Yuze Zhao,Yongcong Ye,Zhaohua Yang
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 5 figures
Abstract:Training terminal agents at scale requires diverse, verifiable terminal tasks and high-quality interaction trajectories, yet acquiring such data remains a significant challenge. Existing synthesis methods face two key limitations: (1) weak reliability caused by the disconnect between task generation and real execution, and (2) limited diversity and scalability due to dependence on existing repositories. We propose Meta-Task, a framework that redefines terminal task synthesis as a Terminal-Bench-format task itself: an agent operates within a real container environment to iteratively generate, execute, and verify tasks, so that synthesized components are checked for internal consistency and executability within the generation loop itself. Building upon this, we decouple the target task requirements along multiple dimensions, introduce a multi-phase mechanism that dynamically designs novel task specifications before producing the actual tasks, and incorporate optional external material support to enhance diversity and realism. We additionally apply LLM-as-Judge filtering to ensure the quality of the final training data. Experiments on Terminal-Bench 2.0 show that fine-tuning on only 3,221 Meta-Task synthesized trajectories achieves 22.5% and 31.8% Avg Pass@1 for Qwen3-14B and Qwen3-32B respectively, outperforming concurrent approaches with significantly less training data.
[AI-68] One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs
链接: https://arxiv.org/abs/2607.27917
作者: Enyi Shi,Fei Shen,Chuancheng Shi,Linxia Zhu,Shuyi Miao,Jinhui Tang,Tat-Seng Chua
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:As large vision-language models (LVLMs) are deployed globally, the combination of multilingual instructions and visual information makes malicious attacks more covert and sophisticated than ever before. However, existing methods isolate language and modality defenses, which, coupled with the scarcity of safety data and high fine-tuning costs, makes it difficult for models to defend against compound attacks. To address this severe challenge, we propose a neuron-level cross-dimensional safety alignment framework driven by modality- and language-shared safety neurons (MLS-Neurons). First, we identify monolingual and unimodal safety neurons by comparing responses to harmful and benign samples, quantifying functional saliency through activation strength and downstream impact. Then, by intersecting these unimodal neurons within each language, we extract modality-shared safety neurons (MS-Neurons) responsive to both visual and textual risks, bridging the safety representation gap between modalities. Furthermore, using English as a semantic anchor, we intersect MS-Neurons across languages to identify modality- and language-shared safety neurons (MLS-Neurons), serving as key defenses against compound attacks. Finally, we update only this minimal subset of shared neurons (~0.03% of parameters), transferring English-only safety supervision to multilingual and multimodal scenarios. Extensive experiments show that our method significantly outperforms state-of-the-art approaches across diverse multilingual and multimodal safety benchmarks while preserving general utility.
[AI-69] A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
链接: https://arxiv.org/abs/2607.27910
作者: Xiangyu Yin,Tora Bodin,Rohan Menon,Chih-Hong Cheng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen decoder layer. We compare five defence candidates across 15 model and layer cells from four architectural families under a magnitude controlled protocol that matches the intervention size for each prompt and pairs every direction with a random control of the same norm. The candidates are the mean image conditioning shift, a CMRM style refusal direction, a ShiftDC style attack specific residual, a prompt instruction to ignore the image, and a random control. No single candidate dominates on both refusal recovery and utility preservation. The image conditioning shift leads on LLaVA 1.5 and Pixtral 12B and is the only candidate whose utility loss remains at the measurement noise floor in every family. The prompt instruction leads on Qwen2.5 VL, while the attack specific residual leads on Qwen2 VL 2B. The image conditioning direction is direction specific in 13 of 15 cells, but strongly architecture specific and nontransferable across the only dimension compatible pair, LLaVA 1.5 13B and Pixtral 12B. We also connect text only and multimodal refusal geometry. The CMRM direction has positive cosine alignment with the image conditioning shift in all 15 cells, with mean 0.35, range 0.17 to 0.65, 15 to 25 times the random vector null, and a sign test p value of about 3e-5. These results show that the two recipes recover partially overlapping geometry and that direction based defences should be calibrated separately for each language decoder family.
[AI-70] Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
链接: https://arxiv.org/abs/2607.27905
作者: Muhammad Adil Saleem,Syed Ali Raza,Mary-Anne Williams
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values that achieve a contrastive outcome. Reinforcement learning (RL) offers a promising approach for CFE generation, enabling efficient exploration of counterfactual instances while ensuring control over key metrics like validity, sparsity, and proximity. Previous studies have formulated RL states exclusively using features derived from the predictors in the supervised dataset. This study explores the impact of including an instance’s predicted class, alongside features derived from the predictors, in the RL state representation for generating CFEs. The hypothesis is that class-awareness enhances exploration efficiency and improves policy optimality. We compare the proposed class-aware RL method with the class-blind RL method, which is similar but excludes the instance’s class information from the state representation. The comparison was conducted using seven datasets from diverse domains, varying in size. The results show that during training, class-aware RL offers benefits in terms of convergence speed, reward optimization, and episode length reduction. Moreover, it generates significantly more valid CFEs compared to class-blind RL. Finally, the instance’s class-based feature consistently ranks among the most influential predictors in RL’s action-selection, as shown by the SHAP and LIME values, underscoring the significance of class-awareness in RL for CFE generation. The impact is heightened clarity, faster learning, improved validity, and more effective counterfactual generation across diverse datasets.
[AI-71] Dynamic Spectral Filtering for Temporal Graph Learning: Learning Evolving Propagation Operators
链接: https://arxiv.org/abs/2607.27891
作者: Yan Kong
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code is available at: this https URL
Abstract:Temporal graph learning is commonly organized around the evolution of node states or the encoding of interaction histories. We study an underexplored, operator-centric question: should the graph propagation mechanism itself evolve over time? We introduce Dynamic Spectral Filtering (DSF), which represents propagation at snapshot t by a Chebyshev polynomial filter with vector-valued, time-dependent coefficients. DSF explicitly treats these compact multi-order coefficients as recurrent temporal states. A recurrent branch proposes updates, while multiplicative global and order-specific gates regulate their magnitude. The temporal state is independent of the number of nodes. On MOOC, Wikipedia, and Reddit temporal link-prediction benchmarks, converged DSF runs attain AP scores of 0.7851, 0.9088, and 0.9860, respectively, with 93K to 133K trainable parameters, 68 to 182 MB peak GPU memory, and 1.6 to 2.1 seconds of training per epoch. Against the closely related DEFT baseline, DSF is better on MOOC, within 0.001 AP on Reddit, and modestly lower on Wikipedia, while using 8.3 to 8.6 times fewer parameters, 25 to 33 times less GPU memory, and 5 to 19 times less time per epoch. Relative to all measured alternatives, it uses 3.3 to 38.6 times less GPU memory. These results support direct spectral-response evolution as a useful temporal inductive bias when computational efficiency is a first-class requirement.
[AI-72] Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
链接: https://arxiv.org/abs/2607.27888
作者: Qiangqiang He,Zhongheng Wu,ZiJian Wang
类目: Artificial Intelligence (cs.AI)
备注: 20 pages, 6 figures, 11 tables
Abstract:Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.
[AI-73] RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents IROS2026
链接: https://arxiv.org/abs/2607.27881
作者: Sihyung Yoon,Minjong Yoo,Sanghyun Ahn,Seojeong Choi,Honguk Woo
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: Accepted to IROS 2026. 8 pages, 6 figures
Abstract:Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments. Existing solutions address these limitations individually through model retraining or environment-specific modules, yet what is needed is a general framework that systematically transforms a pretrained VLA into a robotic agent. We present RoboBRIDGE, a modular framework that provides an orchestration layer over five coordinated modules, namely Monitor, Perceptor, Planner, Controller, and Robot Interface, to compose robust robotic agents from off-the-shelf components, including pretrained VLAs. The Monitor pairs rapid failure detection with hierarchical recovery to correct errors before they cascade. When the environment diverges from the current plan, the Planner triggers replanning while the Perceptor updates scene understanding asynchronously, avoiding execution stalls. Within the Controller, primitive skill fine-tuning factors manipulation into domain-invariant primitives with dedicated LoRA adapters, reducing sensitivity to domain shifts when a VLA is used. Across LIBERO, RoboCasa, and real-world case studies spanning multiple robot platforms and VLA backbones, RoboBRIDGE consistently outperforms both standalone policies and prior augmented VLA deployments. These results suggest that reliable robotic agency does not arise from scaling action predictors alone, but from structured orchestration around them.
[AI-74] ARES: Adaptive Reasoning -Effort Steering for PPA- and Cost-Aware RTL Optimization with LLM Agents
链接: https://arxiv.org/abs/2607.27879
作者: Stef Cuyckens,Mihaela Jivanescu,Jun Yin,Chao Fang,Marian Verhelst
类目: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
备注: 7 pages, 6 figures
Abstract:Large language model (LLM) agents optimize the power, performance, and area (PPA) of register-transfer-level (RTL) designs by iterating over edits, synthesis, and PPA analysis, paying a dollar cost for every LLM call. Prior agents report the quality reached without its normalized cost, attribute that quality to an engineered cross-design memory, and hold the reasoning effort of every call fixed. We propose Ares with three corresponding innovations. (1) We introduce a normalized dollar cost per LLM call reported alongside the figure of merit (FoM), enabling fair comparison across effort levels and optimizers. (2) Using this accounting, we find the construction of the long-term memory matters little. An engineered memory brings no dependable gain over a plain concatenation of the same experience. (3) We instead adapt the per-call reasoning effort by escalating to deeper reasoning only once progress at a lower effort stalls, via a patience counter fit on 21 training designs, allocating reasoning where it pays rather than uniformly across all iterations. On three test designs unseen during training, the effort policy lowers the FoM by 23-27% where the best fixed effort reaches 16-23%, at equal normalized cost. Ares closes up to 83% of the gap from an LLM-drafted multiply-accumulate unit to its highly hand-optimized counterpart, and reaches a 25% deeper FoM than state-of-the-art Dr. RTL at 12% of its tokens.
[AI-75] An Empirical Study of Coordination Mode as the First-Class Citizen in From-Scratch Multi-Agent Coding
链接: https://arxiv.org/abs/2607.27877
作者: Yanyu Ren,Yunfeng Bai,Xizheng Wang,Li Chen,Dan Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-agent vibe coding promises to accelerate software development, yet existing benchmarks rely on synthetic environments that ignore practical time and monetary costs, conflate reasoning with communication, and reward only superficial completion. We introduce multi-agent from-scratch evaluation benchmark, MSEval, evaluating multi-agent coding on real-world tasks. Grounded in 10 authentic, full-stack projects across 10 domains, MSEval scores performance using hierarchical requirements and deterministic rubrics. Its execution engine, LegoGent, tests 10 collaboration topologies where agents coordinate via periodic sync intervals and deploy through native CI/CD pipelines. Concurrently, the automated grader TAgent dynamically probes implementations to jointly measure functional success, latency, and prefix-cached token cost. Across 100 runs, MSEval reveals that organizational topology rivals model capability in shaping the speed–cost–quality trade-off. For identical tasks and models, varying the topology shifts scores by over 30 points and doubles wall-clock time. Structured pipelines converge fastest with the highest quality, whereas heavy managerial oversight degrades performance. Ultimately, MSEval establishes a rigorous, reproducible standard for measuring how multi-agent teams actually build software. The benchmark is released at this https URL. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.27877 [cs.AI] (or arXiv:2607.27877v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.27877 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-76] Search as Computation Allocation
链接: https://arxiv.org/abs/2607.27871
作者: Alexander Tuisov
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Many algorithms spend an internal resource before returning a decision and are evaluated only by the quality of that terminal output. We formalize such procedures as terminal computation-allocation problems: costly computations produce observations, update beliefs about a latent environment, and matter only through terminal decision loss. Bellman equations characterize optimal allocation under fixed budgets, priced computation, and exact certification. We then relate value of computation (VOC) to information. Mutual information equals myopic VOC under log loss, whereas under simple regret VOC is a knowledge-gradient quantity; moreover, information gain can rank computations arbitrarily poorly, although it gives a one-sided upper bound on VOC. Bandit pulls, tree simulations, and node expansions illustrate the same model under different computation topologies. Finally, under an explicit frontier-resolution and heuristic-error model, maximizing approximate VOC recovers weighted A*, with A* and greedy best-first search as limiting cases. The theory identifies a shared decision problem without asserting that one acquisition rule is universally optimal.
[AI-77] Orca: Neural Operators for Causal Reasoning in Continuous Time
链接: https://arxiv.org/abs/2607.27867
作者: Gerrit Großmann,David A. Selby,Sebastian J. Vollmer
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Structural causal models are the standard language for reasoning about interventions and counterfactuals, but they describe static variables, typically measured once, and usually forbid cyclic dependencies. Many systems we care about, such as patients, climates, and economies, instead evolve continuously in time, are observed at irregular time points, and contain feedback loops. We argue that neural operator learning provides a natural foundation for causal reasoning in this setting, and propose Orca, a framework in which each node of the causal graph is a function of time and each mechanism is a learned map between function spaces. We extend existing neural operator architectures to express causal mechanisms: a mechanism computes the function value of a node from its parent nodes by taking several parent functions as input, respects the arrow of time, and treats latent exogenous noise as a function that can be inferred and reused for counterfactuals. We formalize the model class and demonstrate counterfactual reasoning on synthetic continuous-time examples. Code is available at this https URL
[AI-78] Back to All-Entity Ranking: Sampler-Dependent Evaluation in Continuous-Time Dynamic Graphs
链接: https://arxiv.org/abs/2607.27861
作者: Minwoo Yu,Young-guk Ha
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Next-destination prediction in continuous-time dynamic graphs (CTDGs) commonly ranks an observed interaction against sampled negative destinations. The resulting score is conditional on both the negative distribution and the number of candidates chosen by the researcher. We show that a non-uniform negative distribution changes the Bayes-optimal ranking, while even a finite candidate set drawn uniformly can destabilize model rankings and measured module effects. Time-varying source-destination history membership and model operations that use this information directly transmit the sampler’s influence to the evaluation score. We examine this mechanism using a factorial evaluation of repeated and new positives against seen and unseen negatives, a minimal scorer based solely on pair-history membership, and controlled representation interventions. Across six models on LastFM, MOOC, Reddit, and Wikipedia, at least one model pair changes relative order between the expected Uniform-20 metric and the full catalog on three of the four datasets. The measured effect of the same module also changes in magnitude and direction with the candidate-set size and training objective. These results establish that model-superiority and ablation conclusions from sampled-negative benchmarks are conditional on the stated candidate configuration. All-entity ranking evaluates every destination in a fixed catalog, eliminating negative-selection freedom and sampling variation while retaining the original CTDG scorer. We therefore recommend all-entity ranking as the primary evidence for architecture comparisons on CTDG benchmarks with an enumerable, fixed destination catalog. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.27861 [cs.AI] (or arXiv:2607.27861v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.27861 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-79] Virtual Process Dossier: A Process-Aware Data Catalogue
链接: https://arxiv.org/abs/2607.27840
作者: Lukas Kubelka,Alexander Bott,Frank Döhner,Saksham Kiroriwal,Georg Zeeb,Julia Butte,Julius Pfrommer,Jürgen Beyerer,Tobias Käfer
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:We propose the Virtual Process Dossier (VPD), a Knowledge Graph-based data catalogue that also captures workflow provenance. We developed VPD for multi-stage manufacturing use-cases where downstream AI-based optimization tasks require to distinct between datasets generated during individual workflow steps. VPD provides these datasets in a FAIR manner and makes both prospective and retrospective workflow provenance explicit. Our contributions are: (1) the VPD ontology that serves as the catalogue’s semantic core; (2) the VPD provenance framework that integrates ontology instantiation into the production environment; and (3) the VPD user interface that provides human-centered interaction with the VPD Knowledge Graph. The ontology and code are available at this https URL .
[AI-80] Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration
链接: https://arxiv.org/abs/2607.27836
作者: Xiangyu Yin,Jiaxu Liu,Zhen Chen,Chih-Hong Cheng
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recovers held-out forget-set ROUGE for every method we evaluate, and we trace this fragility to optimization geometry. The per-token answer margin of fourteen post-hoc methods spanning gradient, preference, and distillation families converges into a narrow band above the retain reference in 41 of 42 method–size cells, a regularity we call the margin cliff. We prove that this cliff follows whenever the retain coupling holds the diagnostic log-odds of forget content above a floor, a condition that token-saturating losses induce at stationarity and that we verify directly on 34 of 42 cells. Margin Calibration (\textscMC) is a plug-in polish adding a non-saturating margin hinge anchored at the reference’s per-token margin plus a KL probe on a disjoint instruction corpus, restoring forget-side pressure where the native loss saturates. Under a stated gradient-dominance condition, whose on-trajectory gradient signature we measure by instrumenting the polish, its stationary set lies on the cliff-crossing side, yielding an attack-budget upper bound on the relearn margin lift. Across TOFU (three Llama-3 sizes, three forget tiers), MUSE-News on Llama-2-7B-hf, and a Phi-3.5 panel, a single frozen configuration wins all 14 head-to-head forget aggregates and all populated relearn cells (panel-mean post-attack ROUGE-L 0.41 to 0.18 ) and lowers raw membership AUC on 13/14, with reduced retain-side utility as the main cost. A deployment variant matches these gains without a retain-trained reference.
[AI-81] STEREODISCO: Discovering Stereotypicality in LLM s
链接: https://arxiv.org/abs/2607.27824
作者: Farane Jalali Farahani,Corina Dima,Mojtaba Nayyeri,Raphael H. Heiberger,Steffen Staab
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM’s activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work – including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.
[AI-82] MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes
链接: https://arxiv.org/abs/2607.27798
作者: Weihang Wang,Kainan Tu,Jielei Zhang,Run Yang,Boheng Sheng,Yuchen He,Yu Xie,Pengyu Chen,Peiyi Li,Huyang Sun,Longwen Gao,Zhouhui Lian
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 5 figures, and 13 tables
Abstract:Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.
[AI-83] Annotating Topical Legal Insights from Case Proceedings
链接: https://arxiv.org/abs/2607.27792
作者: Subinay Adhikary,Dwaipayan Roy,Debasis Ganguly,Shouvik Kumar Guha,Kripabandhu Ghosh
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:In this paper, we mainly concentrate on finding concepts or topics from the legal case proceedings, since adopting a structured representation for legal documents, as opposed to a mere bag-of-words flat text representation, can significantly enhance processing capabilities. To achieve this objective, we put forward a set of diverse concepts for legal case proceedings. With this motivation, we propose LeDA, a system for Legal Data Annotation. The system offers the generic functionality of annotating and adjudicating entities or concepts within documents via a web-based interface. A novel feature of our system is that it allows to dynamic create new tags for annotation, which is a particularly useful provision for situations where there exists no pre-defined ontology for the entities (concepts) that need to be annotated - these being rather discovered by annotators as they continue examining more documents. The system that we demonstrate is currently in use to annotate a set of concepts from legal documents to construct semantic representations of documents as bags of concepts that can then be used for several downstream tasks, such as prior case retrieval, judgment prediction, and so on. Along with the system features in general, we also describe how LeDA was used by 3 assessors to annotate and adjudicate legal concept names from Indian Supreme Court case proceedings.
[AI-84] SpecCal: Ambiguity-Aware Candidate Calibration for Infrared Spectrum-Based Molecular Structure Reconstruction
链接: https://arxiv.org/abs/2607.27788
作者: Yixuan Chen,Bo Liu,Yusen Tan,Guokun Yang,Wenjie Du,Jun Xia
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:
Abstract:Inferring molecular structures from infrared (IR) spectra is a fundamental yet challenging problem. A key difficulty is that an IR spectrum provides limited structural information: different molecules may share similar functional groups and local vibrational patterns, leading to highly similar spectral responses. Thus, even when an observed spectrum has a unique underlying structure, reconstructing it from the spectrum remains ambiguous. Existing IR-to-molecule models usually generate a ranked set of candidate molecules, but this set is largely determined by the model’s learned generation preference and may not fully capture the structures that best satisfy the observed spectral constraints. To address this limitation, we propose SpecCal, a training-free candidate calibration framework for IR-to-molecule prediction. SpecCal operates on the candidate outputs of existing base models and improves the prediction set by re-ranking current candidates while introducing additional structurally plausible alternatives guided by spectral consistency. The framework is plug-and-play and model-agnostic, requiring no parameter updates for integration with diverse base models. Experiments on multiple benchmarks show that SpecCal consistently improves top-k reconstruction at both SMILES and scaffold levels across different base models. Further analyses demonstrate that calibrating candidate sets under spectral ambiguity provides a practical way to improve molecular reconstruction from IR spectra. The code is available at: this https URL.
[AI-85] LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts
链接: https://arxiv.org/abs/2607.27787
作者: Ken Ding
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on “cliff” prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is identically zero, so GRPO produces no gradient on precisely the prompts at the frontier of the model’s capability. We introduce LoRA Scaffolded Policy Optimization (LSPO), a sampling-time mechanism that recovers this lost gradient. Each RL step, LSPO detects cliff prompts, fits a small low-rank (LoRA) adapter by a brief supervised step on their ground-truth solutions, re-rolls the cliffs with the base-plus-adapter model, splices the now-successful completions back into the RL batch with an importance-sampling correction, and takes a GRPO step on the base alone; the adapter receives only the supervised gradient and is discarded at checkpoint, yielding a base-only model. On DeepMath-103K with DeepSeek-R1-Distill-Qwen-1.5B, evaluated over n=5 paired seeds per arm at a matched 1000-step reporting horizon, LSPO’s 5-seed mean matches or beats a DAPO baseline on all 16 (benchmark, pass@k) cells (15 strict wins and one exact tie), with gains of up to +10.7 points on AIME24/pass@4, +6.7 points on AIME24 and AIME26 at pass@16, and +2.4 points on MATH500/pass@1; averaged over the 16 cells the improvement is +3.8 points.
[AI-86] RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
链接: https://arxiv.org/abs/2607.27782
作者: Zhengyang Yan,Junhao Li,Fangqi Zhu,Zijun Wang,Quanxin Shou,Yikun Miao,Xiaoyi Pang,Zicong Hong,Song Guo
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts during deployment. While offline reinforcement learning (RL) provides a practical way to improve deployed policies using rollout data, existing methods either ignore failure data or exploit it only at the trajectory level, resulting in low learning efficiency and persistent errors. We propose RedFlow, a fine-grained offline RL framework that redirects failure experiences into action-level corrective supervision for flow-matching VLA policies. RedFlow consists of two key components: (1) a Context-Aware Corrective Matching mechanism that identifies failure-inducing actions and retrieves successful alternatives from similar contexts as corrective targets, and (2) an Adaptive Redirection Objective that jointly reinforces successful actions, suppresses undesirable ones, and redirects recoverable failures toward corrective targets. By converting both successful and failed experiences into dense supervision, RedFlow enables robust recovery learning from mixed-quality data. Experiments on the LIBERO benchmark and three real-world manipulation tasks show that RedFlow consistently outperforms state-of-the-art offline RL baselines, improving the real-world success rate from 56.7% to 74.7%. It also matches strong on-policy methods (PPO, GRPO, and DDPO) while requiring roughly an order of magnitude fewer training samples.
[AI-87] VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
链接: https://arxiv.org/abs/2607.27768
作者: Yukun Chen,Tianrui Wang,Zhaoxi Mu,Xinyu Yang,EngSiong Chng
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:
Abstract:Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric–note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by 0.42 points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.
[AI-88] rain Small Deploy Large: Zero-Shot GNN Transfer Through Geometric Renormalization
链接: https://arxiv.org/abs/2607.27767
作者: Robert Jankowski,Pedro Almagro-Blanco,Marián Boguñá,Melanie Weber,M. Ángeles Serrano
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Physics and Society (physics.soc-ph)
备注:
Abstract:Graph neural networks (GNNs) can operate on large graphs but become infrastructure-sensitive at the scale of millions of nodes and typically require scalable training techniques for even larger graphs. This raises a central question: when can a model trained on a smaller, scaled-down replica of a graph be deployed on the full-resolution graph without retraining? We introduce a zero-shot transfer protocol in which a GNN is trained on a graph coarse-grained by geometric renormalization (GR), and the resulting weights are transferred directly to the original network. Across synthetic and real-world networks, training on GR scaled-down replicas preserves much of the original-scale predictive performance while significantly reducing training cost. We further find that learned representations and predictive trajectories remain aligned across scales. These findings suggest that structural similarity may be more important than network size in determining GNN transferability, opening a path toward scale-equivariant graph architectures.
[AI-89] VeriSkill: A Self-Evolution Framework for Program Verification Skills
链接: https://arxiv.org/abs/2607.27733
作者: Changguo Jia,Tianqi Zhao,Zhiyou Xiao,Weiming Zhang,Minghui Zhou
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Automating program verification with LLM agents requires generating specifications, annotations, auxiliary lemmas, and tool invocations, all of which depend on reusable skills. A natural remedy is skill self-evolution: distilling skills from trajectories and refining them through feedback. However, existing evolution methods struggle with program verification tasks because they cannot reliably identify skill-specific failures or extract actionable signals from opaque verifier feedback. In this paper, we propose VeriSkill, a self-evolution framework built for program verification. It attributes verification failures to skill deficiencies, distills diagnostic signatures into reusable lessons, and iteratively refines candidate skills, admitting only revisions that improve verification performance while preserving program semantics. Experiments show that VeriSkill consistently outperforms all baselines across multiple verification tools, agent frameworks, and LLM backends.
[AI-90] owards joint scaling laws with optimal batch size schedules
链接: https://arxiv.org/abs/2607.27731
作者: Jiaxiang Li,Zhiqi Bu,Shiyun Xu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
备注:
Abstract:Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.
[AI-91] New Synchronous Computation Dynamics for Hopfield Networks
链接: https://arxiv.org/abs/2607.27720
作者: Francisco Requena-Domínguez,Rafaela Benítez-Rochel,Ezequiel López-Rubio
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The dynamics of the original Hopfield network is asynchronous (sequential) (updates the state of only one neuron per time step). In this paper, we propose a new tool and a new dynamics to reduce the processing time by updating one or more neurons simultaneously per instant while ensuring process convergence and aiming for the maximum energy decrease at each step, thus guaranteeing the shortest total processing time. From the point of view of synchronous dynamics, calculating the next network state at which energy decreases the most from the current state while ensuring convergence is itself a combinatorial optimization problem. We develop and use a new tool to solve it. We call this new tool Discrete Differential Filter (DDF) and, based upon it, we develop a new synchronous dynamics which we call SD-DDF (Synchronous Dynamics based upon Discrete Differential Filter). In this paper, we review the original asynchronous dynamics for Hopfield networks and present a new tool and a new synchronous dynamics with its theoretical justification and four computational experiments to assess the speed up in processing time empirically.
[AI-92] MECA: A Mechanism-Centered Agent for Constructing Well-Specified and Valuable Mathematical Conjectures
链接: https://arxiv.org/abs/2607.27709
作者: Wentao Long,Yunfei Zhang,Chenyi Li,Zaiwen Wen
类目: Artificial Intelligence (cs.AI)
备注: 78 pages, 5 figures. Includes appendices and the full collection of 100 generated conjectures
Abstract:Automatically constructing well-specified and valuable mathematical conjectures remains a central challenge in AI-assisted mathematical discovery. Many existing open problems and conjectures are often too broad, underspecified, or difficult to connect to plausible proof or refutation strategies. We view a mathematical mechanism as a structure or reasoning principle that connects the assumptions of a candidate problem to its target conclusion, such as an inequality, invariant, decomposition, or reduction to an intermediate claim. We present MECA (MEchanism-centered Conjecture Agent), a multi-agent framework that constructs conjectures by jointly developing candidate statements and their supporting mechanisms. Explorer agents propose mechanisms, test how they apply, and revise the candidate conjecture accordingly, while critic agents assess their mathematical validity and research value. Their feedback guides changes to the assumptions, scope, and conclusion. Through this process, MECA transforms broad research directions into precise conjectures with substantive mathematical support while retaining a clearly identified unresolved core. We evaluate MECA in two complementary settings. First, we compare it with a generate-and-revise baseline on reconstructing preselected target-paper conclusions from target-conditioned but article-blind source materials. Second, we construct 100 semi-open problems from literature-derived seeds and existing open problems and evaluate them through independent proof and refutation attempts by automated provers. Our results indicate that mechanism-centered refinement produces well-specified and research-worthy conjectures that remain challenging for current automated provers.
[AI-93] Albilich: Steerable Proof-State Orchestration for LLM -Based Mathematical Research with CAS Integration
链接: https://arxiv.org/abs/2607.27705
作者: Ting Gong,Michael Ruofan Zeng,Yong Yang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 7 pages, comments welcome!
Abstract:Large language models can contribute useful ideas to mathematical research, yet long-horizon proof attempts remain difficult to coordinate, evaluate, and reproduce. We present Albilich, an open-source agentic harness for autoresearch in mathematics that combines long-horizon reasoning, computer algebra systems (CAS), literature retrieval, and persistent SQLite-based context management. We evaluate Albilich on the RealMath benchmark (Zhang et al. 2025) and on open problems in group theory from the Kourovka Notebook (Khukhro and Mazurov 2026). It solved 10/10 problems on RealMath with CAS and 9/10 with no CAS. On the Kourovka problems, Albilich produced a counterexample to Problem 21.142 and a proof of a strengthening of Problem20.2. Anablation on Problem 17.91 demonstrates 32.0% token reduction when CAS is enabled. An ablation on Problem 21.142 demonstrates higher verifier-rejection rate and failure to synthesize proof routes in the absence of the advisor agent. These results support Albilich as a human-steerable, CAS-boosted environment for scalable AI-assisted mathematical research. Comments: 7 pages, comments welcome! Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2607.27705 [cs.AI] (or arXiv:2607.27705v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.27705 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-94] SpatialCLI: Learning to Reason With Spatial Tools Then Without Them
链接: https://arxiv.org/abs/2607.27703
作者: Yang Zhou,Zixuan Huang,Sunzhu Li,Zhuo Yang,Chen Zhang,Shunian Chen,Caijun Yan,Jianyao Xu,Shunyu Liu,Weijie Fu,Peiliang Li,Xiaozhi Chen,Yuxiang Cai
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often miss the visual details that determine success, while specialist vision models can capture those details but cannot translate them into task-level decisions. In this work, we propose SpatialCLI, a framework that teaches VLMs to reason with spatial tools and progressively internalize the specialist perceptual capabilities they provide. SpatialCLI proceeds in three stages: (1) Call exposes specialist vision models as spatial tools to augment the VLM’s perception; (2) Learn uses Cold-Start SFT and agentic RL to improve tool use; and (3) Internalize verbalizes successful tool-use trajectories to internalize specialist perceptual capabilities. We further introduce SpatialCLI-Bench, a 516-example benchmark for compositional perception across localization, segmentation, depth, and pose. On MindCube, SpatialCLI raises Qwen3-VL-8B-Instruct from 29.3% to 84.6% with tools, surpassing GPT-5.6 Sol with tools (72.1%), while retaining 73.8% without tools after internalization.
[AI-95] Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling
链接: https://arxiv.org/abs/2607.27698
作者: Yuan Tian,Yi Mei,Mengjie Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:In dynamic multi-mode project scheduling, activities have alternative execution modes and uncertain durations, while precedence relations and limited resources constrain their execution. Heuristic priority rules support fast online decisions, but their design requires substantial domain expertise. Genetic programming (GP) hyper-heuristics can automatically evolve such rules. Large language models (LLMs), meanwhile, provide a flexible interface for interpreting scheduling information and explaining decisions. However, zero-shot LLM decisions may lack domain knowledge, consume many tokens, and vary across repeated queries. GP-evolved rules therefore provide a potential source of scheduling knowledge for guiding LLM decisions. Unlike existing LLM–GP hybrids that use LLMs to support heuristic evolution, we transfer knowledge in the reverse direction, using knowledge extracted from high-quality GP rules to guide an online LLM decision maker. We extract knowledge from high-quality GP rules and inject it through Feature Selection, Feature Hint, Rule Reference, and Rule Follow. These mechanisms are evaluated in terms of scheduling performance, token consumption, decision stability, and the feature focus expressed in generated rationales. GP-derived guidance generally improves the unguided LLM, but its representation matters. Simplifying the decision context or supplying explicit decision logic is more effective than highlighting important features. Feature Selection offers the best token efficiency, whereas Rule Follow achieves strong performance at greater token cost. Guidance also improves decision stability and changes the features expressed in generated rationales.
[AI-96] LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
链接: https://arxiv.org/abs/2607.27690
作者: Jingya Wang,Yuyang Gao,Liuzhenghao Lv,Yonghong Tian,Yuyang Liu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experience. LabEvolver couples a state-grounded inner trial loop for adaptive perception, online planning, and safety validation with an outer evolution loop that distills completed trajectories into reusable skill, strategy, and safety experience. On robotic solution-preparation tasks, LabEvolver demonstrates real-world feasibility, reducing pH-regulation completion time and safety-gate intercepts by 48.2% and 60.0%, respectively. On ALFWorld, it further improves cumulative success rate within 20 steps from 76.2% with ReAct to 91.4% over 500 continual tasks, showing generality beyond wet-lab settings. These results support learn-by-doing experience evolution as a feasible path toward closed-loop automated scientific discovery. The project page is available at this https URL.
[AI-97] Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
链接: https://arxiv.org/abs/2607.27687
作者: Jiazhen Ji,Shouhong Ding
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Autoresearch improves machine-learning code by proposing changes, running full training jobs, and keeping changes that improve the metric. The efficiency of this loop depends not only on generating ideas, but also on the agent’s ability to decide, before spending a training run, whether a proposed modification is likely to work. We study how the reliability of this pre-execution judgment changes over the course of an autoresearch trajectory. In public AutoSOTA logs (Li et al., 2026; Tsinghua FIB Lab, 2026), the fraction of helpful modifications falls from 70% in the first two iterations to 43% by iteration 6+. On 296 same-baseline modification pairs from 39 paper-derived AutoSOTA tasks, each containing one modification that improved the metric and one that did not, with measured outcomes hidden, an LLM judge given candidate rationales but no prior-attempt history reaches 79.5% accuracy on the pairs where strict consensus returns a verdict. On the full 366-pair benchmark, however, this ability weakens substantially late in the loop. As successful changes accumulate, selective accuracy - accuracy conditioned on a strict-consensus verdict - falls from 82.8% to 56.9%, while the judge remains willing to decide. We call this operational pattern the confidence cliff. Rehearse implements the loop change as a lightweight skill for autoresearch loops: propose several ideas, compare them before execution, run the most promising, and judge with a focused memory of similar past attempts and outcomes. This focused outcome memory raises late selective accuracy to 83.5%. Across 4,000 budgeted training runs over three loops, Rehearse improves the endpoint under the same training-run budget on nanochat, image classification, and time-series forecasting.
[AI-98] Evaluating and Pricing Advertisements in AI-Generated Responses
链接: https://arxiv.org/abs/2607.27686
作者: John L. Turner-Smith,Zimeng Huang,Yuhan Fu,Yihang Zhang,Tonghan Wang
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 2 figures
Abstract:As search increasingly shifts toward LLM-driven answer engines, advertising is becoming embedded within the generated response itself and should therefore be evaluated for both user utility and commercial value. The key challenge is click-through intent: behavioural logs are unavailable, human annotation resists calibration, and frontier LLM judges conflate intent with linguistic fluency. These gaps compound, as principled pricing presupposes a continuous intent signal, while generating such a signal presupposes supervision that is currently unavailable. We construct the missing supervision through a psychologically grounded agent simulation framework, and distil it into a parameter-efficient evaluator that predicts click-through intent, together with the three companion dimensions of ad quality, as smooth, differentiable estimates. Validated through sign-certain behavioural perturbations, the evaluator surpasses frontier zero-shot judges on relevance sensitivity (79% versus 60-67%), tracks graded content degradation, generalises without error to 103 fictional products, and agrees with human preference in 86% of pairwise judgements across five annotators, with agreement rising in the evaluator’s confidence. Upon its estimates we build the pricing layer directly, deriving the unique payment rule under which truthful bidding is optimal, demonstrating it on a best-of-k allocation, and extending the mechanism to non-monotone allocations. The same differentiable signal stands ready as a training objective for ad generation.
[AI-99] HALO: Heterogeneous Admission through Localized Obligations for Safe Agent ic Execution
链接: https://arxiv.org/abs/2607.27636
作者: Taewoo Park,Kyeonghyun Yoo,Kiseok Kim,Seunghyun Yoo,Hwangnam Kim
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO); Software Engineering (cs.SE)
备注: 16 pages, 2 figures; supplementary material included
Abstract:Recent agentic AI systems may return a heterogeneous response containing notices, requests, handoffs, and actions. Conditions can change before external use, so components from the same response need not remain supported together. Rejecting the whole response discards useful components, whereas checking components independently can leave a dependent without its prerequisite. We present Heterogeneous Admission with Localized Obligations (HALO), a runtime protocol that preserves supported components whose declared prerequisites also remain supported, rechecks each exact action before dispatch, and allows blocked actions to be replaced only by fresh candidates. HALO matched all 96 admission expectations and passed all 20 protocol tests. In structured-response replay, it retained 248/248 supported components, including 128/128 unaffected by unrelated changes, while a whole-response policy retained 0/248. Across ten cold-start PX4/Gazebo sessions, HALO blocked every tested stale route, observed no matching stale setpoint, and completed all fresh recoveries.
[AI-100] HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data
链接: https://arxiv.org/abs/2607.27635
作者: Xiaotong Yu,Joshua Y. Kim,HaeJin Lee,Kalina Yacef
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Wearable sensors continuously capture fine-grained multivariate time-series data, providing opportunities to model behavioural patterns associated with health outcomes. However, existing deep learning methods prioritise predictive accuracy over interpretability, limiting their application in health research. In this study, we present HealthCAT, a flexible framework that integrates an Encoder-only Transformer with an Attentive Class Activation Token (AttentiveCAT) to generate class-specific, time-step-level interpretations. These interpretations can be mapped back onto behavioural cycles that are relevant to the domain (e.g., time-of-day), supporting individual-level analysis of wearable sensor data. We evaluated HealthCAT using two real-world wearable sensor datasets (306 participants in total). HealthCAT outperformed deep learning baselines by up to 17% in F1-score and 12% in accuracy on both datasets ( p0.05 ). In masking experiments, the time steps identified by HealthCAT carried significantly more predictive value than random selection across all masking conditions ( p0.05 ), indicating that the identified time steps are predictively informative. By coupling predictive performance with validated time-step-level interpretability, HealthCAT moves wearable sensor analysis beyond aggregated metrics towards temporal patterns that support health monitoring, behavioural pattern analysis, and intervention design in health research. The significance of this work is that it enables accurate prediction of health indicators from wearable sensor data while providing insights into when and how physical activity patterns occur, rather than relying solely on aggregated summary measures.
[AI-101] SCOPE: Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization
链接: https://arxiv.org/abs/2607.27630
作者: Nguyen Viet Tuan Kiet,Nguyen Huu Duc,Le Cong Bang,Tran Cong Dao,Huynh Thi Thanh Binh
类目: Artificial Intelligence (cs.AI)
备注: 43 pages; Kiet, Duc, and Bang contributed equally
Abstract:Black-box combinatorial optimization requires systematically identifying high-quality solutions under a limited evaluation budget, yet the unknown objective function provides little guidance for deciding where the search should explore next. We introduce SCOPE, a general framework for Synthetic Conditional Objectives for Policy Evolution in Black-Box Combinatorial Optimization. Rather than directly optimizing the inaccessible objective, SCOPE learns a set of synthetic objectives conditioned on the accumulated search history, where each objective is designed to expose a distinct and potentially useful preference over candidate solutions. These objectives are then used to evolve search policies that generate diverse candidates, whose true quality is subsequently assessed through black-box evaluations. The outer loop adaptively updates and selects synthetic objectives according to how effectively their induced policies discover promising regions. In contrast, the inner loop returns a portfolio of top-performing policies to reduce the risk of relying on a single surrogate preference. This formulation reframes objective design as a mechanism for guiding policy exploration, enabling the search process to exploit observed evidence while maintaining structured diversity across discrete solution spaces. Extensive experiments across multiple benchmark problems demonstrate that SCOPE consistently improves black-box search performance under limited evaluation budgets and generalizes well across diverse combinatorial structures.
[AI-102] Arm2Air: Cross-Embodiment Skeleton Transfer for 3D Relay Formation
链接: https://arxiv.org/abs/2607.27627
作者: Dohun Lee,Kyeonghyun Yoo,Seokmin Kim,Byongho Lee,Seungjoo Oh,Hwangnam Kim
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 9 pages, 4 figures
Abstract:Unmanned aerial vehicle (UAV) relay networks can restore connectivity after communication infrastructure is damaged. Urban relay placement is difficult because line-of-sight blockage, communication range, altitude, and three-dimensional obstacles must be considered jointly. Arm2Air transfers obstacle-avoidance skeletons from robot arms to UAV relay placement through cross-embodiment transfer. Source-domain robot-arm motions from a pretrained Neural MP model are converted into ordered skeletons that pretrain a transformer-based transfer platform, which is then adapted to the UAV domain using limited target data and Low-Rank Adaptation. The transferred skeleton initializes a relay chain that is refined for connectivity, bottleneck capacity, delay, and movement cost. On nine held-out high-clutter 3D urban maps, Arm2Air reduced median end-to-end planning runtime by 64.9 percent relative to the fastest conventional planner. On the high-obstruction group of a separate 30-map dense urban holdout, it increased bottleneck capacity by 32.6 percent, reduced capacity variance by 74.7 percent, reduced maximum hop distance by 13.2 percent, reduced hop-distance variance by 75.2 percent, and reduced relay displacement by 16.9 percent relative to IMPC-MD. With only three target-domain training maps, Arm2Air reduced relay-position root mean square error by 53.6 percent relative to training from scratch while updating 0.134 million parameters, compared with 1.383 million for Scratch and Full Fine-tuning. These results demonstrate computationally and data-efficient UAV relay placement and suggest a broader principle for transferring ordered structural priors across heterogeneous embodied tasks.
[AI-103] Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
链接: https://arxiv.org/abs/2607.27617
作者: SiYuan Ma,Yiqin Luo,Zhangji,Canran Xiao,Albert Gao,Wei-Hsing Huang,Wei Wang,Qiwei Wu,Xinran Li,Jinfeng Wei,Qixin Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Identical language-model answers can arise from hidden states that support different future computations, so current-answer probes do not establish a reusable internal interface. We introduce forked futures: future operations are sampled only after a prefix state has formed, and states are compared through the response distributions induced by those operations. This yields an empirical causal quotient over hidden states without requiring researcher-specified latent labels. Shared, Local, Mixture, and Distributed interfaces then compete under prequential causal description length subject to future-signature fidelity and matched capacity constraints. In the two detailed model evaluations, Shared has the lowest held-out description length, with gains of 0.216 nats on Qwen2.5-1.5B and 0.294 nats on Llama-3-8B, while maintaining tightly clustered mean future-signature distortion; a five-backbone sweep preserves the positive direction of Sharedness Gain. The figure-aligned transplantation analysis gives Shared the strongest joint target-correctness, locality, copy-preservation, and composite profile, and API-aligned paths mediate 0.749 of the target effect versus 0.150 for matched null paths. In the blind four-class model-organism test, 14/16 architectures are recovered, with one observed non-Shared to Shared error among 12 non-Shared organisms. These results support an economical reusable causal interface within the tested operation banks, while keeping the claim explicitly conditional on the candidate architectures, interventions, and held-out futures.
[AI-104] CORE: In-Context Reconstruction for Unified Tabular Anomaly Detection
链接: https://arxiv.org/abs/2607.27615
作者: Yunfeng Zhao,Qingfeng Chen,Yue Tan,Shiyuan Li,Yili Wang,Yixin Liu,Shirui Pan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Tabular anomaly detection (TAD), which focuses on identifying abnormal samples that deviate from the majority in tabular data, has received growing attention. Recently, there has been an emerging trend towards unified TAD, which seeks to detect anomalies across different datasets using a single generalizable model. In unified TAD, aligning heterogeneous data remains challenging. While existing methods often rely on distance-based unified feature construction, they may obscure the semantics of the original features. Moreover, existing approaches typically formulate anomaly detection as a binary classification task, which may overlook diverse anomaly patterns from various datasets and be misled by unrepresentative synthetic anomalies. To address these challenges, we propose an in-COntext REconstruction approach for unified TAD (CORE for short). It introduces a decorrelated feature alignment module to directly align heterogeneous features into a unified representation space, which retains their semantic information. Meanwhile, CORE formulates unified TAD as an in-context reconstruction problem, eliminating the need for labeled or synthesized anomalies. Specifically, the in-context reconstruction module reconstructs each sample by leveraging contextual normal samples to capture dataset-specific distributions, such that reconstruction errors reflect its deviation from normality, facilitating unified TAD on arbitrary unseen datasets.
[AI-105] Revisiting the Adversarial Robustness of Graph-Based Traffic Forecasting
链接: https://arxiv.org/abs/2607.27604
作者: Qingzhao Zhang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Traffic forecasting by graph-based AI is a critical component of intelligent transportation systems, motivating security research on robustness to malicious sensor readings. We argue that prior robustness evaluations are largely shaped by unrealistic threat models and untargeted objectives, so both attacks and defenses must be revisited. We study a practical adversary with limited model knowledge and the ability to monitor and manipulate only a few road sensors. More importantly, practical attacks can be localized to specific links or routes, causing incorrect estimated arrival times or unnecessary rerouting while leaving the broader network largely unaffected. This targeted setting remains underexplored, and defenses such as adversarial training do not transfer well from the norm-bounded attacks they train on to structurally different, physics-aware attacks that mimic genuine congestion. We therefore reframe robustness as a detection problem, introducing a learned physics-informed detector whose output is fed to a hardened forecaster as an input feature and trained against adaptive attacks with the forecaster fixed. We evaluate across a variety of model architectures and benchmarks. The physics-aware attack multiplies target-link error several-fold while the network-wide error barely moves, and adversarial training, tuned to norm-bounded perturbations, barely dents it. Our detection–mitigation defense improves even on adversarial training hardened against the physics-aware attack itself, on 13 of 15 model–dataset settings and by the widest margin on a held-out attack, at near-zero clean cost. The results emphasize the need to examine abstracted AI adversarial attacks under application-specific constraints to assess their true security impacts.
[AI-106] World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
链接: https://arxiv.org/abs/2607.27599
作者: Xiangcheng Zhang,Yilun Du
类目: Artificial Intelligence (cs.AI); Robotics (cs.RO)
备注: Project page at this http URL
Abstract:Building generalizable agents for diverse applications remains a fundamental challenge. While imitation learning-based policies succeed in specific training environments, they often fail to generalize to novel scenes and tasks. In this work, we propose World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model. Our system enables an agent to propose initial action plans and iteratively refine them via optimization and search, reasoning over imagined world model rollouts. We demonstrate that our approach achieves superior performance across compositional tasks, new layouts, and zero-shot generalization scenarios, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs. Project website at this http URL
[AI-107] Wiring diagram extraction and gluing: a case study in classifying figure skating jumps using 3D dataset
链接: https://arxiv.org/abs/2607.27598
作者: Jason Lo,Mohammadnima Jafari
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 14 pages
Abstract:Hasse clustering is an algorithm that extracts common patterns in sequential data and represents them in graphical forms. As the number of expected clusters grows, however, the algorithm can become infeasible to run due to combinatorial complexity. In this article, we describe a theory of gluing wiring diagrams, allowing iterative applications of Hasse clustering to achieve the same result as a single application. We test our theory in the context of classifying videos of figure skating jumps.
[AI-108] A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response
链接: https://arxiv.org/abs/2607.27597
作者: Swapnil Saha,Bhuvan Rajanasiriyur Jagadeesha,Karishma Patnaik,Neelakshi Majumdar
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
备注: 10 pages, 8 figures. Author accepted manuscript of AIAA Paper 2026-4010, published in the AIAA AVIATION 2026 Forum
Abstract:Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of sensor data under time pressure. Current VLM applications include social media monitoring for situational awareness, generation of draft action plans, and translation of technical alerts into public-facing messages. While these efforts can accelerate information flow, they remain largely limited to decision-support roles. Such approaches can increase operator burden because humans must still translate outputs into coordinated actions across teams and robotic assets. This study explores the viability of embedding VLMs as coordination agents within the human-UAV loop. The proposed architecture integrates natural language interaction, mission-level task coordination, software-in-the-loop implementation, and communication aligned with the Incident Command System (ICS). Rather than functioning solely as advisory tools, VLMs facilitate communication between human operators, mission control logic, and UAV task execution. The framework was developed using a Model-Based Systems Engineering (MBSE) approach, with use case and block definition diagrams representing system roles, internal structure, and component interactions. Three key elements, the VLM Coordinator Agent, UAV Mission Control, and Task Allocator, were implemented within an integrated simulation and control environment. A preliminary human-factors evaluation with seven participants showed reduced perceived workload across mental demand, effort, and frustration, along with high ratings for AI trust and communication clarity. By integrating MBSE, software-in-the-loop testing, and human-factors evaluation, this work advances scalable human-autonomy teaming for high-stakes disaster response, with broader implications for aerospace autonomy and civil safety.
[AI-109] Is Solving Better Than Evaluating GenAI Solutions?
链接: https://arxiv.org/abs/2607.27586
作者: Ethan Dickey,Marios Mertzanidis,Alexandros Psomas
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注: 11 pages
Abstract:As Generative AI (GenAI) tools become increasingly capable of generating solutions to computing assignments, the computing education community is exploring pedagogical approaches that emphasize solution evaluation, verification, and critique alongside traditional solution generation. However, evidence regarding the impact of such evaluation-centered tasks on student learning remains limited, particularly in upper-division, theory-heavy courses. We conducted a randomized A/B crossover study (N=220) in a junior-level algorithms course to compare evaluating GenAI-generated solutions with traditional problem solving. Across six assignments, student working groups either solved challenging algorithmic problems directly or evaluated often-flawed GenAI-generated solutions, with roles reversed midway through the semester. We found no statistically significant differences between groups in midterm scores, final exam scores, overall course grades, or exam problems structurally aligned with the homework interventions. Students received significantly higher homework scores when evaluating GenAI-generated solutions, but this localized advantage did not translate into downstream summative gains. Survey data further indicated that most students reported no change in study habits in response to the intervention; however, those who reported adapting their study strategies rated the GenAI-evaluation assignments as significantly more helpful. These findings suggest that GenAI evaluation redistributes student effort from open-ended solution construction toward verification, diagnosis, and judgment, but does not automatically produce stronger conceptual transfer. We conclude that GenAI-evaluation activities can be incorporated into algorithms coursework without broad performance losses, but meaningful learning gains may require deliberate scaffolding that pushes students beyond simple error diagnosis. Comments: 11 pages Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI) ACMclasses: K.3.2; I.2.0 Cite as: arXiv:2607.27586 [cs.CY] (or arXiv:2607.27586v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2607.27586 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-110] From Minds to Models: The Intersection of Psychology and LLM Behaviours
链接: https://arxiv.org/abs/2607.27579
作者: Oliver Guidetti,Reza Ryan
类目: Artificial Intelligence (cs.AI)
备注: 32 pages, 1 Figure
Abstract:Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis Comments: 32 pages, 1 Figure Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2607.27579 [cs.AI] (or arXiv:2607.27579v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2607.27579 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-111] What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
链接: https://arxiv.org/abs/2607.27578
作者: Sandeco Macedo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Prompts stopped being isolated strings some time ago. In real systems, one model call feeds another, retrieval interleaves with generation, routers branch, and aggregators merge parallel results. Practice converged on a single structure to hold this together: the graph. Frameworks such as LangGraph, DSPy, and Prompt Flow expose it openly, and research systems already optimize it automatically. The vocabulary, however, lags behind. Graph names, variously, a reasoning topology inside one sampling strategy, a multi-agent conversation, or an orchestration artifact, while prompt engineering still evokes writing one good string. What is missing is a reference definition treating prompts as nodes of an explicit, executable, improvable graph. We build that definition through conceptual analysis over sources with persistent identifiers, complemented by primary grey literature. We reconstruct the genealogy of the idea, from dataflow graphs and build systems, through prompt chaining and the thought topologies (chain, tree, graph), to graphs compiled and optimized as artifacts. We then propose a constitutive definition of prompt graph engineering, state its four conditions (explicit structure, separation between structure and prompt content, executable semantics, and the graph as a first-class engineering artifact), and operationalize them as an inclusion and exclusion test. We draw the boundary against six neighboring concepts and apply the test to six real systems (LangGraph, DSPy, Prompt Flow, AutoGen, CrewAI, and Claude Code subagents); it includes and excludes consistently. We close with a research agenda organized along four design tension axes. The contribution is an operational definition and a shared vocabulary for a practice that industry already exercises daily without naming precisely.
[AI-112] DeepResearch Agent System
链接: https://arxiv.org/abs/2607.27562
作者: Yong Huang,Yulu Huang, for theteam Collaboration
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The DeepResearch Agent System is a large language model system engineered for deep information retrieval, multi-step reasoning, and autonomous research tasks. Built upon a sparse activation architecture with 30 billion total parameters of which only 3 billion are activated per token, the system achieves state-of-the-art performance on multiple agent search benchmarks while delivering 3.2 times faster inference compared to dense counterparts of equivalent scale. The system supports a 128K-token context window with hierarchical attention mechanisms that yield 18.7% accuracy and 23.4% recall improvements over standard long-context approaches. A dual-mode reasoning engine provides both a ReAct paradigm for basic multi-step problem solving and an IterResearch mode for high-performance iterative research with up to 20 reasoning steps, collectively delivering a 31.2% accuracy improvement over single-pass baselines. Multi-tool coordination integrates retrieval, computation, web search, and file parsing modules to achieve 92.1% tool-use accuracy. A reinforcement learning optimization framework based on the GRPO algorithm provides token-level policy gradients that improve training stability by 35% and accelerate convergence by 42%. An automated data synthesis pipeline with seed-based expansion achieves a 92.5% usability rate. Benchmark results include 87.3% on Humanity’s Last Exam, 85.3% on BrowserComp Chinese, and 91.2% on WebWalkerQA. The system is fully open-sourced, including data synthesis, training, and inference code, and supports applications in academic research, business analysis, RD support, and education.
[AI-113] AI Literacy: An Exercise in Power-Knowledge
链接: https://arxiv.org/abs/2607.27547
作者: Brady D. Lund,Zoë Abbie Teel
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:
Abstract:As generative artificial intelligence becomes one of the most significant systems of knowledge production in our society today, questions relating to who can access and shape that production grow increasingly important in our discourse. This paper argues that the existing frameworks for AI literacy, which are dominated by technical competency and responsible-use principles, are insufficient because they enforce a “consumer” orientation toward AI rather than fostering genuine epistemic agency. Based upon Foucault’s concept of power-knowledge, Freire’s pedagogy of critical consciousness, and scholarship of digital literacy, this paper proposes a reconceptualization of AI literacy as a critical practice that equips individuals not just to use AI systems, but to critically evaluate them, resist their structuring assumptions, and participate in their governance. The paper further argues that unequal access to AI tools in society recapitulates longstanding epistemic injustices, and that a literacy framework oriented toward empowerment must account for these structural inequities. A three-part framework of AI literacy based on the notions of contextual use, critical interrogation, and participatory governance frames this literacy as a cultivation of epistemic “agents” rather than the training of competent consumers of AI-generated information.
[AI-114] Expanding Data-Agnostic Pivotal Instances Selection Models with Proximity Trees and Ensemble Learning
链接: https://arxiv.org/abs/2607.27522
作者: Alessio Cascione,Mattia Setzu,Cristiano Landi,Paolo Maria Mancarella,Riccardo Guidotti
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:As decision-making processes grow more complex, machine learning tools have become essential for tackling business and societal challenges. However, many existing methods rely on decision-making procedures that are difficult to interpret. Since humans naturally make decisions by comparing new cases with a few representative examples, we aim to design an approach that selects such pivots to construct an interpretable predictive model. Inspired by decision trees, we propose a hierarchical, interpretable-by-design pivot selection model based on the similarity between pivots and input instances. Our method functions both as a pivot selection technique and a standalone predictive model. Extending beyond single pivots, we incorporate pairs of pivots that are used by proximity and oblique trees, as well as ensembles, which enhance the versatility and effectiveness of our proposal. Additionally, our approach is data modality-agnostic, leveraging pre-trained networks for data transformation. Experiments across diverse datasets, including tabular data, text, images, and time series, demonstrate the effectiveness of our approach, outperforming alternative instance selection strategies and achieving competitive results against state-of-the-art interpretable models while maintaining a minimal number of pivots.
[AI-115] Automated Transcript Analysis for Detecting Flaws in Agent ic Benchmarks
链接: https://arxiv.org/abs/2607.27518
作者: Jeff Mohl,Nelson Gardner-Challis,Magda Dubois,Harry Coppock,Benjamin Allan-Rahill,Kaelan Yim,Damian Sójka,James Mann,Justin Olive
类目: Artificial Intelligence (cs.AI)
备注: 49 pages, 19 figures, Preprint
Abstract:Capabilities of frontier models are often assessed using agentic benchmarks. To trust these results, benchmarks must accurately measure what they claim to and be free from invalidating flaws. Previous manual audits of benchmarks such as SWE-Bench-Verified have uncovered several validity issues in transcripts. However, manual review is difficult to scale, and it is unclear whether automated methods can reliably surface flaws that compromise benchmark validity. In this paper, we developed AI scanners to detect four types of validity issues: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. We produced grading rubrics for each to instruct human labeling, and evaluated the scanners against human labels on a held-out test set of Inspect Evals benchmarks. Our scanners identified several verified quality issues in five widely used benchmarks, including cases unlikely to be caught by random manual inspection. Not all cases were identified, and scanner performance varied substantially across benchmarks, criteria and models. We highlight several open challenges to be addressed to improve scanners for stronger quality assurance claims, including broader standardization gaps in the evaluation field that degrade scanner performance. Together, these results serve as a proof of concept for using automated transcript analysis to audit benchmark quality more broadly.
[AI-116] A dataset of rated conceptual arguments
链接: https://arxiv.org/abs/2607.27499
作者: Emery Cooper,Caspar Oesterheld,Linh Chi Nguyen,Alexander Kastner,Ethan Perez
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models have improved rapidly on tasks with verifiable answers, such as mathematics and programming. Much less is known about their ability to reason about what we call conceptual questions: questions for which no ground truth is realistically accessible and no widely accepted resolution methodology exists, but on which progress can still be made by debating arguments. Most philosophical questions are of this kind, as are central components of questions in AI safety, decision theory, and social choice. Our approach is based on the view that while bottom-line conclusions on such questions are hard to evaluate, individual contextualized arguments can be evaluated far more reliably. We therefore introduce a dataset of 951 argumentative critiques of 442 position texts, spanning topics from AI safety and decision theory to ethics and politics, with 1,458 ratings by six expert raters along dimensions including centrality, strength, correctness, and clarity. We propose two scoring functions and benchmark a range of models. Performance tracks general capability rankings.
[AI-117] MedLLM : An Open Medical Language Model at the Sub-Billion Scale
链接: https://arxiv.org/abs/2607.27490
作者: Maxx Richard Rahman,Asim Ahmed,Mihan Mohagheghzadeh,Wolfgang Maass
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Open medical language models have converged on a single scale: every widely used system runs at 7B parameters or more, leaving the sub-billion regime uncharacterized. We present MedLLM, an open 0.1B-parameter medical language model trained through a fully open three-phase pipeline: general pretraining with curriculum sequence-length scheduling, domain fine-tuning on MedFineWeb, a reference-guided medical corpus we release that is selected from general web data by embedding similarity to medical question-answering (QA) data, and preference-aligned fine-tuning combining SFT with direct preference optimization (DPO). Across medical benchmarks, MedLLM shows a pattern visible only at sub-billion scale: medical competence does not degrade uniformly under compression but splits by task type. On context-grounded QA it comes within 2.9 pp of a medically adapted 7B model and surpasses the instruction-tuned and general-purpose 7B baselines; on knowledge-recall QA it stays near the task floor on clinical-vignette MedQA yet significantly exceeds every 7B and sub-7B baseline on MedMCQA, indicating that where recall fails the constraint is model capacity rather than adaptation. This dissociation is masked at 7B, where both capabilities are present, and surfaces only when capacity is scarce.
[AI-118] INCLAIR: Inception-Based Longitudinal Clinical Anomaly Detection with Informed Reasoning
链接: https://arxiv.org/abs/2607.27487
作者: Maxx Richard Rahman,Wolfgang Maass
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Detecting anomalies in longitudinal clinical profiles is clinically important but difficult: abnormal evidence is often sparse, patient histories have unequal length, and expert explanations are costly. We propose INCLAIR, a framework that scores each observation against multiple historical contexts, aggregates evidence at the profile level, and generates grounded natural-language explanations under limited expert supervision. Under stated within-profile exchangeability assumptions, the complete mean subsequence score takes an order- l U-statistic form, yielding a variance decomposition and an incomplete-subset approximation that controls combinatorial inference cost independently of profile length. The same analysis shows that mean aggregation attenuates localized anomalies by a factor set by the anomaly support and profile length, motivating validation-selected top- k pooling. Across three clinical datasets, INCLAIR consistently outperforms state-of-the-art baselines. We further validate practical relevance through a case study on longitudinal steroid profiles, comparing INCLAIR’s predictions and explanations against domain-expert assessments supported by DNA analysis. The results show that INCLAIR enables clinically actionable anomaly detection under limited expert supervision.
[AI-119] VAmoS Bench: Voice Agent Simulation Bench
链接: https://arxiv.org/abs/2607.27453
作者: Joshua Meyer,Sahar Shayegan,Ritiz Tambi,Ali Khan,Sun Kim,Victor Shih,Mehdi Jamei,Andi Partovi
类目: Artificial Intelligence (cs.AI)
备注: 12 pages, 3 figures, 2 tables. Agent implementations: this https URL
Abstract:Production voice agents span cascaded, speech-to-speech, and hybrid architectures. Voice-agent benchmarks typically measure component quality and conversational properties such as word error rate, latency, naturalness, and turn-taking. Fewer measure whether the agent handled a phone call correctly on its own. Contact centers refer to this as ``containment’': the share of phone calls the automated system resolves without handing off to a human. On some phone calls the right outcome is refusal or a redirect. To address this gap, we introduce VAmoS Bench, the Voice Agent Simulation Bench. It measures complete voice-agent systems end to end in a stateful customer-support task. The agent is Riley, a credit-card support representative for a fictional bank who can freeze, cancel, replace, or activate a card. Each of 100 scenarios supplies a simulated caller with a private goal and a seeded PostgreSQL backend. The platform uses each scenario to populate and activate an isolated simulation in which the caller reaches Riley over audio; roughly one-third apply adversarial pressure. The agent can use five tools that execute real SQL against the backend. Each scenario also defines binary assertions. A grader evaluates them against the complete trace of what the caller and agent said and what the agent did, including tool invocations, arguments, and returned rows. This catches an agent that claims to have changed a card without updating the database, as well as one that makes the right database change while disclosing protected information. This first benchmark version focuses on financial services. Its evaluation protocol supports an evolving leaderboard: additional voice agents can be evaluated on the same version, while later versions can expand the tasks and scenarios.
[AI-120] Leverag ing Trajectory Graphs for Pre-Execution Error Diagnosis in Agent ic LLM Systems
链接: https://arxiv.org/abs/2607.27443
作者: Xu Zheng,Zhuomin Chen,Chaohao Lin,Hua Wei,Haifeng Chen,Wei Cheng,Dongsheng Luo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Large Language Model~(LLM)-based agents have demonstrated exceptional performance across a wide range of complex interactive tasks. However, they often struggle with long-horizon interactive tasks common in domains, such as embodied AI. The complexity and vast action spaces in these settings lead to compounding errors, where a single suboptimal action can derail an entire trajectory, causing the agent to exhaust its limited step budget on inefficient or unrecoverable paths. To overcome this without costly fine-tuning, we draw inspiration from software debugging, where execution logs are analyzed to preemptively catch errors. We propose \textitTrajectory Graph Copilot, a novel framework that acts as a ``copilot’’ for LLM agents by diagnosing potential action errors before they are executed. At its core,\textitGraph Debugger models historical trajectories as a probabilistic graph and uses a Graph Neural Network to identify sequential action patterns that frequently lead to failure. Functioning as a proactive diagnostic sandbox, our method provides early warnings on potentially flawed actions, prompting the agent to self-correct. This pre-action error diagnosis prevents costly mistakes, significantly enhancing the agent’s ability to complete long-horizon tasks successfully. The extensive experiments on four benchmarks with three LLM agents demonstrate a 14.69% pass ratio improvement on average.
[AI-121] SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
链接: https://arxiv.org/abs/2607.27431
作者: Yikun Bai,Binghang Lu,Yikai Liu,Elaheh Akbari,Soheil Kolouri,Linxuan Wang,Ping He,Shuchan Wang,Ruqi Zhang,Guang Lin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Existing diffusion and flow-matching models produce high-quality backbones on SE(3)^N, but inference requires numerically integrating an ODE over hundreds of network evaluations, each involving a Lie group exponential map - a bottleneck for high-throughput design campaigns. We introduce SE(3)-MeanFlow, a few-step generative framework that extends MeanFlow from Euclidean space to the Lie group geometry of protein frames. Working natively in the Lie algebra so(3) and in R^3, we derive closed-form average-velocity identities for rotations and translations, giving simulation-free training targets. We further introduce an SE(3) alpha-Flow objective that removes the Jacobian-vector product from the rotation branch and serves as a warm-up stage, after which training switches to a small-t stabilized MeanFlow loss that is used for the remainder of pretraining and for rectification-based post-training. In protein backbone generation, SE(3)-MeanFlow matches or exceeds flow-matching baselines that use several times more sampling steps, and its advantage widens in the few-step regime, where rectification lets it lead at every matched budget - at a modest cost in diversity.
[AI-122] Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
链接: https://arxiv.org/abs/2607.27415
作者: Xu Zheng,Chaohao Lin,Zhuomin Chen,Weijieying Ren,Haifeng Chen,Wei Cheng,Dongsheng Luo
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Recent advancements in inference-time scaling have significantly unlocked the complex reasoning capabilities of Large Language Models~(LLMs). However, for agents, these approaches suffer from a critical inefficiency, operating in a stateless manner and engaging in redundant search processes. Existing memory mechanisms largely rely on the reasoning capabilities of LLMs, leading to prohibitive computational costs. In this paper, we propose a novel framework, \textitGAMER~(Graph-based Action-centric Memory with Episodic Reasoning), that bridges the gap between inference scaling and episodic memory. Our approach models historical reasoning as a dynamic \textitAction-Centric Graph. By decoupling the memory mechanism from LLMs, our method can save token/money usage by providing less memory context than memory mechanism baselines. To extract knowledge from the graph effectively, we use a dual-stream Temporal Difference learning mechanism to estimate the positive~(suggestion) and negative~(avoidance) value of action nodes based on past successes and failures. During the inference phase, this learned value function optimizes decision-making bi-directionally, so that positive values provide action suggestions, while negative values indicate high-risk actions. By performing efficient searches on the graph, our method significantly improves the efficiency of inference scaling. Experiments on multiple benchmarks demonstrate that \textitGAMER achieves superior performance by \textbf20.81%/6.17% for success/progress rate compared to vanilla baselines.
[AI-123] SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements
链接: https://arxiv.org/abs/2607.27409
作者: Pengyu Xue,He Yang Yuan,Xin Wang,Junkai Chen,Haonan Zhang,Boyuan Chen,Zishuo Ding,Zhenhao Li,Weiyi Shang
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored. In real-world software development, developers continuously improve software quality without changing observable behavior, yet existing benchmarks primarily evaluate functional correctness and provide limited support for assessing these non-functional improvements. In this paper, we present SWE-NFI, a benchmark for evaluating coding agents on NFIs beyond functional correctness. Our benchmark contains 188 tasks constructed from real merged pull requests in open-source Python projects. We operationalize developer-oriented NFIs into 92 executable rules and develop a comprehensive evaluation suite that combines functional correctness testing with rule-based NFI evaluation. We evaluate state-of-the-art commercial and open-source coding agents. Although the best-performing agent achieves a 70.0% functional correctness rate, all evaluated agents generally fall short of human developers in overall NFI capability. The gap is particularly evident for structural code improvements, where agents’ NFI scores range from 0.0 to 1.3, compared with 1.5 for the human reference. Our benchmark and findings provide a reproducible foundation for evaluating and advancing coding agents beyond functional correctness. Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI) Cite as: arXiv:2607.27409 [cs.SE] (or arXiv:2607.27409v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2607.27409 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-124] ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders
链接: https://arxiv.org/abs/2607.27404
作者: Yixuan Duan,Wei Qiu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Research paper. Includes supplementary material and publicly available source code
Abstract:Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses. We introduce ECG-InterpBench, a benchmark designed to systematically evaluate the interpretability of ECG foundation-model representations. ECG-InterpBench uses sparse autoencoders as standardized measurement instruments and matches their capacity across models to enable controlled comparisons. We evaluate six frozen ECG foundation models across five standardized encoder depths, five matched dictionary widths, and three random seeds, producing a 450-cell interpretability atlas comprising 75 exactly matched six-model comparison blocks. The benchmark evaluates complementary dimensions of representation interpretability, including sparse reconstruction fidelity, single-feature accessibility and coverage of 49 clinically meaningful ECG measurements, and cross-seed feature reproducibility. The evaluation further quantifies patient-sampling uncertainty, depth- and seed-dependent variation, and sensitivity to the sparsity parameterization. The benchmark reveals that ECG foundation models exhibit distinct interpretability profiles. A matched replication on MIMIC-IV-ECG confirms that reconstruction fidelity and clinical accessibility identify different leading models. The benchmark is accompanied by executable evaluation code, standardized manifests, cell-level metrics, and reproducibility audits. ECG-InterpBench complements performance-centered ECG benchmarks by providing a capacity-controlled and reproducible framework for comparing ECG foundation models across distinct dimensions of representation interpretability.
[AI-125] FunL2O: LLM -Guided Feature Function Design for Learning to Optimize
链接: https://arxiv.org/abs/2607.27389
作者: Bingheng Li,Junyang Cai,Yupeng Zhang,Bistra Dilkina,Jayant Kalagnanam,Dzung T. Phan
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Learning-to-optimize (L2O) methods accelerate repeated optimization by training models to predict solutions, warm starts, branching decisions, or other forms of solver guidance. A critical yet largely overlooked component of these pipelines is the feature function that maps problem instances to inputs for machine learning models. Existing L2O methods typically rely on hand-crafted features, making representation design manual and largely fixed across domains. We introduce FunL2O, the first unified framework for automating feature design through LLM-driven program evolution for L2O. In a FunSearch-style loop, an LLM proposes executable feature functions, while a fixed evaluation process retrains the original L2O model and measures downstream optimization performance. We evaluate FunL2O on linear and quadratic programming tasks involving solution prediction and warm-starting, as well as on mixed-integer optimization tasks using GNN-guided backdoor branching and Predict-and-Search. Across continuous and discrete optimization tasks and four LLMs, the evolved features consistently outperform hand-crafted representations. These results establish LLM-driven feature evolution as a general and effective approach to automating representation design in L2O.
[AI-126] Beyond the Bidirectional Promise: Re-evaluating the Robustness of Diffusion Language Models
链接: https://arxiv.org/abs/2607.27386
作者: Saurabh Yadav,Badri Narayana Patro,Vijay Srinivas Agneeswaran
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Diffusion Language Models (DLMs) offer a compelling alternative to autoregressive (AR) generation by enabling bidirectional context and iterative refinement. However, their reliability under natural input noise and adversarial attacks remains under-explored. To address this, we systematically evaluate DLM robustness and calibration against AR baselines, using two parameter-matched pairs (LLaDA-8B vs. LLaMA-3-8B and Dream-7B vs. Qwen2.5-7B) across 32 natural perturbation conditions, adversarial gradient probes, and mechanistic hidden-state analyses. This paired design effectively isolates architecture-intrinsic properties from weight-dependent behaviors. We find a nuanced robustness profile: while highly stochastic DLM loss landscapes naturally resist gradient-based adversarial suffixes, they provide no guaranteed defense against natural noise, proving that everyday robustness is weight-dependent rather than inherently architectural. Furthermore, DLMs exhibit systematic overconfidence, presenting a practical deployment hazard. Most crucially, mechanistic probing reveals that all models perfectly encode input corruption, isolating behavioral fragility entirely to a decoder routing failure. Consistent with this diagnosis, we show that surface-level prompt patching fails to improve over noisy baselines. Ultimately, DLM robustness cannot be patched on; it must be fundamentally integrated into the iterative decoding loop.
[AI-127] RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
链接: https://arxiv.org/abs/2607.27373
作者: Benyamin Tafreshian,Prathamesh Dhake
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: This manuscript supersedes the preliminary version available as arXiv:2511.18790 . The work has been substantially revised, expanded, and reorganized, with a refined threat model, revised methodology, clearer stage-level evaluation criteria, and expanded analysis of moderation bypass, instruction reconstruction, and execution
Abstract:Large language models (LLMs) are becoming increasingly integrated into mainstream development platforms and daily technological workflows, typically behind moderation and safety controls. Despite these controls, preventing prompt-based policy evasion remains challenging, and adversaries continue to “jailbreak” LLMs by crafting prompts that circumvent implemented safety mechanisms. Prior work has established cipher-mediated interaction, code-embedded decryption, prompt decomposition and reconstruction, and layered custom encryption as viable attack primitives. However, reported evaluations generally collapse visible acceptance, successful recovery of the concealed request, and subsequent execution into an aggregate attack-success outcome. This leaves limited evidence about where multistage prompt-transformation attacks fail within an observable black-box interaction. This paper introduces RoguePrompt, a jailbreak pipeline that partitions a forbidden prompt and applies two nested encodings, Vigenere followed by ROT13, along with natural-language reconstruction instructions. RoguePrompt was developed and evaluated under a black-box threat model, with only API or user-interface access to the hosted models, and was tested on 313 real-world, hard-rejected prompts. Success was measured in terms of moderation bypass, instruction reconstruction, and execution when the relevant stage exceeded its automated criterion. RoguePrompt achieved average rates of 93.93% for filter bypass, 79.02% for reconstruction, and 70.18% for execution. These results demonstrate the effectiveness of layered prompt encoding while providing stage-level evidence of where multistage jailbreaks fail during moderation bypass, instruction reconstruction, and execution.
[AI-128] SkillM entor: LLM Agent Self-Evolution via Learning Blind-Spot Diagnosis
链接: https://arxiv.org/abs/2607.27360
作者: Xiaoyi Bao,Yuanzhen Xie,Yunzhi Tan,Jinghang Gu,Zhongqing Wang,Chu-Ren Huang,Bo Hu,Zang Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Agent self-evolution has primarily focused on learning how to act, while overlooking an equally important capability: learning to discover what an agent does not know. Existing approaches typically assume that failure discovery is given, focusing on how to repair failures once they are identified. We ask whether blind-spot diagnosis itself can be learned. We thus study diagnosis as an agent capability separate from execution, and exclude two alternative sources of progress: executor adaptation and human supervision. Under these constraints, performance cannot improve through executor updates or annotated examples, forcing all improvements to originate from the learned diagnostic capability. We propose SkillMentor, which trains a Mentor policy via reinforcement learning to generate diagnostic tasks, identify recurrent failure modes, and curate them into reusable corrective skills. Across AppWorld and BFCLv3, SkillMentor improves executor performance by an average of 44.2%. These results suggest that blind-spot diagnosis is a learnable capability, enabling self-evolution without updating executor weights or relying on human-curated data.
[AI-129] PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
链接: https://arxiv.org/abs/2607.27354
作者: Haoyu Chen,Xirui Shi,Yuyao Wang,Jerry Chen,Di Niu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Personal AI assistants are increasingly deployed as task-oriented, tool-augmented agents that operate within unified service environments to support everyday user activities. In realistic settings, such assistants must reason over persistent user state, respect user-specific configurations and permissions, and sustain long-horizon, constraint-aware interactions across multiple services. Existing benchmarks, however, often fragment service contexts or abstract away user state, limiting their ability to evaluate user-centric personal assistant behavior in realistic service settings. We introduce PAUSE, a user-centric benchmark for evaluating personal AI assistants in stateful, service-integrated environments. PAUSE captures core challenges of real-world assistant deployment by requiring agents to coordinate actions across heterogeneous user-owned resources while maintaining consistency with environment state, authorization constraints over multi-turn interactions. The benchmark incorporates explicit user-agent interaction via realistic user simulation, enabling evaluation beyond static tool execution. To support principled and reproducible evaluation, PAUSE adopts a multi-regime evaluation framework aligned with task characteristics. Open-ended service management tasks are assessed using semantic and trajectory-level behavioral metrics, while constraint-intensive tasks admit deterministic, state-based verification. Benchmark results show that even state-of-the-art proprietary models fail to reach 70% task completion on scenarios requiring stateful reasoning and configuration awareness, revealing consistent and interpretable failure patterns. Finally, we present a user-centric synthesis pipeline that enables scalable generation of coherent service environments, user configurations, and reliably annotated tasks, supporting benchmark extensibility and future research.
[AI-130] PIE-APT: A Unified Framework for Temporal Planning and Contradiction Hunting via Incremental Direct-Derivation Abduction
链接: https://arxiv.org/abs/2607.27287
作者: Amir Hossein Sharafi,Alireza Shahbazi
类目: Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)
备注:
Abstract:Reasoning and planning over Dynamic Knowledge Graphs (DKGs) present significant challenges, especially in open-world environments with incomplete information. Existing action formalisms often face decidability issues and the Ramification Problem, while managing incomplete knowledge via structural abduction requires expansive combinatorial search. This paper introduces a unified framework with two integrated modules—\textbfPIE-Abducer (incremental direct-derivation abduction) and \textbfPIE-APT (Abductive Planning for Temporal KGs)—operating natively on the highly expressive Description Logic. We model state transitions along a linear timeline as non-monotonic updates to deductively closed DL theories. Treating the incremental reasoner as a black-box and representing actions natively in OWL without external modal operators preserves logical decidability. To address incomplete knowledge, \textbfPIE-Abducer circumvents traditional Minimal Hitting Set (MHS) enumeration. Instead of combinatorial syntactic search, it injects the logical negation of a target goal into a consistent branch and extracts missing premises via direct refutation consequences. \textbfPIE-APT then employs a recursive \textitGenerate-and-Test architecture, interleaving backward-chaining A* search with \textbfPIE-Abducer up to a bounded causal depth, followed by strict validation via forward-chaining Temporal Projection. We evaluate four OWL benchmarks stressing semantic abilities absent in classical planning: parameterized goals with witness search, mid-search DL entailment, open-world assumption injection, and adversarial contradiction hunting. Results demonstrate qualitative superiority over classical planners and prove our direct-derivation approach quantitatively outperforms an MHS-faithful baseline during abductive enrichment.
[AI-131] Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance
链接: https://arxiv.org/abs/2607.27283
作者: Chao Peng,Zhiheng Lyu,Peijie Dong,Hande Dong,Qiang Lin
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
备注:
Abstract:Long-horizon benchmarks often show that agents fail more as tasks become longer. This observation is useful for deployment, but it does not by itself explain why failure occurs. More stages create more opportunities for ordinary errors to compound; longer tasks may also contain harder individual decisions or become harder as conversation history, tool outputs, and environment changes accumulate. We use trajectory-induced degradation to mean this last possibility: earlier execution makes later work harder. When the harmful accumulation is specifically the text visible to the model, it is often called context rot. In this position paper, we argue that to claim a “long-horizon failure”, benchmarks must compare actual full-task success against a baseline prediction built from short, individual stages. We call the log-ratio between this prediction and actual success the horizon residual. The comparison must use the same agent configuration and specify in advance how stages, checkpoints, information, and budgets will be chosen. The residual shows that the full rollout differs from the chosen baseline; targeted experiments are still needed to explain why.
[AI-132] Flat Score Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
链接: https://arxiv.org/abs/2607.27275
作者: Jiwon Jang,Kisu Yang,Heuiseok Lim,Hyunwoo Park
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: preprint
Abstract:Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On \tau^2 -bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison correction, and in the cell that carries the largest process damage, equivalence testing bounds the change within \pm 7.5 points. The process tells a different story. Quantization amplifies the failure the model already exhibits at full precision (tool-name hallucination in telecom, with the same directional trend in retail entity errors) by up to 2.5 \times in volume (+17.6 points per task), while creating essentially no new failures. The failure set is the same at every precision (rank correlation \geq 0.94, 0.18% novel events). The score stays flat because the benchmark’s ten-error budget absorbs the extra failures. Shrinking the budget to two errors re-exposes a score gap of 17 points, and it does so only in the one cell where quantization added error volume, exactly as the masking account predicts. A targeted error-repair prompt, run for five telecom models at every precision, removes the damage exactly and only where it lives. Both diagnostics, the per-channel error rate and success under a shrinking budget, come from logs benchmarks already collect; we suggest reporting them alongside task reward.
[AI-133] FAVA: Formal Authorization for Verified Agents with Evidence-Backed Permission Graphs
链接: https://arxiv.org/abs/2607.27267
作者: Yifan Zhang,Xinkui Zhao,Sai Liu,Hengxuan Lou,Guanjie Cheng,Chang Liu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents autonomously interleave semantic reasoning with complex system operations. In these dynamic environments, static tool-level permissions are fundamentally insufficient; safe authorization is highly context-dependent and heavily reliant on evolving runtime states and data flows. We present FAVA (Formal Authorization for Verified Agents), a permission-carrying authorization framework for agent execution. FAVA utilizes an LLM-guided Permission Intermediate Representation (IR) to translate ambiguous natural-language tasks into structured constraints. A deterministic lowering pass then converts this IR into an evidence-backed permission graph that explicitly tracks data flows, dependencies, and contextual labels. To provide strict security guarantees, a Satisfiability Modulo Theories (SMT) authorizer mathematically verifies the current graph against security policies before any effectful action executes. A runtime gateway then enforces the solver’s result, either authorizing the execution or intercepting it with a precise counterexample. We evaluate FAVA across OpenAgentSafety, OctoBench, and ActPlane scenarios. Our evaluation demonstrates that FAVA achieves a 90.5% Decision Compliance Rate (DCR) over the aggregate dataset, successfully intercepting dynamic violating traces in the evaluated trace-conditioned scenarios.
[AI-134] Recursive transformers for semiconductor thermo-mechanical reliability
链接: https://arxiv.org/abs/2607.27251
作者: Kart-leong Lim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design. But conventional transformer architectures are often over parameterized for the small, low-dimensional datasets typical of engineering design spaces, where large simulation data is expensive to generate. Under these conditions, excess parameter capacity leads to overfitting rather than improved accuracy, while also incurring unnecessary memory and compute overhead. This motivates a shift towards architectures that focus on additional compute rather than additional learnable parameters. This paper presents a hardware-aware evaluation of three recursive transformer paradigms for surrogate thermo-mechanical analysis of advanced packages: a)Tiny Recursive Model, b) our proposed Depth Recursive transformer, c) and a simple recursive transformer. We systematically compare their predictive performance (Recall, Mean Reciprocal Rank), parameter count, computational complexity (FLOPs), providing practical design guidelines for selecting recursive transformer architectures under resource-constrained scenarios. We validate this principle on two low-dimensional engineering prediction tasks: 1) thermo-mechanical reliability analysis of advanced semiconductor packages, where stress and warpage from thermal cycling must be evaluated repeatedly across a design-of-experiments sweep under costly finite element analysis (FEA). 2) Laplace PDE iterative numerical solver for capacitance field. Overall, recursive weight-sharing transformers provide an effective and generalizable trade-off between prediction accuracy, parameter efficiency, and computational cost for small data engineering surrogate modeling.
[AI-135] Do Context Files Help Coding Agents ? A Two-Agent Ablation Study on Real Repositories
链接: https://arxiv.org/abs/2607.27250
作者: Prakhar Khatri
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Persistent context files (this http URL, this http URL) are standard practice for guiding AI coding agents, yet evidence for their effectiveness is contradictory. We present a controlled ablation of context-injection strategy across two frontier agents (Claude Code and Codex), 17 real tasks from 3 repositories (15 shared + 2 Codex-only), and 288 evaluated runs with gold-test evaluation. Context strategy does not measurably move correctness on either agent (bounded to =10-15pp via equivalence testing). A failure-mode triage reveals why: agents fail on implementation skill—feature design, pattern selection, exact wiring—not missing repository knowledge that a context file could supply; a manipulation probe confirms the real this http URL never converts a near-miss to a pass on either agent. We further show that borderline task difficulty is agent-specific (Spearman rho=0.75), offering a candidate explanation for prior contradictions: single-agent studies draw tasks from different agents’ informative bands. We release all code, data, and analysis.
[AI-136] Divergence Decoding: Training-Free Capability Fusion
链接: https://arxiv.org/abs/2607.27248
作者: Yimi Wang,Hao Li,Shuo Yang,He Cao,Dechen Zhang,Ziang Wu,Zhiyuan Yan,Fanyang Mo,Li Yuan
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:While large language models excel in reasoning, these generalists often lack knowledge for specialized scientific domains. Conversely, domain models~(specialists), while knowledgeable, suffer from specialization side-effects including diminished logic and reduced this http URL address this dilemma, we introduce Divergence Decoding, a training-free framework for capability fusion. It reconstructs the “draft-and-verify” skeleton of speculative decoding into an adaptive routing mechanism. The core is using Jensen-Shannon divergence to monitor the distributional disagreement between the two models at each token. When the specialist exhibits significant divergence, our method identifies it as a potential reasoning risk and instantaneously routes control to the generalist. This allows the dynamic injection of general reasoning while preserving domain expertise, achieving inference-time policy composition of the generalist and the this http URL evaluate Divergence Decoding across diverse model families (Qwen and Llama series) on challenging scientific benchmarks (GPQA, ChemBench, and ChemCoTBench). Experimental results demonstrate that Divergence Decoding outperforms both the domain-specialized and general-purpose models, effectively surpassing the performance of most single-model baseline. This suggests that Divergence Decoding provides a general, training-free paradigm for fusing diverse LLM capabilities through adaptive inference-time collaboration.
[AI-137] Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage
链接: https://arxiv.org/abs/2607.27245
作者: Vivek Senthil,Zhiqiang Tao,Ernest Fokoué
类目: ound (cs.SD); Artificial Intelligence (cs.AI)
备注:
Abstract:Modern policing faces a “visibility paradox” where law enforcement agencies possess petabytes of Body-Worn Camera (BWC) footage that remains largely unutilized for accountability or systemic review due to the prohibitive labor costs of manual transcription. This research presents a framework for adapting the OpenAI Whisper architecture to the unique acoustic and linguistic challenges of the policing environment. By employing Parameter-Efficient Fine-Tuning (PEFT) through Low-Rank Adaptation (LoRA), we address the significant performance degradation observed in zero-shot models when confronted with high-stress scenarios, sirens, and radio interference. Crucially, we demonstrate that this adaptation is feasible on consumer-grade hardware (Acer Nitro local machine with NVIDIA 4GB GTX GPU) using 8-bit quantization and gradient checkpointing. We further integrate these transcriptions into a symbolic reasoning pipeline using a domain-specific ontology to transform raw audio into evidence-linked incident graphs, achieving a 93.7% lexicon mapping rate for the advancement of procedural justice and transparency.
[AI-138] Asymmetric Collapse in Model Merging: When Refusal Over- writes Recognition
链接: https://arxiv.org/abs/2607.27240
作者: Aarnav Choudhary,Matheus Fonseca Rocha,Jiwon Seo,Vasu Sharma,Maheep Chaudhary
类目: Artificial Intelligence (cs.AI)
备注: Accepted into COLM AIW, waiting decision for AdvML-Frontiers x CoTMA and HAIPS
Abstract:Model merging is often used to combine capabilities from separately fine-tuned models without additional training, but it is unclear whether standard merging methods preserve multiple safety-relevant behaviors simultaneously. We study this question through a controlled case study using two Gemma-3-1B-IT finetunes on two complementary safety objectives: CARES harm-level classification and WildJailbreak adversarial refusal. We merge the two fine-tunes using Linear, SLERP, TIES, and DARE-TIES, and evaluate the merged models on classification accuracy, attack resistance, and benign compliance. Across all four methods, attack resistance transfers significantly more than classification accuracy: merged models retain 81-85% jailbreak refusal rates while CARES accuracy falls to at most 12.9%. Weight-space measurements suggest that this asymmetry is not caused by strongly opposing task-vector directions: the two task vectors are nearly orthogonal (cosine similarity 0.011). Instead, the refusal fine-tune induces consistently larger per-layer task-vector magnitudes, causing magnitude-sensitive methods to favor refusal updates. These results show that standard model merging can collapse safety recognition into broad refusal when safety-relevant task vectors differ substantially in scale.
[AI-139] KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM -based Kernel Generation
链接: https://arxiv.org/abs/2607.27231
作者: Peiyu Zang,Jian Tao,Jialing Zhang,Yichen Yuan,Wentao Zhang,Guang Liu,Yonghua Lin
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 10 pages, 4 figures. Code and data are publicly available at this https URL
Abstract:Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task. The recent rise of LLMs and agentic frameworks offers a promising pathway toward automatic kernel generation. However, despite rapid progress, there is still no comprehensive benchmark to rigorously evaluate LLM-generated kernels across diverse operator sources or heterogeneous hardware platforms. We present KernelGenBench, a unified benchmark for systematically evaluating LLM- and agent-generated Triton kernels across diverse operator sources and heterogeneous hardware platforms. It comprises two complementary sub-benchmarks: KernelGenBench-MS (Multi-Source), evaluating 210 operators from three sources beyond standard PyTorch-centric tasks, and KernelGenBench-MC (Multi-Chip), measuring performance portability across six heterogeneous hardware platforms using a 110-operator subset. Our large-scale evaluation, consuming over 15 billion tokens, shows: (1) agent-based methods consistently outperform pure LLM sampling methods, while cuBLAS operators are the most challenging across all methods; (2) generation performance varies significantly across hardware platforms, with even recent kernel-specialized agents experiencing severe cross-platform degradation (e.g., AutoKernel drops from 87% on NVIDIA to 25% on Platform E); (3) autonomous kernel generation remains highly cost-intensive, with specialized agent methods averaging 5.11 million tokens per successful operator (AKO4all reaches 5.19 million), orders of magnitude higher than simple LLM sampling approaches.
[AI-140] Multi-Head Attention Residuals
链接: https://arxiv.org/abs/2607.27230
作者: Cheng Luo,Zefan Cai,Junjie Hu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by letting each sublayer attend, through a learned softmax. However, that read uses a single query shared across the entire width, so every feature subspace must read the depth history through one distribution. The cost of this forced compromise grows with how much the subspaces disagree about which layers to read, and disagreement grows with model width. We introduce Multi-Head Attention Residuals (MHAR): the routing query is reshaped into H per-subspace heads, each with its own softmax over the depth history. The read becomes block-diagonal, the reshape adds zero parameters and negligible compute, and H = 1 recovers attention residuals exactly. Trained from scratch on a deduplicated Nemotron-based anneal corpus that is quality-filtered and STEM- and code-heavy, MHAR improves validation loss over a standard Transformer at 100M, 350M, and 1B (-0.061, -0.149, and -0.140). It achieves the best result among four methods in every setting, with the gain increasing from 100M to the larger scales. The head count is a real design axis rather than a free knob: validation loss is U-shaped with respect to H, with a flat optimum at H = 4 or H = 8 across scales. We adopt H = 8 for large-scale models; over-splitting beyond this point (H = 16) consistently gives back part of the gain. A direct probe of the trained queries confirms that learned subspace disagreement is the underlying driver. Fused Triton routing kernels increase attention-residual training throughput from 0.2-0.5x to 0.55-0.88x of the baseline while maintaining near-baseline peak memory. An identity-preserving conversion using delta attention residuals supports 8B mid-training, yielding improvements of +3.2 on GSM8K and +3.1 on GPQA.
[AI-141] Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
链接: https://arxiv.org/abs/2607.27209
作者: Binyan Xu,Fan Yang,Xilin Dai,Kehuan Zhang
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 20 pages, 7 figures, 4 tables
Abstract:Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no systematic audit has examined whether this instrument functions uniformly across research areas, or whether acceptance outcomes are in practice shaped by forces that reviewer scores neither capture nor control. This position paper argues that acceptance outcomes are shaped by forces beyond reviewer scores, and that the underlying cause is a measurement design failure, not individual bias. When a fixed numerical scale aggregates quality judgments across communities with structurally non-uniform reviewer pools, absolute scores become incomparable across areas, and area chairs must substitute community priors for score-based decisions. Using ICLR 2021–2026 data covering 50,289 papers across 219 research topics, we show that at any given reviewer score, a paper’s acceptance probability varies by up to 8x depending on its topic. We rule out scoring culture, expert reviewer standards, rational area chair reweighting, and quality dilution as alternative explanations. We call on program committees to adopt inherently calibrated review signals and publish topic-stratified, score-conditional acceptance rates as a first-class fairness metric.
[AI-142] Learning to Trace Seiberg Dualities
链接: https://arxiv.org/abs/2607.28628
作者: Jonathan J. Heckman,Shani Meynet,Alessandro Mininno,Gary Shiu
类目: High Energy Physics - Theory (hep-th); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); High Energy Physics - Phenomenology (hep-ph)
备注: 59 pages + appendices, 38 figures. Code and tools available at this https URL
Abstract:Dualities play an important role in establishing both microscopic and emergent phenomena in a wide range of physical systems. In practice, though, it can often be computationally challenging to establish when two systems are dual, even when all of the “rules of the game” are well-known. Said differently, when confronted with two systems, how can one efficiently establish that they are in fact dual? In this paper we use machine learning methods to address this question for Seiberg dualities of supersymmetric quiver gauge theories. Mathematically, this involves establishing mutations of quivers, which is in turn a variation on the theme of “learning to unknot”. On the one hand, this leads us to a practical tool for establishing the computational complexity of different dualities. On the other hand, it also allows us to study how different network architectures learn how to trace Seiberg dualities. We find that for quivers with a modest number of quiver nodes (of order 10 ), different network architectures consisting of transformers and multi-layer perceptrons tend to outperform deterministic algorithms. Supplementing the network by well-established pathfinder algorithms (essentially “Google Maps for quivers”) leads to an additional improvement in the efficiency and accuracy of the search strategy. We anticipate that this class of questions can serve as a useful benchmark for frontier AI models applied to theoretical physics.
[AI-143] Vibe-FDTR: An agent -oriented framework for reproducible frequency-domain thermoreflectance data analysis
链接: https://arxiv.org/abs/2607.28200
作者: Fuwei Yang,Weiheng Li,Bai Song
类目: Applied Physics (physics.app-ph); Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI)
备注:
Abstract:Frequency-domain thermoreflectance (FDTR) is a laser pump-probe technique widely used to measure thermal properties at the micro- and nanoscale; however, it relies on a complex data analysis procedure that demands substantial domain expertise and is susceptible to subtle human errors. Here, we present Vibe-FDTR, an agent-oriented framework that enables large language model (LLM) agents to perform reliable and reproducible FDTR analyses directly from natural language requests. This framework couples a configuration-driven FDTR code package, which enforces physical and parametric consistency, with procedural agent skills that translate user intentions into organized and verifiable analysis steps. We evaluate Vibe-FDTR using a controlled benchmark with two levels: synthetic single-step tasks and real-data multi-step tasks based on measurements of gold-coated graphite samples. Across the two levels, agents using Vibe-FDTR achieve success rates of 100% and 98.9%, respectively. In sharp contrast, ablating skills (Code-agent) reduces performance to 91.4% and 36.7%, which drops further to 38.6% and 0% when the domain package is also omitted (Agent-only). Beyond success rate, Vibe-FDTR also reduces computational cost by 87.7% relative to the Code-agent variant and cuts execution time by more than 60%. Finally, an optional expert mode supports experimental planning via autonomous sensitivity and uncertainty evaluations, and formulates physically grounded recommendations for underspecified tasks. These results demonstrate that encapsulating domain code and expert knowledge into agent skills offers a promising route toward low-barrier, autonomous, and trustworthy thermal metrology.
[AI-144] On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems
链接: https://arxiv.org/abs/2607.28080
作者: Illia Horenko
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and subspaces of relevant features in nonstationary and nonlinear regression problems. It is shown that the proposed extension - that we coin as Entropy-Optimal Manifold Regression (EOMR) - allows a robust learning with linearly-scaling iteration and memory complexities. EOMR is compared to the most complete set of state-of-the-art tools from the Artificial Intelligence (AI) and Machine Learning (ML) that is available to the author, on the very challenging problems from chaotic and fluid dynamics: (i) on predicting the Lorenz-96 systems dynamics in strongly- and very-strongly chaotic regimes (with forcing parameter being F=8 and F=12 , respectively); and, (ii) on a data from the Hasegawa-Wakatani model on the edge of the tokamak plasma. It is demonstrated that the proposed benchmarks (i) and (ii), indeed, are the very challenging problems for the state of the art ML and AI tools - since both the general-purpose gradient boosted random forests and deep neuronal networks, as well as transformer-based AI tools like TabPFN v.03 (more spezialised for large-dimensional small data learning problems) - result in orders of magnitude inferior root mean squared prediction errors, and orders of magnitude larger model complexities, when compared to the EOMR. For a Hasegawa-Wakatani example, EOMR distills a very simple entropy-optimal and skilful description of the leading Essential Orthogonal Function (EOF) dynamics, given by linear, causal and weakly-stationary autoregressive process described by just 8 parameters.
[AI-145] Stimulus-Evoked Network Dynamics in Human Cortical Organoids: From a Graph-Computational Framework to Repeated-Stimulation Depression
链接: https://arxiv.org/abs/2607.28068
作者: Esmaeil S. Nadimi,Vinay C. Gogineni,Jan-Matthias Braun,Martin Røssel Larsen,Victoria Blanes-Vidal,Helle Bogetofte Barnkob
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI)
备注:
Abstract:Human cortical organoids provide an experimentally accessible model of early neural circuit formation, yet whether their activity reflects structured information processing rather than spontaneous synchronization is unclear. We developed a graph-computational framework to quantify stimulus-evoked propagation. This includes stimulus-conditioned functional graphs, a graph-constrained dynamical (graph-neural-network) model used as a system-identification tool, a biological message-passing principle bounding integration depth by observable propagation depth, and a suite of graph-level metrics. We carried this program out in full on longitudinal HD-MEA recordings from three organoids. Once the true acquisition sampling rate and stimulus timing were recovered, the evoked response proved to be a fast, near-synchronous network burst with no measurable outward propagation (peak-latency vs. distance slope = 0). The propagation/integration-depth metrics (Deff ,reachability index, dmax) therefore do not apply, and per-day connectivity graphs were not reliably estimable at the available trial count, a negative result with methodological consequences for applying such metrics to organoid data. Reframing around synchrony, response-population size and shared variability revealed a control-validated phenomenon, i.e., repeated daily stimulation progressively depressed and spatially contracted the evoked response. That repeated stimulation reshapes organoid networks is established, but longitudinal designs in which every preparation is stimulated cannot separate this from developmental maturation. We break that confound with a developmentally-matched, stimulation-naive control, where at day 7, an organoid receiving its first-ever stimulation engaged 93% of the array, whereas organoids with five prior sessions engaged 10%.
[AI-146] Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecasting
链接: https://arxiv.org/abs/2607.27945
作者: Kuo-Chung Peng,Samuel Yen-Chi Chen,Jiun-Cheng Jiang,Chen-Yu Liu,En-Jui Kuo,Yun-Yuan Wang,Tzung-Chi Huang,Prayag Tiwari,Chi-Sheng Chen,Chun-Hua Lin,Yu-Chao Hsu,Tai-Yue Li,Saif Al-Kuwari,Simon See,Kuan-Cheng Chen,Nan-Yow Chen,Hsi-Sheng Goan
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 8 pages, 7 figures
Abstract:Sequence models must decide what to write into memory and what to retain. In quantum and quantum-inspired sequence learning, nonlinear recurrent updates often require repeated circuit evaluations and sequential backpropagation through time, making long contexts costly. Gated fast-weight programmers (FWPs) based on quantum-inspired Kolmogorov-Arnold networks (QKANs) alleviate this bottleneck by storing context in time-varying fast parameters. However, their scalar gate applies one retention-write balance to every fast-state coordinate, forcing all parameters to share a memory timescale. We introduce Self-Modulating QKAN-based FWPs, which replace this broadcast gate with low-rank-generated element-wise modulation of the new-proposal branch, a bounded old-state branch, or both. We further propose Complementary Matrix Gating (CMG), which uses one sigmoid matrix gate to retain the old state and its complement to write the new proposal. CMG provides coordinate-wise memory control while preserving the bounded convex update and affine prefix-scan structure of scalar gating, at the modulation-head cost of a single-branch rule. We compare four self-modulating rules with scalar gating across four FWP architectures combining classical and QKAN-based slow and fast programmers. Across seven single-step forecasting benchmarks and five sequence lengths, CMG gives the most consistent improvements for architectures whose fast programmer incorporates a QKAN-based module. In direct multi-step forecasting of Jaynes-Cummings and transmon-resonator dynamics simulated with CUDA-Q Dynamics, CMG models maintain mean-squared errors on the order of 0.001 or lower across forecasting horizons of 4, 8, and 16 steps, while improving on their scalar-gated counterparts by at least 91.2%. These results establish coordinate-wise complementary modulation as a stable and effective update for QKAN-based FWPs.
[AI-147] Deep Learning for Accelerated Long-Horizon Forecasting of Multicomponent Multiphase Microstructure Evolution in High-Entropy Alloys
链接: https://arxiv.org/abs/2607.27820
作者: Hamidreza Razavi,Nele Moelans
类目: Materials Science (cond-mat.mtrl-sci); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE)
备注:
Abstract:Phase-field modeling provides a powerful approach for predicting microstructure evolution but becomes computationally prohibitive for multicomponent and multiphase systems over large spatial and temporal scales. This work presents an AE-GCN-LSTM surrogate framework for long-horizon forecasting of microstructure evolution in the multicomponent AlCrFeNi high-entropy alloy system containing coexisting BCC and FCC phases. A multi-head autoencoder compresses the four elemental concentration fields and phase-field order parameter into latent representations, which are formulated as graphs for learning their spatial and temporal evolution. The framework accurately forecasts microstructure evolution over horizons extending to 3,000,000 simulation timesteps. Its robustness is systematically evaluated under previously unseen conditions without retraining, fine-tuning, or parameter adaptation. These evaluations include variations in FCC precipitate size and initial position, microstructures containing one, two, and five FCC precipitates, and complex phase interactions involving precipitate merging and splitting. Although trained only on 100 x 100 computational domains containing a single nominal alloy composition, the framework is successfully transferred to larger 256 x 256 and 512 x 512 systems and to previously unseen AlCrFeNi compositions. Across the evaluated configurations, the model preserves the dominant phase morphology and compositional evolution while providing computational speedups ranging from approximately 7200 to 62300 relative to conventional phase-field simulations. These results demonstrate that latent graph-based AE-GCN-LSTM forecasting provides a scalable and computationally efficient surrogate for long-horizon simulation of multicomponent, multiphase microstructures and offers a promising foundation for high-throughput alloy design.
[AI-148] Can AI Follow In Einsteins Footsteps?
链接: https://arxiv.org/abs/2607.27794
作者: Michael Shalyt,Nathan Regev,Marin Soljačić,Ido Kaminer
类目: History and Philosophy of Physics (physics.hist-ph); Artificial Intelligence (cs.AI); Popular Physics (physics.pop-ph)
备注:
Abstract:AI is accelerating physics discovery, but perhaps away from Einstein-level theory building. To understand this gap, we must recognize a striking trend: while being very successful, the most visible AI contributions to physics discovery appear to mirror the historical development of physics, but in reverse. Human discovery in physics progressed, in broad strokes, from ancient pattern prediction, through phenomenological laws such as Kepler’s, to principle-based universal theories such as relativity and the Standard Model. On the AI side, prominent contributions to physics discovery point in the opposite direction: early milestones emphasized explicit equation-discovery methods, such as symbolic regression, whereas more recent frontier contributions are powerful predictors such as AlphaFold and GraphCast, which can be remarkably accurate yet do not provide clear theoretical understanding. If this trend continues, AI would become extraordinarily good at prediction but may struggle to ever propose its first serious contender to quantum gravity or other paradigm-level theories. We review the current landscape of AI for physics discovery and highlight a critical missing skill: the ability to pose the right questions or invent the right principles to guide the development of new theories and the tests to falsify them. This mode of discovery has driven many of the deepest advances since the 17th century, where symmetry, simplicity, and new mathematical frameworks guided theory construction before experimental tests. Equipping AI systems with such skills could move them from predicting within known frameworks to proposing the next paradigm-level discovery in physics.
[AI-149] Rethinking Artificial Intelligence in Medical Imaging: Assumptions Reality and Reframing
链接: https://arxiv.org/abs/2607.27428
作者: Arman Rahmim,Nourhan Bayasi,Xiaoxiao Li,Babak Saboury,Fereshteh Yousefirizi
类目: Medical Physics (physics.med-ph); Artificial Intelligence (cs.AI)
备注: 10 pages, 2 figures
Abstract:Medical imaging has served as primary proving ground for clinical artificial intelligence (AI), yet a decade of intense research has not translated into proportionate bedside impact. We argue that this gap is not primarily a product of insufficient algorithmic performance, inadequate regulation, or limited explainability. Rather, it reflects a structural misalignment, between how AI systems are designed and evaluated, and how clinical decisions are made. This Perspective identifies six interconnected dimensions of this misalignment: the dominance of pixel-only models in a multimodal clinical world; the erosion of physician trust through opaque and inflexible systems; the unfulfilled promise of foundation models in data-sparse medical domains; the persistent bottleneck of non-shareable, under-curated datasets; the gap between validated algorithms and deployable clinical platforms; and the failure of prediction-centric AI to generate actionable clinical guidance. For each dimension, we reframe the problem and propose a path forward, culminating in a vision of agentic, physician-aligned AI that extends, rather than replaces, clinical judgment.
[AI-150] LLM -Guided Initialization for Accelerated Hybrid Quantum-Classical Medical Image Classification
链接: https://arxiv.org/abs/2607.27262
作者: Riza Alaudin Syah,Irwan Alnarus Kautsar,Haza Nuzly Bin Abdull Hamed
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
备注: Presented at IAICT 2026, Bali, Indonesia
Abstract:Variational quantum algorithms often encounter barren plateaus, where cost gradients decay rapidly with increasing circuit depth, undermining the trainability of parameterized quantum circuits. This paper evaluates AdaInit (Adaptive Initialization), proposed by Zhuang and Cunningham, which uses large language models to propose initial parameters for quantum neural networks. We study a simplified single-query AdaInit variant paired with GPU-accelerated simulation in NVIDIA CUDA-Q and apply it to binary classification on the DMR-IR mammography dataset. AdaInit delivers 14.6 times higher gradient variance at initialization than random initialization (0.0095 vs. 0.0006), producing 160 times faster convergence (1.1s vs. 176 s) while maintaining the same classification accuracy of 61.4 percent. We provide theoretical analysis grounded in the geometry of parameterized circuit landscapes and show empirically that LLM-guided initialization places the optimizer in trainable regions of parameter space. Beyond performance, our results indicate that a single LLM query can yield informative parameters without iterative refinement, suggesting a low-overhead path to improved trainability. The findings validate AdaInit in a medical imaging setting and demonstrate its compatibility with GPU-accelerated quantum backends for practical speedups.
[AI-151] Estimating Treatment Effects for Depression in Longitudinal Therapy Switching Settings
链接: https://arxiv.org/abs/2607.27214
作者: Xinyu Qin,Martin Katzman,Alexandria Greifenberger,Elssa Toumeh,Sachinthya Lokuge,Tia Sternat,Ruiheng Yu,Lu Wang
类目: Applications (stat.AP); Artificial Intelligence (cs.AI)
备注: Accepted as an Oral Presentation at the IEEE Engineering in Medicine and Biology Conference (EMBC) 2026
Abstract:Depression treatment often requires switching medications due to inadequate response or adverse effects. Estimating individualized treatment effects in this setting is challenging because treatment assignment is confounded by patient characteristics, switching induces time-varying selection, and counterfactual outcomes are not observed in follow-up data. Using a proprietary longitudinal major depressive disorder (MDD) clinical trial dataset, we formulate a next-visit counterfactual prediction task to estimate Hamilton Depression Rating Scale (HAMD-17) total scores under alternative treatments. We benchmark 8 estimators, including meta-learners, residual-based methods, and tree-based approaches. Causal Forest (CF) demonstrates the most favorable and consistent performance across all criteria. Our analysis shows that symptom benefits concentrate in specific switch directions, with dose intensification being generally beneficial. Notably, we identify a counterintuitive exception where a lower-intensity regimen outperforms a higher-intensity alternative for specific patient subsets. While crude observational comparisons substantially overstate gains, confounding-adjusted estimates yield modest, actionable magnitudes. These findings provide prospectively testable candidates for clinical decision support in depression care.
机器学习
[LG-0] KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models
链接: https://arxiv.org/abs/2607.28608
作者: Sparsh Roy,Samuel Girmachew,Nishita Chavan
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注:
Abstract:Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pipelines have been proposed to catch this, but their components are rarely stress-tested, so it is unclear which parts of an audit can be trusted and under what conditions. We present KAISEN, a five-phase audit pipeline covering subgroup stratification, disparity measurement, mechanism diagnostics, post-hoc mitigation, and drift monitoring, evaluated to the point of failure on a synthetic benchmark of 16 disease tasks, 15 social-determinant axes from Healthy People 2030, and three prespecified intersections. Four findings follow. (i) Significance tracks each axis’s gap against its own minimum detectable effect: rank correlation between significance count and raw equalized-odds difference (EOD) across the 15 axes is rho = 0.56, rising to rho = 0.78 once EOD is standardized by that floor. (ii) Per-group threshold optimization reduces EOD in 48 of 48 held-out runs (paired delta = -0.285, 95% CI [-0.313, -0.252]), while group-wise Platt scaling – the better calibrator – behaves as a coin flip on EOD (19 of 48 runs improved, 95% CI [0.26, 0.55]) with mean effect near zero, so what an audit should report is the variance, not the average. (iii) The mechanism diagnostic classifies 144 of 144 controlled cases correctly but recovers none of 48 model-driven cases under proxy misspecification, with no signal that it failed. (iv) CUSUM failures and false alarms track cohort realization far more than disease: at the reference threshold, all 27 false alarms and 7 of 8 missed shifts come from different seeds (chi-squared p = 0.002), so a threshold tuned on one cohort fails to transfer. All results are synthetic with known ground truth and do not establish clinical validity. Code, artifacts, and scripts reproducing every number are released.
[LG-1] β-OPSD: Deriving with Policy Optimization Training with Self-Distillation
链接: https://arxiv.org/abs/2607.28582
作者: Jiawei Xu,Minghui Liu,Juzheng Zhang,Tom Goldstein,Furong Huang
类目: Machine Learning (cs.LG)
*备注:
Abstract:On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the \beta=1 member of a broader policy-optimization family, where \beta weights the KL penalty anchoring the student to a reference policy. This equivalence turns \beta from an implicit value fixed at one into a controllable regularization parameter, yielding a more general formulation that trades off proximity to a reference policy against privileged teacher guidance. We introduce \beta -OPSD and derive its optimal policy as a geometric interpolation between the reference policy and the privileged teacher. Directly optimizing this objective with reinforcement learning, however, would be costly and high-variance. Rather than optimize the RL objective directly, we turn its closed-form solution into a distillation target. Each value of \beta selects a target along the reference-to-teacher path, which we implement efficiently by mixing their token-level logits. In this way, inexpensive distillation approximates the solution of expensive policy optimization. Return-to-go credit assignment further aligns token updates with the sequence-level objective while retaining the simplicity of OPSD. Experiments on mathematical reasoning benchmarks show that \beta -OPSD consistently outperforms vanilla OPSD, improving optimization stability and downstream reasoning performance. Our results provide a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.
[LG-2] Same Graph Cross-Task Transfer in GNNs: Protocols and Predictors
链接: https://arxiv.org/abs/2607.28525
作者: Neelam Akula,Surbhi Kumar,Murat Kantarcioglu,Baris Coskunuzer
类目: Machine Learning (cs.LG); Social and Information Networks (cs.SI)
*备注: 17 pages, 2 figures
Abstract:Many real-world graphs support multiple predictive tasks over the same underlying structure, creating an opportunity to reuse supervision across node classification (NC) and link prediction (LP). However, existing evaluations often rely on incompatible splits, observed-graph assumptions, and negative sampling rules, making conclusions about same-graph cross-task transfer unreliable. We formalize same-graph NC-LP transfer and propose a leakage-free protocol that fixes node and edge splits, uses a shared message-passing graph that excludes evaluated edges, and employs fixed negatives for LP. Across three backbones (GCN, GraphSAGE, GPS), we find that transfer is strongly directional and predictable: NC \to LP is consistently beneficial on homophilic graphs, while LP \to NC is fragile and can even degrade accuracy under naive representation reuse. LP \to NC becomes reliably positive mainly in a structure-dominant regime where LP is easy but NC is unsaturated, suggesting that LP acts as structural pretraining. Finally, we introduce the CoTask Score (CTS) to summarize joint NC+LP utility when a shared encoder must serve both tasks, and show that simple dataset statistics, especially homophily, can guide mechanism choice and help avoid negative transfer.
[LG-3] he Role of Causality in Algorithmic Recourse
链接: https://arxiv.org/abs/2607.28497
作者: Srikanth Avasarala,Varun Gupta,Shahin Jabbari,Saber Salehkaleybar,Juba Ziani
类目: Machine Learning (cs.LG); Computers and Society (cs.CY); Computer Science and Game Theory (cs.GT)
*备注:
Abstract:Algorithmic recourse aims to provide individuals with actionable changes to improve their predicted outcomes in high-stakes classification settings, such as loan and mortgage applications. However, most existing approaches focus only on flipping a model’s prediction, without accounting for whether the recommended changes lead to genuine improvement in an individual’s true qualifications or merely enable strategic gaming of the classifier. Consequently, deployed recourse policies can induce behavioral responses that degrade predictive accuracy and become ineffective after model retraining. In this work, we formalize this failure mode through a causal performative framework for recourse. We model how recourse actions propagate through a structural causal model, capturing interactions among features as well as their effect on the true label. These causal responses induce a non-convex optimization problem, even under standard convex losses. We characterize conditions under which performatively stable solutions exist and can be efficiently computed via simple iterative dynamics. Our analysis reveals that recourse policies that ignore causal structure can induce large, misaligned behavioral responses, whereas causal recourse leads to stable equilibria that reduce incentives for gaming. Experiments on both semi-synthetic and real credit datasets demonstrate that our approach consistently outperforms standard empirical risk minimization while reducing the need for repeated model retraining to accommodate distribution shifts caused by strategic agent behavior. Subjects: Machine Learning (cs.LG); Computers and Society (cs.CY); Computer Science and Game Theory (cs.GT) Cite as: arXiv:2607.28497 [cs.LG] (or arXiv:2607.28497v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.28497 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-4] Cybersecurity Detection Classification with Reasoning -enabled Language Models
链接: https://arxiv.org/abs/2607.28460
作者: Amol Khanna,Manu Nandan,Cristian Viorel Popa,Joan Pujol-Roig,Diana Bolocan,Laura Vasilie,Alexandru Apostu,Chase Helwig,Mihaela Gaman,Michael Brautbar,Edward Raff,Chase Midler,Sven Krasser
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
*备注:
Abstract:A major issue in Security Operations Centers (SOCs) is alert fatigue, as the number of detections reported is more than staff can triage in a given day. Prior work prompts or fine-tunes large language models (LLMs) to emit a triage label directly, but does not train them to reason about whether a detection is a genuine threat. We train a chain-of-thought (CoT) reasoning-enabled triage classifier on real, human-labeled Windows endpoint detections by combining automated prompt optimization, self-training, and reinforcement learning with verifiable rewards. We find that CoT reasoning also degrades the label-token probabilities that automated triage relies on, so we separately train a calibrator that reads the full reasoning trace and estimates the probability that the verdict is correct. Our system reaches 82.6% test accuracy and, at the high-confidence operating point that governs automated triage, improves benign recall by 43.0% and malicious recall by 18.3% over a direct-label LLM classifier. We further show that the trained calibrator is necessary - an untrained confidence judge collapses high-confidence recall to zero - and that a finetuned 30B model significantly outperforms frontier general-purpose models, motivating targeted training over scale.
[LG-5] Graph Neural Multilevel Preconditioners for Iterative Solvers KDD2026
链接: https://arxiv.org/abs/2607.28456
作者: Zechen Zhang,Rui Peng Li,Yousef Saad
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注: Accepted at KDD 2026
Abstract:Solving large, sparse linear systems is a core task in scientific computing, and efficient iterative solvers rely critically on effective and robust preconditioning. While classical methods such as algebraic multigrid (AMG) are highly scalable, their robustness can degrade on indefinite or nonsymmetric systems where heuristics originally developed for elliptic PDEs are less reliable. Recently, Graph Neural Networks (GNNs) have emerged as data-driven preconditioners; yet, the practical impact of imposing an AMG-style hierarchy remains underexplored for general sparse matrices. In this work, we propose a Graph Neural Multilevel Preconditioner (GMP) that adopts an AMG hierarchy as a structural prior and learns smoothing, restriction, and interpolation operators in a unified framework. Our method targets general sparse systems and is instantiated as a drop-in preconditioner for standard Krylov solvers. On a benchmark of over 800 sparse matrices, we compare against classical AMG, single-level ILUT, and state-of-the-art GNN preconditioners, and characterize the regimes where multilevel graph neural preconditioning improves convergence or, conversely, introduces overhead relative to strong single-level baselines. These results highlight both the promise and the limitations of enforcing AMG-style multilevel structure in learned preconditioners for large-scale scientific simulations.
[LG-6] Oracle-Budgeted Molecular Optimization with Short-Term Graph Memory
链接: https://arxiv.org/abs/2607.28437
作者: Jiannan Yang,Veronika Thost,Xiang Ling,Tengfei Ma
类目: Machine Learning (cs.LG)
*备注: 12 pages, 5 figures
Abstract:Molecular optimization is commonly performed under a limited oracle budget, which makes deciding what to evaluate as important as deciding what to generate. We introduce short-term graph memory, a plug-in module that preserves the generator architecture and native update rule while learning from previously evaluated molecules to prioritize subsequent oracle queries. The module maintains an online graph neural surrogate that pre-screens each round’s candidate pool, so the fixed oracle budget is spent on molecules with higher predicted utility. Applied to a fragment-based generator on a standard molecular optimization benchmark, it improves the mean top-10 score at no extra oracle cost and never falls behind the base on any oracle; the gain extends to all four generators we tested at a tight budget of one thousand calls. We then analyze how surrogate-guided selection interacts with the exploration and exploitation behavior of different generators. Its benefit at larger budgets is consistent with two properties of the backbone: how broadly it searches, and how effectively its native search already exploits oracle feedback. We provide a simple way to spend a fixed oracle budget more selectively, and evidence on which generators benefit from it.
[LG-7] QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for Quantum Error Correction
链接: https://arxiv.org/abs/2607.28422
作者: Ran Miao,Rui Luo,Xiaohan Shan,Xiaoming Sun
类目: Machine Learning (cs.LG)
*备注: 11 pages, 6 figures, 6 tables
Abstract:Fault-tolerant quantum computing (FTQC) relies on quantum error correction to suppress physical errors and preserve logical information at scale. In practice, however, performance is constrained not only by physical noise but also by the latency of classical decoders processing rapidly generated syndrome data. This challenge is exacerbated by hardware noise that is strong, heterogeneous, and nonstationary, as well as by the simulation-to-hardware distribution shift that can substantially degrade fixed neural decoders. We present QAdapt, a noise-adaptive neural pre-decoding framework for surface-code quantum error correction. QAdapt captures local spatiotemporal correlations in syndrome data, sequentially adapts to evolving noise conditions while mitigating catastrophic forgetting, and forwards the residual syndrome to a conventional global decoder. Across 110 synthetic out-of-distribution noise configurations for rotated surface-code memory circuits, QAdapt consistently reduces the logical error rate relative to the neural pre-decoding baseline. On Google’s Willow benchmark data, without target-domain fine-tuning, it achieves reductions of up to 5.79 percent in logical error rate and 9.32 percent in backend decoding latency on the residual syndrome. These results demonstrate that QAdapt provides a practical and decoder-compatible approach to improving the robustness and backend decoding efficiency of quantum error correction under evolving hardware noise.
[LG-8] Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
链接: https://arxiv.org/abs/2607.28413
作者: Jianfeng Lu,Yinchen Luo
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Probability (math.PR); Statistics Theory (math.ST)
*备注:
Abstract:Let \mu(d x)\propto e^-U(x) d x on \R^d , where U is m -strongly convex and L -smooth, and denote by \kappa=L/m the condition number. We consider windowed thinning, an exact simulation method for the bouncy particle sampler and the coordinate Zigzag process. The method divides a trajectory into deterministic windows and uses a gradient evaluation at the beginning of each window to construct a tractable local envelope for the event rate. Combining this construction with quantitative mixing estimates and finite-time bounds on the expected numbers of bounces and flips yields query complexity guarantees from a Gaussian cold start. For total-variation error \varepsilon , the expected query counts are O(\kappa^1/2d,(d\log\kappa+\log\frac1\varepsilon)) gradient queries for the bouncy particle sampler and O(\kappa d^1/4(d\log\kappa+\log\frac1\varepsilon)) full-gradient equivalents for Zigzag, where d coordinate-partial queries count as one equivalent.
[LG-9] Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path Tested with Pre-Compiled Policy Trees
链接: https://arxiv.org/abs/2607.28399
作者: Zihan Dong,Rui Qian,Qishi Zhan,Dongshen Peng,Kaixin Li,Yu Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory Policy Trees (AAPT), which eliminates this delay without modifying the underlying model. During idle screen periods, the same frozen multimodal model constructs a bounded conditional policy tree with observable guards, pre-authorized actions, and branch-specific deadlines. The tree is sized to cover the model’s own decoding latency. When an event occurs, a lightweight observer matches change-gated frames to a prepared branch and immediately executes the corresponding action without generating new text. In paired trials with pre-registered endpoints and exact McNemar tests, AAPT improves the success rate from 0.50 to 0.79 within a contested decision window ( p=1.8\times10^-3 ), while producing no incorrect actions. Both open-loop and predict-and-replan baselines achieve zero success because they still decode during execution. A preparation-time sweep shows that the gain emerges where the latency-based tree-sizing rule predicts, and ablations reveal three key requirements: fast observer decoding, valid tree planning, and accurate branch routing. A pre-registered oracle probe rejects our initial hypothesis and instead points to branch routing as the causal bottleneck. We further reproduce the effect on an independent general-purpose multimodal model over 126 paired trials ( p=4.9\times10^-13 ). On an external benchmark, AAPT matches the overall performance of a reactive baseline, although the two methods exhibit complementary strengths. Together, these results suggest that AAPT performs best when candidate actions can be enumerated in advance, whereas reactive execution remains stronger when they cannot.
[LG-10] Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Averag e-Reward CMDPs
链接: https://arxiv.org/abs/2607.28390
作者: Ankur Naskar,Vaneet Aggarwal
类目: Machine Learning (cs.LG)
*备注:
Abstract:Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maximize long-term reward while satisfying long-term constraints. Although primal-dual actor-critic methods with linear critics are well understood, extending order-optimal convergence guarantees to neural critics in average-reward CMDPs has remained open. The main challenge is a fundamental bias-cost trade-off in neural critic estimation: under Neural Tangent Kernel (NTK) analysis, reducing critic bias substantially increases critic optimization cost, preventing order-optimal convergence in the primal-dual framework. We resolve this bottleneck by introducing a hierarchical Multilevel Monte Carlo (MLMC) neural critic that performs debiasing simultaneously across trajectory sampling and critic optimization. The resulting estimator attains the bias of a long critic optimization run with only logarithmic expected sample cost. Building on this estimator, we develop a primal-dual Natural Actor-Critic algorithm that achieves both an optimality gap and a constraint violation of order \tildeO(T^-1/2) . This establishes the first order-optimal convergence guarantees for infinite-horizon average-reward CMDPs with general policy parameterization and neural critics, while eliminating the need to know the underlying mixing time. Our results are novel even in the unconstrained setting.
[LG-11] LEDGERMIND: Provenance-Constrained Multimodal Agent ic Reasoning with a Structured Evidence Ledger
链接: https://arxiv.org/abs/2607.28374
作者: Enjun Du,Hange Zhou,Chenxu Du,Siyi Liu,Zirong Chen,Ziyu Zheng,Yongqi Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.
[LG-12] Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata
链接: https://arxiv.org/abs/2607.28338
作者: Michael Ben Ali,Imen Megdiche,André Péninou,Olivier Teste
类目: Machine Learning (cs.LG); Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (stat.ML)
*备注:
Abstract:Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation, communication cost, and computational efficiency. We formalize this as the CFL trilemma, according to which improving two of these dimensions comes at the expense of the third. A prominent paradigm relies on metadata (i.e., low-dimensional representations of client datasets shared with the server) to enable communication- and computation-efficient clustering. However, such approaches are not compatible with standard FL privacy-preserving mechanisms. To address this limitation, we propose FLAMECHE, which reformulates metadata-based CFL as a distributed Expectation-Maximization (EM) procedure, restricting server updates to additive operations while preserving efficiency. This design enables compatibility with practical secure FL schemes. We conducted extensive experiments on multiple datasets under various heterogeneous scenarios. Results show that FLAMECHE improves the effectiveness of client models. It enables encryption-compatible metadata-based clustering, enhancing its positioning within the CFL trilemma.
[LG-13] Measuring Distortion in the Empty Regions of Dimensionality Reduction Scatterplots with the Gap Index
链接: https://arxiv.org/abs/2607.28324
作者: Jaume Ros,Alessio Arleo,Fernando Paulovich
类目: Machine Learning (cs.LG)
*备注: 11 pages, 13 figures
Abstract:Quality metrics play a crucial role in the proper use of dimensionality reduction projections for visual analysis of high-dimensional data. They quantify the degree of distortion of a projection compared to the high-dimensional data and provide a reliable indication of how confident users can be in the structures they see in the resulting layouts. However, most popular metrics focus on capturing direct relationships between points (e.g., distances or neighborhoods) while neglecting distortions in empty areas of the layout, even though these often compose visually relevant features of a 2D layout. In this paper, we introduce the Gap Index (GI), a quality metric for 2D projections that captures visual distortion by measuring spatial distortion in empty areas of a projection. It does so by decomposing the space into empty triangles, which are then compared to their high-dimensional counterparts to compute the deformation. This per-triangle deformation can be aggregated into a single scalar value or overlaid on a projection to visualize regional distortion patterns. Results show that, contrary to popular quality metrics, the GI is sensitive to small structural deformations that have high visual impact. It is also fast to compute and interpretable.
[LG-14] Fully Inductive Cardinality Estimation ISWC2026
链接: https://arxiv.org/abs/2607.28311
作者: Tim Schwabe,Lukas Ketzer,Maribel Acosta
类目: Databases (cs.DB); Machine Learning (cs.LG)
*备注: Extended version of a paper accepted at ISWC 2026. 34 pages, 8 figures
Abstract:Query optimization of Basic Graph Patterns (BGP) SPARQL queries over Knowledge Graphs (KG) requires accurate cardinality estimation. Recently published learned estimators outperform statistics- and sampling-based approaches, but share a limitation preventing their adoption in real-world triplestores: they are transductive and require retraining when the underlying graph changes or when applied to new graphs. We present FICE (Fully Inductive Cardinality Estimation), the first learned cardinality estimator for BGP queries over KGs that generalizes to entirely unseen graphs (including unseen relations), without any retraining. FICE is a graph neural network (GNN) with two coupled components. First, an encoder GNN over a factor-graph view of the KG produces entity and relation embeddings. We prove that BGP cardinality is a local function of the 2-hop neighborhood around bound terms in this view, motivating the local message-passing encoder. A decoder GNN then composes these embeddings along the join topology of the query to predict log-cardinality. The encoder and decoder are trained jointly, making the embeddings specialized for cardinality estimation. FICE is trained using neighborhood sampling to scale to KGs with millions of triples, and decouples embedding generation from cardinality decoding to enable estimation latency below a millisecond. Compared to learned and non-learned baselines over 10 KGs, FICE reduces the overall median q-error from 13.54 (for the best competitor) to 5.34 and dominates all approaches in tail behavior.
[LG-15] Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
链接: https://arxiv.org/abs/2607.28308
作者: Huiyuan Tian,Bonan Xu,Shijian Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled 2\times2 factorial; frozen-route interventions and a controlled Top- k study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.
[LG-16] Semi-Supervised Learning for Molecular Graphs via Ensemble Consensus ICML
链接: https://arxiv.org/abs/2607.28304
作者: Rasmus Tirsgaard,Laurits Fredsgaard,Marisa Wodrich,Mikkel Jordahn,Mikkel N. Schmidt
类目: Machine Learning (cs.LG)
*备注: ICML
Abstract:Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.
[LG-17] HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLM s on HPC Tasks
链接: https://arxiv.org/abs/2607.28301
作者: Tiangang Li,Xiangbo Tian
类目: Machine Learning (cs.LG)
*备注:
Abstract:Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race detection and benchmark question answering. However, knowledge alone does not guarantee task-appropriate behavior: the same SFT model that correctly classifies 88.65% of C/C++ data race samples produces verbose, imprecise answers to factual queries, with 65.9% of MLPerf responses exceeding 40 characters. Reinforcement learning (RL) post-training addresses this gap by optimizing for task-specific rewards rather than token-level imitation. Yet HPC tasks exhibit extreme heterogeneity, with binary classification, factual QA, and semantic generation differing by 58x in answer length, spanning three distinct reward distributions, and showing widely varying SFT accuracy. This makes uniform-weight RL methods such as GRPO suboptimal. We propose HARGO, Heterogeneity-Aware Reward-Guided Optimization, which introduces per-response importance weighting via confidence-modulated advantage: computing a discrimination signal from group-level reward contrast and a confidence signal from reference model log-probabilities, then modulating the advantage before computing per-response weights, without requiring task-type labels. Across four HPC tasks and nine methods, HARGO achieves the best performance on all three primary metrics: WinRate 54.62%, Data Race F1 91.30%, and PLP Similarity 0.8558. Ablation confirms complementary contributions from both signals. HARGO establishes the best overall alignment quality among compared methods for heterogeneous HPC tasks.
[LG-18] opoFormer: Topology Meets Attention for Graph Learning
链接: https://arxiv.org/abs/2607.28259
作者: Md Joshem Uddin,Astrit Tola,Cuneyt Gurcan Akcora,Baris Coskunuzer
类目: Machine Learning (cs.LG); Algebraic Topology (math.AT)
*备注: 26 pages, 5 figures
Abstract:We introduce Topoformer, a lightweight and scalable framework for graph representation learning that encodes topological structure into attention-friendly sequences. At the core of our method is Topo-Scan, a novel module that decomposes a graph into a short, ordered sequence of topological tokens by slicing over node or edge filtrations. These sequences capture multi-scale structural patterns, from local motifs to global organization, and are processed by a Transformer to produce expressive graph-level embeddings. Unlike traditional persistent homology pipelines, Topo-Scan is parallelizable, avoids costly diagram computations, and integrates seamlessly with standard deep learning architectures. We provide theoretical guarantees on the stability of our topological encodings and demonstrate state-of-the-art performance across graph classification and molecular property prediction benchmarks. Our results show that Topoformer matches or exceeds strong GNN and topology-based baselines while offering predictable and efficient compute. This work opens a new path for parallelizable and unifying approaches to graph representation learning that integrate topological inductive biases into attention frameworks.
[LG-19] Secure Aggregation for Privacy-Preserving Federated Learning on Clinical EEG Data ESORICS2026
链接: https://arxiv.org/abs/2607.28191
作者: Pouya Rajabi,Mohsen Toorani
类目: Cryptography and Security (cs.CR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 27 pages, 6 figures, 7 tables. A version of this manuscript has been accepted for presentation at the International Workshop on Hot Topics at the Intersection of Distributed Machine Learning and Security (HotDiSec 2026), co-located with ESORICS 2026
Abstract:Federated learning enables multiple institutions to train shared models without exchanging raw clinical EEG data, but it does not fully prevent privacy leakage from individual model updates. This paper presents a privacy-preserving federated learning framework for clinical EEG data using masking-based secure aggregation as the core protection mechanism. The framework combines graph-based communication, threshold secret sharing, dropout-resilient aggregation, local update clipping, an optional Bloom filter-based privacy-preserving record-linkage initialization module, and auxiliary-notary-based verifiability. It supports both semi-honest and malicious aggregation settings and is implemented using the Flower federated learning framework. The secure-aggregation variants are evaluated in a simulated cross-silo healthcare setting using TUH EEG-derived data under different client configurations. Under the stated assumptions, the secure variants hide individual updates from the aggregation server. The results show that these variants remain compatible with federated model training, although malicious-setting safeguards and lightweight consistency-checking mechanisms introduce additional computation, communication, and round-duration overhead. The semi-honest variant provides the lowest overhead among the secure configurations, while malicious and auxiliary-notary variants offer stronger consistency, integrity, and lightweight verification support at higher cost.
[LG-20] Multi-channel Uplift Policy Learning
链接: https://arxiv.org/abs/2607.28182
作者: Changjian Liu,Tianyu Wang,Xiaoxuan Deng,WenTao Zhu,Yuwei Xu,Jungqi Jin,Yong Gao,Chuan Yu,Jian Xu,Bo Zheng
类目: Machine Learning (cs.LG)
*备注:
Abstract:E-commerce platforms must allocate fixed marketing budgets across multiple channels to maximize business utility. However, standard predict-then-optimize (PTO) paradigms fail in this compositional space due to observational confounding and severe extrapolation. We formulate this challenge as a simplex-constrained uplift decision problem and propose ReAlloc, a fast-slow causal framework. Specifically, an agile Orthogonal Teacher extracts unbiased local gradients from short-term logs, while an Explanation-Guided Student distills them into a structured marginal field over long-term horizons. This design enables support-aware, conservative decisions that capture cross-channel substitutions. Extensive simulations and large-scale online A/B tests on Taobao platform demonstrate that ReAlloc achieves simultaneous lifts in both pay order and income.
[LG-21] LM-GRASP: Instance-Specific Language Models for Combinatorial Construction via Online Imitation Learning
链接: https://arxiv.org/abs/2607.28135
作者: Mohand Mezmaz,Grégoire Danoy
类目: Machine Learning (cs.LG)
*备注:
Abstract:Machine learning for combinatorial optimization typically relies on neural constructors trained via reinforcement learning on large offline datasets for a fixed problem class-incurring high pretraining costs and generalizing poorly outside the training distribution. We propose an alternative: a metaheuristic framework that reformulates the randomized constructive phase of GRASP as an online imitation learning task, trained from scratch on each problem instance. A local search procedure acts as an expert oracle, while a decoder-only Transformer serves as the constructive policy. Unlike classical GRASP, which relies on static, myopic heuristic rules based on localized scalar costs, our approach is fully data-driven: the construction policy emerges from high-quality solutions discovered during the search itself, with no problem-specific feature engineering required. We instantiate this as LM-GRASP, a hybrid metaheuristic following an iterative learn-infer-improve cycle, training the policy online via behavioral cloning on a dynamic archive of elite trajectories-no external data or offline pretraining needed. The pipeline interfaces with the domain solely through the objective evaluator used by local search. Evaluated on the Taillard PFSP benchmark (ta51-ta60), the most discriminating block due to half its optima being unknown, LM-GRASP outperforms GPU-GRASP by 28.4 makespan units on average-comparable to the gain from GPU acceleration over sequential execution (27.2 units), though with overlapping standard deviations. This suggests instance-specific, online-trained language models are a promising, practical alternative to hand-engineered constructors, especially for landscapes resistant to classical greedy construction. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.28135 [cs.LG] (or arXiv:2607.28135v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.28135 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-22] From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE Inference
链接: https://arxiv.org/abs/2607.28097
作者: Tianyang Zhu
类目: Machine Learning (cs.LG)
*备注: 32 pages, 3 figures
Abstract:Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-Flash by freezing local MoE state and varying only aggregation semantics. Four schemes separate operand representation from accumulator precision. At one layer-5 fork, 720 A-mode orders yield 10 continuation basins; 720 B-mode orders form 360 exact structural classes and 11 basins. Under one Chinese prompt, the B classes split into 202 layoffs, 113 hiring, and 45 other continuations. Maximum-L-infinity B-branch selection separates 12, 24, and 36 of 50 prompts by 8, 16, and 32 tokens. Across 192 persistent trajectories per scheme, P32, A, and B change every native-reference route trajectory, while C preserves routes, token sequences, and texts. A separate 192-trajectory C check matches native MoE, post-mHC, next-router, and LM states bitwise. For one controlled B branch, exact post-mHC endpoint reconstruction reproduces the measured downstream trajectory. At the next decode boundary, exact FP64 reconstruction of the branch’s full persistent state yields agreement for 301 downstream post-mHC states, 301 persistent-state checkpoints, 301 routes, predictions, and text over seven steps, given the same naturally generated next input. These controls identify post-mHC as an intra-token boundary and full persistent state as a cross-token continuation boundary. Identical tokens need not imply identical autoregressive state: divergence can survive a token boundary and become visible later. These results make expert operand conversion, accumulator precision, and reduction order part of a numerical compatibility contract for sparse-MoE runtimes and hardware backends. They establish controlled causal possibility, not deployment incidence; C’s order invariance is limited to evaluated six-term states and schedules.
[LG-23] A Query-Efficient Stochastic Volume Rendering Framework for Time-Varying Implicit Neural Volumes
链接: https://arxiv.org/abs/2607.28047
作者: Alper Sahistan,Haichao Miao,Zhimin Li,Peer-Timo Bremer,Joshua A Levine,Valerio Pascucci
类目: Graphics (cs.GR); Machine Learning (cs.LG)
*备注:
Abstract:Time-varying implicit neural representations (INRs) provide a compact representation of scientific volumes and, for modalities such as dynamic X-ray computed tomography (CT), are often the only practical way to represent the data. However, interactive volume rendering of INRs is challenging, as cheap memory lookups are replaced by expensive neural inferences, hindering the performance. Therefore, conventional volume rendering methods such as ray marching with dense sampling are often impractical. While resampling, caching, and retraining can mitigate this cost, they compromise convenience and accuracy and become impractical for time-varying data. We tackle these challenges using a query-efficient stochastic volume rendering framework based on delta tracking. Our system employs a four-stage pipeline that exploits heterogeneous parallelism, using ray tracing cores for traversal and tensor cores for batched neural evaluation. Furthermore, we present strategies to reduce INR queries via ray budgeting and query pruning, thereby increasing per-frame performance. Using our renderer, many time-varying INRs can be rendered directly from their original representation. The system achieves ~30-40 FPS at 1024x1024 resolution on an RTX 4090 GPU and converges to high-fidelity images. Moreover, the system enables interactive temporal exploration of the continuous domain, with timestep updates taking approximately 1-2 ms.
[LG-24] ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
链接: https://arxiv.org/abs/2607.28037
作者: Xingjian Wu,Xuhang Zhu,Xingchen Liu,Junlin Liu,Jianing Wang,Linsen Guo,Xiaoyu Li,Xuezhi Cao,Xunliang Cai
类目: Machine Learning (cs.LG)
*备注:
Abstract:As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks. In this work, we present ClawTrack, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task Score) and how it achieves it (Process Score). ClawTrack comprises 320 tasks across 8 domains with 25+ deterministic mock services. A Process Grader scores each reasoning turn along four dimensions (goal alignment, efficiency, information utilization, and result verification), anchored by 12,541 task-specific rubric items. Evaluating 21 models over 16,000+ trials, we find that: (1) process scores effectively attribute success and failure to specific reasoning dimensions, filtering lucky passes invisible to outcome-only evaluation; (2) the four dimensions are complementary, with result verification as the systematic bottleneck; (3) the framework is robust to evaluator choice across different judge LLMs; and (4) process-based trajectory filtering yields consistent post-training improvements across model scales. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.28037 [cs.LG] (or arXiv:2607.28037v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.28037 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-25] Learning features from Newtons algorithm: a way to accelerate nonlinear parametrized PDE solvers
链接: https://arxiv.org/abs/2607.28036
作者: Rémy Vallot(CB, Michelin),Florian de Vuyst(BMBI),Thibault Dairay(CB, Michelin),Mathilde Mougeot(CB, ENSIIE, ENS Paris Saclay)
类目: Machine Learning (cs.LG); Analysis of PDEs (math.AP); Numerical Analysis (math.NA)
*备注:
Abstract:It is well known that Newton’s method converges faster when the initial guess is closer to a root of a system of nonlinear equations. In this paper, a two-stage Newton initial guess strategy is proposed by learning features from a parameter-space sampling and a database of precomputed solutions. The method uses discrete Newton trajectories to construct two complementary reduced spaces: a solution feature space, built from converged states, and a corrective search direction feature space, built from intermediate Newton increments. For an unseen parameter, a regression model is used to predict a surrogate solution approximation. Then, in a second step, a residual-minimizing correction is computed using a dedicated GMRES-based approach. The resulting state is then used as an initial guess for the high-fidelity Newton method, which completes convergence. The corrective step is computationally inexpensive since it only requires residual evaluations and the solution of a small least-squares problem. The methodology is weakly intrusive once the high-fidelity residual fields and a script-based programming interface are available. This strategy reduces the number of Newton iterations and decreases the overall CPU time. Numerical experiments on representative PDE problems show quantifiable speedups compared with standalone surrogate initialization. Significant speedups are observed. This generic approach can be applied to a broad class of large-scale nonlinear problems.
[LG-26] Enhancing Irregular Time Series Forecasting with Continuous-Time Modeling Framework
链接: https://arxiv.org/abs/2607.28035
作者: Tianen Shen,Zhengyu Li,Yutong Li,Xiangfei Qiu,Xingjian Wu,Bin Yang,Jilin Hu
类目: Machine Learning (cs.LG)
*备注: 13 pages, 5 figures
Abstract:Irregular multivariate time series are widely encountered in applications such as healthcare monitoring, human activity recognition, and environmental sensing. Their core challenges stem from asynchronous observations, non-uniform sampling intervals, and the fact that temporal patterns themselves carry critical dynamic information. Existing approaches either rely on discretization-based preprocessing (e.g., interpolation, imputation, or aggregation), which disrupts the underlying continuous-time semantics, or adopt continuous-time modeling via ODE-based frameworks, which typically require specialized architectures and incur substantial computational overhead due to numerical solvers. To address these limitations, we propose WrapFlow, a continuous-time modeling framework for irregular time series forecasting. On the input side, WrapFlow introduces Continuous-Time Tokenization, which directly encodes raw observation events and explicitly models long unobserved intervals via gap-aware tokens. The resulting continuous-time tokens are then processed by a standard Transformer backbone to capture long-range temporal dependencies. On the output side, we develop a simulation-free training paradigm for Residual Flow Matching, which learns conditional residual vector fields around base predictions while avoiding numerical-solver simulation and backpropagation during training. This design enables high-quality continuous forecasting using only a small number of fixed rollout steps at inference. Extensive experiments on multiple real-world datasets demonstrate that WrapFlow achieves state-of-the-art performance.
[LG-27] Contrastive Reinforced Policy Optimization via Privileged Self-Distillation
链接: https://arxiv.org/abs/2607.28026
作者: Xingjian Wu,Junlin Liu,Xingchen Liu,Xuhang Zhu,Jianing Wang,Linsen Guo,Xiaoyu Li,Xuezhi Cao,Xunliang Cai
类目: Machine Learning (cs.LG)
*备注:
Abstract:Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Self-Distillation (OPSD). While OPSD provides dense, logit-level supervision, it inherently suffers from exposure bias due to the privileged information of the self-teacher. In multi-turn agentic settings, this leads to reasoning route convergence and the loss of clear optimization directions. To tackle these challenges, we introduce Contrastive Reinforced Policy Optimization (CRPO), which reformulates agentic OPSD from a contrastive learning perspective. By leveraging predictive entropy to distinguish between positive positions (reflective exploration) and negative positions (exposure bias), CRPO conducts group-wise contrast to preserve reliable, fine-grained optimization signals. Extensive evaluations across 13 challenging reasoning and deep-search benchmarks demonstrate that CRPO consistently outperforms existing reinforcement learning and self-distillation baselines, significantly enhancing training stability and generalization in long-horizon interactions.
[LG-28] Building a User Foundation Model for the Open Web RECSYS’26
链接: https://arxiv.org/abs/2607.28019
作者: Solal Vernier,Ivan Can Arisoy,Merwan Barlier,Blaž Škrlj
类目: Machine Learning (cs.LG)
*备注: RecSys’26
Abstract:User foundation models have demonstrated strong results in e-commerce and social recommendation, but most industrial deployments assume environments where user identity is stable and persistent. Open-web real-time bidding (RTB) operates on a structurally different data distribution: user identity is fragmented and non-persistent across browsing sessions, and the availability of browsing history depends on user privacy choices. Consequently, a significant portion of traffic carries no historical data, and available records often consist of relatively short, disjointed sessions. As a result, historical signals in this domain are typically represented as aggregated counters and recency buckets, leaving the sequential structure unexploited. To address this limitation, we present a user foundation model that applies self-supervised learning on user browsing histories and show that the learned representation improves multiple downstream production tasks, demonstrating the viability of this approach on the open web. We pre-train a Transformer encoder with masked language modeling and a sequence-level contrastive objective, then fine-tune it on the click prediction task. We optimize the encoder’s pre-training pipeline with an LLM-in-the-loop search over a curated catalog of reviewable, code-level edits (lifters), instantiating the LLM-as-optimizer paradigm in an industrial setting. The same encoder representation yields +1.197% RIG on the production bid win-rate model and +1.354% RIG on the production CTR ranker; a 7-day live A/B test confirms +2.13% CTR, -1.13% eCPC (80% CI excluding zero on both metrics).
[LG-29] Its All Just Vectorization: einx a Universal Notation for Tensor Operations ICLR2026
链接: https://arxiv.org/abs/2607.27987
作者: Florian Fervers,Sebastian Bullinger,Christoph Bodensteiner,Michael Arens
类目: Machine Learning (cs.LG)
*备注: Published at ICLR 2026 (oral)
Abstract:Tensor operations represent a cornerstone of modern scientific computing. However, the Numpy-like notation adopted by predominant tensor frameworks is often difficult to read and write and prone to so-called shape errors, i.a., due to following inconsistent rules across a large, complex collection of operations. Alternatives like einsum and einops have gained popularity, but are inherently restricted to few operations and lack the generality required for a universal model of tensor programming. To derive a better paradigm, we revisit vectorization as a function for transforming tensor operations, and use it to both lift lower-order operations to higher-order operations, and conceptually decompose higher-order operations to lower-order operations and their vectorization. Building on the universal nature of vectorization, we introduce einx, a universal notation for tensor operations. It uses declarative, pointful expressions that are defined by analogy with loop notation and represent the vectorization of tensor operations. The notation reduces the large APIs of existing frameworks to a small set of elementary operations, applies consistent rules across all operations, and enables a clean, readable and writable representation in code. We provide an implementation of einx that is embedded in Python and integrates seamlessly with existing tensor frameworks: this https URL Comments: Published at ICLR 2026 (oral) Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.27987 [cs.LG] (or arXiv:2607.27987v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27987 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-30] Generalization Bounds on Optimal Control for Transformer Training and Wasserstein Distributional Robustness
链接: https://arxiv.org/abs/2607.27975
作者: Kağan Akman,Naci Saldi,Serdar Yüksel
类目: Machine Learning (cs.LG); Optimization and Control (math.OC)
*备注: 25 pages
Abstract:We derive finite-sample generalization bounds for Transformers trained with dynamic programming recursions. Building on the doubly lifted, measure-valued formulation of Transformer dynamics, we view data sets as probability laws on pairs of empirical input-output measures, allowing us to interpret the training problem as a finite-horizon Markovian control problem. We then analyze a quantized model, derived by quantizing the state, action, and measure-state spaces, and derive explicit finite-sample generalization bounds using concentration inequalities for empirical laws on finite metric spaces together with a Lipschitz stability estimate for the value function. These bounds are transferred to the base model at the cost of an explicit approximation error. Finally, we show that the same machinery yields a distributionally robust control formulation of the training problem, connecting Transformer generalization to Wasserstein distributionally robust optimization.
[LG-31] Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning ECML-PKDD2026
链接: https://arxiv.org/abs/2607.27968
作者: Efstratios Zaradoukas,Davide Gabrielli,Bardh Prenkaj,Gjergji Kasneci
类目: Machine Learning (cs.LG)
*备注: Accepted to WIPE-OUT 2 @ ECML-PKDD 2026
Abstract:Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR and the EU AI Act. Recent work has reformulated unlearning as a Reinforcement Learning with Verifiable Rewards (RLVR) problem, where models are optimized against verifiable rewards computed directly from their outputs. However, existing methods rely on sparse binary rewards that provide minimal learning signal, indicating only whether forbidden content was avoided, and limiting convergence speed. In this paper, we study how reward design affects unlearning efficiency within the Reinforcement Unlearning (RUL) framework. We introduce a principled reward decomposition framework that decouples verifiability from sparsity, and propose two new reward functions: an exponential reward that provides graded penalties based on the count of forbidden-concept occurrences, and a PageRank inspired reward that weights penalties by semantic importance. We conduct experiments on the Real World Knowledge Unlearning (RWKU) benchmark, demonstrating that both rewards consistently outperform the binary setting, while reaching similar forgetting performance up to 3\times faster and preserving general model utility. Our results show that reward design is a key driver of unlearning efficiency offering a practical path toward scalable and efficient machine unlearning.
[LG-32] What Makes Graph Unified? Principles and Generative Sliding-Window Transformer for Graph Foundation Models
链接: https://arxiv.org/abs/2607.27966
作者: Dongxiao He,Siqi Liu,Jitao Zhao,Yawen Li,Yi Wang,Di Jin
类目: Machine Learning (cs.LG)
*备注:
Abstract:Graph Foundation Models (GFMs) have recently emerged as a promising paradigm for general-purpose graph learning, aiming to learn reusable knowledge that generalizes across diverse graph domains and downstream tasks, reducing the need for specific model development. Achieving this goal requires reconciling the substantial heterogeneity in node features, graph structures, and semantic information across domains. Among them, heterogeneous node features constitute a fundamental input-level barrier, as their dimensionality and semantics vary substantially across datasets. Existing studies typically project or map heterogeneous node features into a fixed-dimensional space, often implicitly equating dimensional uniformity with effective feature unification. Yet dimensional consistency alone does not ensure that the unified features preserve informative semantics and capture transferable patterns that can support cross-domain knowledge transfer. To bridge this conceptual gap, we distill four desiderata for cross-domain graph feature unification: formal uniformity, cross-domain transferability, information preservation, and backbone compatibility. Guided by these principles, we propose SliGFM, a graph foundation model built upon topology-aware sliding-window feature encoding and generative reconstruction. SliGFM orders feature dimensions by topological smoothness and scans the reordered features with a shared sliding-window feature encoder, transforming heterogeneous features into a common space of ordered fixed-dimensional feature tokens. This formulation enables a smoothness-aware transformer to capture transferable relational patterns among feature tokens within each node, while the generative reconstruction objective encourages preservation of the original feature information.
[LG-33] AutoPref: Automatic Discovery of Task-Specific Preference Objectives for Neural Combinatorial Optimization
链接: https://arxiv.org/abs/2607.27953
作者: Shengda Gu,Kai Li,Xinyi Ke,Haobo Fu,Yifan Zhang,Jian Cheng
类目: Machine Learning (cs.LG)
*备注: 8pages, 2figures
Abstract:Combinatorial optimization problems (COPs) underpin many real-world decisions, but their exponentially large search spaces make high-quality solutions costly to obtain. Neural combinatorial optimization (NCO) learns fast construction policies, typically with reinforcement learning (RL), while preference-based NCO improves sample efficiency by learning from relative solution quality. However, existing preference objectives combine two distinct design choices in manually specified, one-size-fits-all formulations: what learning signal to extract from each solution pair and how to weight each pair relative to the sampled set. We present AutoPref, the first LLM-guided framework for automated preference-objective discovery in NCO. AutoPref factorizes the objective into a pairwise loss program, which defines the learning signal, and a set-aware weighting program, which determines each pair’s relative contribution. Their composition forms a unified programmatic objective space containing existing preference objectives as special cases. To make its search tractable, we introduce a staged conditional search strategy with behavioral gates that filter inadmissible programs before short-horizon training and evaluation. Across TSP, CVRP, FFSP, and JSSP, AutoPref consistently outperforms strong hand-designed baselines across problem scales, demonstrating the benefits and scalability of automated objective discovery for NCO.
[LG-34] Harnessing the Potential of Optimizing Data Mixtures via Bayesian Domain Reweighting
链接: https://arxiv.org/abs/2607.27928
作者: Xiang Yuan,Kaiqing Lei,Zhenyu Jin,Jun Shu,Deyu Meng,Zongben Xu
类目: Machine Learning (cs.LG)
*备注:
Abstract:The performance of Large Language Models (LLMs) is fundamentally influenced by the distributional composition of multi-domain pre-training data. While manual heuristics were prevalent in early models, they increasingly fail to capture the intricate synergies between domains as data complexity grows. To overcome the issue, a dominant approach seeks to fit a proxy function mapping between domain weights and their corresponding validation losses, and then find the optimal domain weights to minimize validation losses. These methods rely on strong structural assumptions, such as rank invariance or scaling laws, which are often violated, resulting in non-negligible estimation bias. A promising approach is to directly optimize the weighting scheme from data. However, it suffers from unstable optimization trajectory and prohibitive computational overhead, limiting its potential to search better domain weights configurations. This paper presents a Bayesian domain weighting method to infer the weights from a Dirichlet distribution via introducing Gamma prior information learned from observations. Experimental results demonstrate that proposed method could achieve stable and efficient domain weights learning, and identifies optimal mixtures while consuming substantially less data than search-based function-fitting methods, revitalizing optimization-based domain weighting for large-scale applications.
[LG-35] Exact Action Values Are Not Enough: Rollout-Verified Reinforcement Fine-Tuning of a Reasoning Model for Multi-Zone VAV Control
链接: https://arxiv.org/abs/2607.27914
作者: Takumi Shioda,Kohei Terashima,Tatsuo Nagai
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注: 34 pages, 14 figures
Abstract:Multi-zone variable-air-volume control must balance thermal comfort, indoor air quality, and electricity use across several continuous actuators. Model predictive control and reinforcement learning are widely studied, but deployment typically requires building-specific modeling or training, limiting scalability. We first test whether a frontier reasoning model (an LLM trained to use additional inference-time computation) can achieve competitive VAV control from text without building-specific training. With that capability established, we then test whether TD3-guided reinforcement fine-tuning (RFT) can transfer control knowledge into a locally deployable open-weight model. Five controllers are evaluated over three summer days in a physics-based four-zone emulator. Relative to a Guideline 36-based baseline, TD3 reduced HVAC electricity by 4.5% while improving temperature and CO _2 compliance. Without building-specific training, GPT-5 achieved the largest reduction (6.2%) but reduced the ventilation margin. For RFT, deterministic rollouts restore a saved state, apply one candidate, and follow TD3 to score each action. Auditing a learned critic against these rollouts exposed a failure hidden by its near-perfect across-time correlation ( r=0.9998 ): within-state ranking was unreliable; the critic selected the rollout-best candidate in only 5 of 10 states. Even with the rollout verifier, 200 RFT steps produced no sustained improvement in sampled-action return; the open-weight controller used more electricity than the baseline before and after training, and its five-minute predictions remained worse than persistence. GPT-5 predicted transitions far better. Exact rollout scores rank sampled actions but reveal neither next-state effects nor an improvement direction. The unchanged transition errors motivate transition-focused supervised fine-tuning before value-based RFT.
[LG-36] S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring MICCAI2026
链接: https://arxiv.org/abs/2607.27913
作者: Glenn Anta Bucagu,Thorir Mar Ingolfsson,Yawei Li,Luca Benini
类目: Machine Learning (cs.LG)
*备注: This is the pre-rebuttal version of a paper accepted at MICCAI 2026. The camera-ready version will be posted following the embargo
Abstract:Foundation models offer a promising paradigm for Electroencephalography (EEG) analysis, leveraging generalizable representations from vast unlabeled datasets. Yet, Transformer-based architectures face a critical bottleneck: global attention mechanisms couple the attention memory state to the signal duration, causing memory overflow during continuous monitoring. To address this, we introduce S-CEReBrO (Streaming CEReBrO), an evolution of the CEReBrO architecture designed for continuous monitoring. Our novel Windowed Alternating Attention mechanism factorizes attention computation into fixed-size spatiotemporal windows, guaranteeing constant KV cache memory as only the active window requires resident attention maps. Empirical scaling analysis confirms that windowed alternating attention can process signals 100X longer than full self-attention and 3X longer than low-rank linear attention. Compared to low-rank linear attention on long contexts, windowed alternating attention requires 55% of the memory while increasing inference throughput by 2.1X. Pre-trained on 25,000 hours of recordings from 12,000 subjects, S-CEReBrO achieves state-of-the-art performance on 7 of 11 downstream tasks, with up to 60% fewer parameters. This work represents a significant step toward the realization of efficient, generalizable, and continuous EEG monitoring. An accompanying code repository is available.
[LG-37] Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
链接: https://arxiv.org/abs/2607.27909
作者: Dmitrii Gavrilev,Ilya Borovik,Vladimir Viro
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: Accepted at ISMIR 2026
Abstract:Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances. In generative applications, the wide variety of expressive attributes makes it difficult to aggregate them into a single scalar metric for model selection. In this work, we reexamine attribute-scoped metrics and explore the perceptual properties of contextual embeddings from self-supervised symbolic music models, Aria and CLaMP3. Results from our listening study indicate that these models can be used as perceptual proxies, showing agreement with per-sample human ratings on par with traditional metrics. To measure conditional distributional similarity, we adapt Kernel Audio Distance to the symbolic music domain. Unlike Pearson correlation and reconstruction error, kernel-based methods on contextual embeddings do not require note alignment and are sensitive to contextual perturbations. To facilitate reproducibility, we release Pereval, an open-source library that integrates performance evaluation utilities, including both attribute-scoped and deep feature metrics.
[LG-38] Contrastive Concept Importance: Explaining Pairwise Class Decisions Through Automatically Extracted Concept Representations
链接: https://arxiv.org/abs/2607.27904
作者: Roel Visser,Isaac Roberts,Barbara Hammer
类目: Machine Learning (cs.LG)
*备注:
Abstract:Concept-based explanations are a prevalent way to explain the decisions of complex black-box methods through semantically meaningful, human-interpretable concepts. To attribute the contribution of such concepts to a model’s decisions, feature attribution methods are used to quantify how strongly each concept contributes to a model output. These attributions are typically computed for a single output class and therefore answer a non-contrastive “why P?” question. In many situations, however, such as cases of misclassification, class confusion, and low-margin predictions, the more natural question to ask is “why P rather than Q?”. We introduce contrastive concept importance (CCI), which attributes the logit margin between a target class and a contrast, or foil, class to concepts in an automatically extracted visual concept basis. The resulting scores are signed, indicating whether a concept supports the target over the foil or the foil over the target, and can be decomposed into target-logit and foil-logit effects. This makes it possible to distinguish globally important concepts from concepts that specifically influence a class-pair distinction, including whether their effect is shared, one-sided, or directly contrastive. We evaluate the method on ImageNet class pairs using CRAFT-style concept bases, insertion and deletion curves, logit-wise decomposition analysis, and semantic class hierarchy. The results show that contrastive concept importance reveals class-pair-specific model behavior that is not captured by ordinary concept importance alone, and that highly contrastive concepts can be evaluated against semantic superclass structure to assess whether they affect fine-grained distinctions rather than broad category evidence. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.27904 [cs.LG] (or arXiv:2607.27904v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27904 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-39] Safety-Gated Agent ic Supervisory Control on a Coupled Distillation Benchmark: Regime Map Auditable Gate and Co-Design Findings
链接: https://arxiv.org/abs/2607.27849
作者: Christian Rosenthal
类目: ystems and Control (eess.SY); Machine Learning (cs.LG)
*备注: 31 pages, 8 figures. Code and data: this https URL . Sole author; independent research
Abstract:An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, and an admit/block decision before the regulatory layer moves. This paper puts that check in a rule-based forked-twin counterfactual gate (nine pinned constraints) and leaves the regulatory layer unchanged. On Skogestad’s Column A the ladder is PID-only (C0), linear MPC (C1), ungated agent (C2), and gated agent (C3) under one contract: identical level closure (M_D, M_B), scenarios, and seeds; C2/C3 share the linear-MPC backend. The split is not subtle. Off-nominal target acquisition: the agent beats Pareto-tuned linear MPC in the strong band (C2/C1 IAE ratio 0.361 at the upper CI). Disturbance rejection on the same 16-point grid inverts by 16.03 at the upper CI (10.18 at the point estimate), where an ungated LLM supervisor does not belong. The gate compresses a specification-abandonment attractor into a bounded offset (d approx. -1.4; P95 cell IAE 11.5 to 0.77). A one-line prompt fix removes the attractor at source (6/10 to 0/10; sensitivity only, not a new headline). In a 250-cell statistical pass, 534 of 590 gate interventions are spec-on-bound geometry: the operating specification sits on a safety limit, so a well-behaved OP becomes inoperable while misbehaving ones are only contained; 318 blocks still correct actively harmful proposals. Headlines are single-column and model-conditional on DeepSeek-V4-Flash. A second-family sweep (NVIDIA Nemotron-3-Super) keeps the disturbance-rejection fails band and plant-side failure geography; magnitudes and protocol operability stay model-conditional, and Super target-acquisition strong cells are survivors only (not confirmation). Transfer means twin, constraint envelope, and setpoint interface, not a second plant class measured here. Comments: 31 pages, 8 figures. Code and data: this https URL. Sole author; independent research Subjects: Systems and Control (eess.SY); Machine Learning (cs.LG) MSC classes: 93C95, 68T05 ACMclasses: I.2.6; I.2.8; J.2 Cite as: arXiv:2607.27849 [eess.SY] (or arXiv:2607.27849v1 [eess.SY] for this version) https://doi.org/10.48550/arXiv.2607.27849 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-40] Nanoparticle Networks for Neuromorphic Computing
链接: https://arxiv.org/abs/2607.27844
作者: Jonas Mensing,Wilfred G. van der Wiel,Andreas Heuer
类目: Emerging Technologies (cs.ET); Mesoscale and Nanoscale Physics (cond-mat.mes-hall); Hardware Architecture (cs.AR); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:
Abstract:Physical computing leverages complex dynamical systems for energy-efficient data processing. In this work, we present a neuromorphic architecture based on metallic nanoparticles interconnected by molecular junctions on a \textSiO_2 /Si substrate. We demonstrate that surrounding static control electrodes transform this nanoparticle network from a passive reservoir into a tunable nonlinear dynamical system. By analyzing how these electrodes route simple one-dimensional voltage inputs into multidimensional signal responses, we establish three core design rules to maximize computational performance. First, operating near the system’s cutoff frequency achieves an optimal balance between nonlinear charge tunneling and linear capacitive memory. Second, tuning the underlying \textSiO_2 thickness sets the electrostatic screening length and dictates the memory type. Thick oxide layers reduce the screening length, causing networks larger than this length to transition into a persistent, non-volatile-like regime. Conversely, networks smaller than the screening length exhibit only fading memory. Third, introducing structural disorder via heterogeneous molecular junctions overcomes inherent limits on expressivity. While a network’s computational expressivity scales with its physical size, it is ultimately capped by the screening length. Breaking internal spatial symmetries with localized disorder bypasses this saturation, allowing control voltages to independently manipulate specific signal amplitudes and phases, universally maximizing performance for dynamic neuromorphic applications.
[LG-41] Learning-Augmented and Randomized Algorithms for Line Aggregation with Delays
链接: https://arxiv.org/abs/2607.27807
作者: Tianhang Lu,Runtian Ren,Shengcai Liu,Ke Tang
类目: Machine Learning (cs.LG); Computational Complexity (cs.CC)
*备注:
Abstract:This paper studies learning-augmented and randomized online aggregation with delays on a line metric. We consider advice given as online suggested service lengths, and evaluate the algorithms in terms of robustness and consistency. For each \lambda \in (0,1] , we first propose a deterministic learning-augmented \textscBalance algorithm that is (4/\lambda+1/\lambda^2) -robust and (4+\lambda) -consistent. We also propose a randomized algorithm for the problem in the classical adversarial model, which is (e+1) -competitive against an oblivious adversary, improving over the deterministic 5 -competitive \textscBalance benchmark~\citebienkowski2013chain. Notably, this competitive ratio is even lower than the lower bound of 4 for deterministic online algorithms. Moreover, we establish a lower bound of e on the competitive ratio of randomized online algorithms, improving the previous lower bound of e/(e-1) . Besides, we combine the two ideas and obtain a randomized learning-augmented algorithm that is (e/\lambda+1/\lambda^2) -robust and (e+\lambda) -consistent. Finally, we conduct numerical experiments to complement our theoretical analysis and evaluate the empirical performance of our algorithms.
[LG-42] Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence Tabular and LLM Approaches ECML KDD2026
链接: https://arxiv.org/abs/2607.27797
作者: Lennart Fertig,Lukas Kirchdorfer,Tobias Sesterhenn
类目: Machine Learning (cs.LG)
*备注: Accepted at ECML PKDD 2026 Workshops
Abstract:Predictive process monitoring (PPM) leverages event logs to forecast the future of running process instances, for instance, predicting the next activity, the remaining time until case completion, or the time to the next event. While PPM research in recent years has been dominated by deep sequence models trained from scratch, such as Long Short-Term Memory (LSTM) models, foundation-model approaches—particularly large language models (LLMs)—are increasingly explored for PPM. At the same time, tabular foundation models with in-context learning capabilities offer a promising alternative but have not yet been systematically benchmarked for PPM. Thus, it remains unclear whether classical sequence-based models remain competitive in this evolving landscape. This paper compares the three modeling paradigms both conceptually and empirically through a controlled benchmark across multiple datasets and prediction tasks. The results show that sequence models consistently perform best for next activity prediction, whereas tabular foundation models are competitive on temporal tasks, with LLMs usually lagging behind despite higher cost.
[LG-43] RIPPLE: Generating Multi-Channel Phase Not Recovering It
链接: https://arxiv.org/abs/2607.27775
作者: Jaehyuk Lee,Yeajin Lee,Dayeon Shin,Donghun Lee
类目: Machine Learning (cs.LG); Sound (cs.SD)
*备注:
Abstract:Generative models synthesize magnitude spectra with high fidelity, while phase is delegated to a recovery module—Griffin–Lim, a vocoder, or a latent decoder—applied independently to each channel. For multi-channel waveforms this delegation is costly: the physical content of spatial audio and three-component seismograms lives in the phase relationships between channels, precisely what channel-independent recovery cannot produce. The cost is also invisible, since the magnitude-based metrics common to both fields barely move when inter-channel phase coherence collapses—so a pipeline can discard the physical information in its output while still scoring well. We argue that phase should be generated, not recovered, and present RIPPLE (Rectified Inter-channel Phase with Prior-based LEarning), which reinterprets Griffin–Lim as a phase prior rather than a final estimator: initialized from the source phase, this prior carries the inter-channel structure to be preserved, and a rectified flow refines it toward the target under an explicit inter-channel phase loss. Tested on first-order ambisonics environment transfer and seismic cross-station translation—two physically unrelated domains—RIPPLE outperforms recovery-based pipelines on the coherence metrics that downstream analyses consume. The seismic case is decisive: across architecturally distinct generators, per-channel recovery leaves S-wave polarization error near the 57.3^\circ random expectation, whereas learned phase reduces it to 33.8^\circ .
[LG-44] Beyond the Best Teacher: Expanding and Compressing the Reasoning Solution Manifold
链接: https://arxiv.org/abs/2607.27770
作者: Songshuo Lu,Zhi Chen,Yaohua Tang
类目: Machine Learning (cs.LG)
*备注:
Abstract:A single reinforcement-learning run can produce a strong reasoner yet an incomplete teacher: it often amplifies only a subset of the valid solution modes. We argue that reinforcement learning (RL)-trained policies should therefore be viewed as local probes of a multi-basin reasoning solution manifold, rather than as globally reliable supervisors. Based on this view, we propose an expand-then-compress framework that couples teacher construction with multi-teacher policy distillation. In the expansion stage, Residual Group Relative Policy Optimization (RGRPO) trains a sequence of teachers from a common initialization and redirects each later round toward examples not yet covered by the accumulated teacher union. In the compression stage, reliability-gated Teacher-Union On-policy Distillation (TU-OPD) lets the student learn from its own response prefixes. For each example, only reliable teachers contribute, and their sampled-token OPD losses are weighted by their per-example quality. We further introduce Consensus-Residual Decomposition, which preserves a winner teacher’s excess token preferences over its reliable peers, preventing specialist behavior from being suppressed during teacher aggregation. Experiments on mathematical reasoning, code generation, and instruction following show that the resulting Qwen3-1.7B student consistently outperforms the strongest individual teacher across all three domains, yielding relative improvements of 2.0%, 8.3%, and 6.9%, respectively, while retaining single-model inference. These results establish a simple but powerful principle: stronger students can be obtained not by selecting a single better teacher, but by deliberately constructing and compressing a complementary teacher union.
[LG-45] Improving the Robustness/Accuracy Tradeoff Against Adversarial Attacks Using Information Bottleneck Distillation Through Dual Teachers
链接: https://arxiv.org/abs/2607.27737
作者: Vincent Ryusuke Takahashi,Yoshinari Takeishi,Jun’ichi Takeuchi,Kave Salamatian
类目: Machine Learning (cs.LG)
*备注: 9 pages, 5 figures
Abstract:Deep neural networks (DNNs) have achieved remarkable success in classical machine learning problems. However, they are known to be vulnerable to adversarial attacks. Countermeasures proposed in the literature, notably Information Bottleneck Distillation (IBD) introduced by Kuang et al., degrade the classification accuracy on clean inputs while improving the robustness to adversarial inputs. In this work, we extend the IBD framework by introducing an extra teacher model (clean teacher) trained with only clean inputs, into the distillation process from a robust teacher model trained by adversarial training. The features of both clean and robust teachers are transferred to the student through a cross-layer attention matrix. Experimental results on the CIFAR-10 and CIFAR-100 datasets show that the proposed method improves classification accuracy on clean samples compared to the original IBD, while maintaining similar accuracy on adversarial samples. Furthermore, our methods are competitive with state-of-the-art approaches, including the recent dual-teacher distillation framework B-MTARD, particularly in terms of the harmonic mean between clean and robust accuracy. We also analyze the impact of different training settings that have different influences on the attention module.
[LG-46] VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers Validated on Ancient DNA Reconstruction
链接: https://arxiv.org/abs/2607.27712
作者: Angshuman Chakravertty,Rahul Maheshwari
类目: Machine Learning (cs.LG)
*备注: 18 pages, 9 figures, 6 tables
Abstract:Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is position-agnostic. When the degradation process is characterised and concentrated at predictable positions, this assumption fails: at peak damage sites the model can underperform a frequency-matched random predictor. We introduce VESTIGE, a parameter-free, drop-in replacement for the standard MLM collator that aligns the masking distribution with an empirically measured per-position corruption profile. We apply it to ancient DNA (aDNA) reconstruction, where cytosine deamination produces a position-dependent C-to-T / G-to-A gradient quantified per-position by mapDamage2. Rescaling so the mean C/G masking rate equals 15% - identical to standard MLM - isolates spatial redistribution as the sole variable, with model, data, seed, and hyperparameters held fixed across both DNABERT-2 runs on a mammoth CDS corpus (two specimens, seven genes). Across six terminal-zone widths and 626 paired windows, VESTIGE leads standard MLM at every width (Delta = +4.18 to +10.35 pp, all p 10^-8), cuts validation cross-entropy by 13% (3.274 vs. 3.757), and yields ESMFold reconstructions with TM-score 0.95 across all six reconstructions (three genes) even under damage amplified 10-30x beyond authentic PMD rates. A 1D CNN biosecurity classifier returns AUC = 0.935 and clears 98.2% of reconstructed windows, the 1.76% remainder attributable to reference-genome features, not reconstruction artefacts. The principle is domain-agnostic: any measurable position- or context-specific corruption profile - FFPE, bisulfite, metagenomic, or nanopore - substitutes directly for the PMD array, making VESTIGE a knowledge-guided training routine for intelligent systems operating on degraded or noisy sequence inputs.
[LG-47] NMINE: Normalized Mutual Information Neural Estimation
链接: https://arxiv.org/abs/2607.27710
作者: Petra Eerikinharju,Marko Tuononen,Ville Hautamäki
类目: Machine Learning (cs.LG)
*备注:
Abstract:Mutual information is a general measure of statistical dependence that captures both linear and nonlinear relationships between random variables. For continuous and multidimensional variables For continuous multidimensional variables, mutual information must be estimated from samples. Because mutual information is unbounded, its values are not directly comparable across datasets, dimensions, or applications. Normalized mutual information addresses this limitation by converting mutual information into a normalized dependency score. Recent work has demonstrated the practical value of normalized mutual information in applications such as molecular dynamics arXiv:2405.04980 and interpretable machine learning arXiv:2409.16768, but existing estimators remain sensitive to dimensionality and numerical stability arXiv:2410.07642. In this paper, we propose a fully neural normalized mutual information estimator for continuous variables. The proposed approach combines a MINE-based neural mutual information estimator arXiv:1801.04062 with MI-NEE-inspired neural marginal entropy estimators arXiv:1905.12957. Mutual information is estimated using the Donsker–Varadhan representation, while marginal entropies are estimated by learning the divergence between each marginal distribution and a uniform reference distribution, from which entropy is recovered. The resulting estimator provides a neural alternative to k-nearest-neighbor-based normalized mutual information estimation arXiv:2405.04980. Experiments on Gaussian data from one to eight dimensions show that the proposed estimator improves accuracy over a KSG-based normalized mutual information baseline. These results indicate that neural estimation is a promising direction for normalized dependency measurement in continuous multidimensional settings. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.27710 [cs.LG] (or arXiv:2607.27710v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27710 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-48] LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
链接: https://arxiv.org/abs/2607.27704
作者: Sangjin Kim,Yuseon Choi,Jungjun Oh,Byeongcheol Kim,Hoi-Jun Yoo
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: 13 pages, journal version. Published in IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS), vol. 15, no. 2, pp. 231-243, 2025, DOI: https://doi.org/10.1109/JETCAS.2025.3558300
Abstract:As large language models (LLMs) continue to demonstrate exceptional capabilities across various domains, the challenge of achieving energy-efficient and accurate inference becomes increasingly critical. This work presents LightRot, a lightweight rotation scheme and dedicated hardware accelerator designed for low-bit LLM inference. The proposed architecture integrates Grouped Local Rotation (GLR) and Outlier Direction Aligning (ODA) algorithms with a hierarchical Fast Hadamard Transform (FHT)-based rotation unit to address key challenges in low-bit quantization, including the energy overhead of rotation operations. The proposed accelerator, implemented in a 28nm CMOS process, achieves a peak energy efficiency of 27.4 TOPS/W for 4-bit inference, surpassing prior state-of-the-art designs. Unlike conventional approaches that rely on higher-precision inference or evaluate on basic language modeling tasks like GPT-2, LightRot is optimized for advanced models such as LLaMA2-13B and LLaMA3-8B. Its performance is further validated on MT-Bench, demonstrating robust applicability to real-world conversational scenarios and redefining benchmarks for chat-based AI systems. By synergizing algorithmic innovations and hardware efficiency, this work sets a new paradigm for scalable, low-bit LLM inference, paving the way for sustainable AI advancements.
[LG-49] GyRot: Leverag ing Hidden Synergy between Rotation and Fine-grained Group Quantization for Low-bit LLM Inference HPCA HPCA68181
链接: https://arxiv.org/abs/2607.27694
作者: Sangjin Kim,Yuseon Choi,Byeongcheol Kim,Jungjun Oh,Hoi-jun Yoo
类目: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
*备注: 15 pages, 12 figures. Published in 2026 IEEE International Symposium on High-Performance Computer Architecture (HPCA), Sydney, Australia, pp. 1-15, DOI: https://doi.org/10.1109/HPCA68181.2026.11408453
Abstract:Low-bit quantization is essential for efficient LLM inference, and both rotation and fine-grained group quantization have shown individual promise. However, their combination often leads to accuracy degradation or hardware overhead due to a mismatch between the global nature of rotation and the localized behavior of group scaling. We propose GyRot, a quantization framework and hardware accelerator that bridges this gap through algorithm-hardware co-design. GyRot introduces Coarse Rotation, Fine Grouping (CoRFiG) and Harmonic-Aligned Permutation (HAP) to enable cooperative integration of rotation and group quantization, enhancing quantizability while relaxing scaling factor precision. To further reduce hardware cost, we reformulate asymmetric quantization and introduce a zero-point rounding strategy that enables fully integer dequantization. Implemented on an INT4-based tensor PE architecture, GyRot achieves state-of-the-art 4-bit accuracy across LLaMA-family models, while delivering up to 3.4x speedup and 3.6x energy efficiency over baseline LLM accelerators. These results validate GyRot’s practical effectiveness for scalable and energy-efficient LLM deployment.
[LG-50] Event-Structured Physics-Informed Neural Networks for Differentiable Critical Clearing Boundaries
链接: https://arxiv.org/abs/2607.27681
作者: Baoli Hao,Chenxi Hu,Ming Zhong,Ren Wang
类目: Machine Learning (cs.LG); Systems and Control (eess.SY)
*备注:
Abstract:Transient-stability assessment determines whether a power system can recover after a disturbance and is therefore essential to preventing generator trips and cascading outages. A key metric is the critical clearing time (CCT), which specifies the maximum time available to clear a fault before synchronism is lost. Reliable CCT estimation is challenging because complicated fault-clearing dynamics require repeated simulations over many fault severities and clearing times. We propose an event-structured physics-informed neural network (ES-PINN) that aligns its representation with the pre-fault, fault-on, and post-clearing swing dynamics and enforces exact state chaining across event interfaces. A smooth trajectory-induced stability margin defines a differentiable approximation of the CCT boundary, enabling accurate boundary extraction, local sensitivity analysis, and optional direct CCT prediction through a distilled readout. We further prove a local residual-to-trajectory-to-CCT error estimate, in which exact event chaining eliminates separate state-interface defect terms. Experiments on IEEE 9-, 14-, and 30-bus systems show that ES-PINN consistently improves held-out trajectory and stability-boundary accuracy over matched neural-surrogate baselines across mechanical and electrical contingencies with multiple clearing configurations. Additional full-network DAE validation, multi-fault experiments, and runtime analyses further demonstrate the effectiveness and computational efficiency of the proposed framework.
[LG-51] FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning
链接: https://arxiv.org/abs/2607.27665
作者: Zekai Chen,Haodong Lu,Shihao Li,Weiwei Ji,Xunkai Li,Xun Wu,Yinlin Zhu,Rong-Hua Li
类目: Machine Learning (cs.LG)
*备注: 8 pages, 7 figures
Abstract:Federated graph learning enables collaborative training over decentralized graph data without sharing raw graph information. As such risks evolve, clients must learn emerging classes from private multimodal graph streams, retain historical categories, and reject samples outside the known class space. In this setting, clients must learn emerging classes from private multimodal graph streams while preserving historical categories and rejecting samples outside the current known class space. The core challenge is catastrophic forgetting, which in federated multimodal graphs is not merely a classifier-level failure: old knowledge can be erased through modality-semantic overwriting, topology-induced structural erosion, and federated memory fragmentation. To address this challenge, we propose \textbfFedOGL, a semantic-structural memory preservation framework. On the client side, FedOGL preserves historical decision behavior through replay and task-start distillation, while protecting graph-propagation memory via projection onto a globally shared structure basis. On the server side, FedOGL maintains and transfers compact category prototypes to facilitate cross-client knowledge sharing without exposing raw graph data. Extensive experiments demonstrate that, compared with the best-performing baselines, FedOGL reduces performance degradation caused by catastrophic forgetting by \textbf42.67%, while maintaining or improving performance on downstream tasks.
[LG-52] Evaluation Protocols and Cross-Subject Generalization in EEG Emotion Recognition
链接: https://arxiv.org/abs/2607.27655
作者: Hanting Suo,Yuwen Li
类目: Machine Learning (cs.LG)
*备注: 41 pages, 6 figures, 11 tables including supplementary information. Submitted to Applied Intelligence
Abstract:Reported accuracy in electroencephalography (EEG) emotion recognition depends on the complete evaluation procedure, not only the classifier. We separate the target quantity, development procedure, and reporting rule, then use one archived dynamical graph convolutional neural network (DGCNN) pathway on SEED and SEED-IV as an illustrative case. In a protocol-matched subject-dependent check, the SEED result was within 1.47 percentage points of the public reference value; the 3.40-point SEED-IV difference remained unresolved. Across 30 matched SEED subject-session trajectories, checkpoint selection based on repeated test-set evaluation increased mean window accuracy from 0.7855 at epoch 80 to 0.8892. Under five-fold subject-disjoint evaluation, validation-selected checkpoints achieved training-participant trial accuracies of 0.9990 on SEED and 0.9920 on SEED-IV. Accuracy for entirely held-out participants was 0.5348 (95% conditional subject-level bias-corrected and accelerated [BCa] interval [0.4667, 0.5985]) on SEED. The SEED-IV estimate was 0.3954 ([0.3343, 0.4648]) and is reported only as secondary sensitivity evidence because its protocol-matched compatibility check remained unresolved. The observed train-to-held-out-subject gaps are inconsistent with simple optimization underfitting, but they do not isolate subject identity from implementation, preprocessing, representation, or distributional factors. Supporting analyses further showed that participant rankings depended on representation and time scale, while a development-selected tail-risk ensemble did not establish a positive gain in a separate final evaluation. Subject-dependent, subject-disjoint, and cross-session results should therefore be reported as answers to different questions.
[LG-53] Certifying when decision-time information justifies adaptive experimentation
链接: https://arxiv.org/abs/2607.27651
作者: Jia Bi,Samuel Pinilla,Chenyang Zhu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted. We introduce Opportunity-aware Policy Authorization for Laboratories (\OPAL), a framework that decides whether adaptation should be enabled at all. \OPAL uses a precommitted contract to require non-trivial adaptation, controlled target risk and positive executed value after cost. We establish an impossibility boundary: source outcomes and unlabelled target covariates cannot uniformly support non-trivial authorization under unrestricted conditional outcome shift, and derive a target-calibrated recovery. Applied to an unseen 11,265-compound Cell Painting partition, the frozen gate selected 595 compounds, captured 384 positive opportunities and achieved strictly positive executed value under least-favourable completion; its 5.18% false-activation upper bound remained below a 7.5% limit. Among six methods, only \OPAL combined non-zero activation with this risk control. Locked pharmacogenomic and finite-campaign studies distinguish policy misalignment from non-certifiability, establishing authorization as a distinct layer for safe adaptive science.
[LG-54] First-order Constrained Trilevel Optimization Over Distributed Networks for Robust Coreset Selection
链接: https://arxiv.org/abs/2607.27632
作者: Yang Jiao,Kaixuan Jiao,Kai Yang,Nadjib Aitsaadi,Ilhem Fajjari,Renwei(Richard)Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:With the rapid advancement of the Internet of Things (IoT), massive amounts of data are generated across distributed edge networks. Training models on full data incurs significant computational overhead and storage bottlenecks, rendering coreset selection a critical paradigm. Furthermore, given the privacy-sensitive nature of local data and the escalating demand for model robustness in real-world deployments, developing an effective distributed optimization framework for robust coreset selection is vital, yet remains largely unexplored. To this end, this work first characterizes the hierarchical dependencies among coreset selection, robust optimization, and distributed learning, and formulates the distributed robust coreset selection as a trilevel optimization problem with level-wise constraints. Furthermore, to effectively solve the trilevel problem in a distributed manner, the \underlineFederated \underlineFirst-order \underlineConstrained \underlineTrilevel \underlineOptimization (F ^2 CTO) is proposed, which synergistically integrates a hierarchical composite value-function reformulation and a distributed alternating projected gradient algorithm. To the best of our knowledge, F ^2 CTO is the first method developed for distributed robust coreset selection, as well as the first distributed optimization approach for trilevel optimization problems with level-wise constraints. Additionally, we prove that the proposed method achieves a non-asymptotic convergence rate of \mathcalO(\epsilon^-3/2) for finding an \epsilon -stationary point. Extensive empirical evaluations on reliable continual learning demonstrate the effectiveness and efficiency of the proposed F ^2 CTO.
[LG-55] Real-Time Hard Peak Age-of-Information Safety with No-Regret Learning
链接: https://arxiv.org/abs/2607.27626
作者: Wentao Zhang,Wentao Mo
类目: Machine Learning (cs.LG)
*备注: Accepted to 2026 IEEE Real-Time Systems Symposium (RTSS)
Abstract:Safety-critical IoT systems such as industrial closed-loop control, V2X coordination, and remote teleoperation require every sensor’s peak Age of Information (peak AoI, also abbreviated PAoI) to stay below a hard per-slot deadline, not merely an average bound. Existing approaches meet this requirement only under restrictive assumptions: stochastic channels for Whittle-index AoI, simulator rollouts for deep reinforcement learning, or sublinear cumulative violation for long-term constrained online convex optimization. Under adversarial coefficients, OCO-PAoI-Hard guarantees zero per-slot violation of the modeled AoI state under one-step viability and O(sqrt(T)) regret against any static safe comparator; packet-level safety requires stronger service assumptions. Our key observation is that the fractional peak-AoI deadline collapses exactly to an affine half-space constraint on the resource-allocation vector, turning hard real-time scheduling into time-varying constrained online convex optimization over a polyhedral safe set. A strictly causal proposal-shield-update loop enforces feasibility through one Euclidean projection per slot, the gradient step preserves no-regret behavior, and the classical virtual queue is reduced to an a-posteriori certificate. We establish closed-form static and dynamic regret bounds, a matching Omega(sqrt(T)) minimax lower bound, a margin-safe variant against execution noise, and a deadline-induced competitive ratio. On a four-sensor adversarial fluid-model trap channel, OCO-PAoI-Hard attains zero modeled-state deadline violations across all ten seeds, while four representative baselines miss between 1.65 percent and 64.0 percent of slots, and the empirical normalized regret stays below the theoretical envelope across two orders of magnitude in T.
[LG-56] Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning
链接: https://arxiv.org/abs/2607.27610
作者: Haodong Zhu,Yangyang Ren,Yanjing Li,Sheng Xu,Haiguang Liu,Linlin Yang,Baochang Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy. This is challenging because prompt difficulty evolves throughout training. Existing online methods therefore face a trade-off: evaluation-based approaches are accurate but expensive, while prediction-based approaches are efficient but typically assume stationary difficulty, making them ill-suited to RL’s non-stationary training dynamics. To address these issues, we propose a Kalman-Guided Prompt Selection method (KGPS), which reformulates prompt selection as a dynamic state estimation problem rather than static difficulty prediction. KGPS models each prompt’s latent success rate in logit space using a linear-Gaussian state-space model, with process noise coupled to the magnitude of policy updates so that uncertainty increases when the policy changes more substantially. A Kalman filter then maintains a calibrated Gaussian posterior over prompt difficulty, and prompts are selected by maximizing a posterior-expected training utility that favors intermediate-difficulty prompts while naturally revisiting uncertain ones. The resulting procedure is adaptive to policy drift and requires no additional rollouts beyond standard policy training. Extensive experiments across mathematics, planning, and geometry reasoning benchmarks, as well as multiple RL algorithms, show that KGPS consistently improves both final accuracy and rollout efficiency over strong baselines, establishing state-of-the-art performance among online prompt selection methods. For example, on DeepSeek-R1-Distill-7B, KGPS uses 83% fewer rollouts than DS while even improving the average performance by 0.12 point across six math reasoning benchmarks.
[LG-57] Back from the Future: Key-Value Cache Management by Counter-Causal Surprise
链接: https://arxiv.org/abs/2607.27600
作者: Stephen Gould,Anton van den Hengel
类目: Machine Learning (cs.LG)
*备注: 17 pages, 5 figures
Abstract:Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years. Computational demands of large language models (LLMs) and their multi-modal variants during output generation can be partially alleviated by caching previous key and value calculations needed by subsequent scaled dot-product attention operations. However, this leads to another problem: the size of the resulting KV cache grows linearly with context length and quickly consumes all available GPU memory when either the prompt or the generated output are long. KV cache management periodically prunes entries from the cache thereby reducing its memory footprint while attempting to retain sufficient information for accurate generation. A by-product is faster inference speed. We propose a simple yet effective KV eviction scheme motivated by the insight that past tokens which can be well-predicted from more recent tokens are redundant and their associated keys and values can be removed from the cache. To score entries for eviction we run the model on the tokens in their original order, reusing the key and value representations already stored in the KV cache, and applying a counter-causal attention mask so that each position attends only to its future context. This is in-distribution, tied directly to the actual cache contents, and requires no additional training. To further reduce cost, we additionally propose a fast single-layer approximation that restricts the counter-causal pass to the last transformer layer, achieving a significant speedup per refresh cycle at marginal accuracy cost. We evaluate our strategy on various open-source LLMs and benchmark datasets showing competitive or improved performance over other state-of-the-art methods. Reference code is available at this https URL.
[LG-58] Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
链接: https://arxiv.org/abs/2607.27594
作者: Pankayaraj Pathmanathan,Furong Huang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, as LRMs personalization for downstream users takes center stage, the demand for varying levels of policy compliance grows as different user-specific LRMs must adhere to distinct subsets of safety policies. Training a separate LRM for each policy subset introduces severe combinatorial overhead. While in context learning methods overcome this combinatorial overhead, they introduce additional computational challenges associated with long context generation. To address this challenge, we propose \ours, a unified adaptive hypernetwork-based framework for multi-policy compliance. In our framework, safety policies serve as customizable inputs to a LoRA adapter generator, which learns to produce policy compliant LoRA weights for downstream LRM. When added to the LRM these weights enable the generation of responses compliant with the specified policy subsets. In this work, we demonstrate that training such a hypernetwork enables on-demand policy adjustments on a single LRM without sacrificing task performance across reasoning models of different sized and different evaluation datasets. This highlights the effectiveness and practicality of adaptive hypernetwork based alignment in LRMs.
[LG-59] MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
链接: https://arxiv.org/abs/2607.27581
作者: Zhankai Ye,Yukai Jin,Bingyang Wei,Bofan Li,Yusen Wu,Fangyi Li,Shangqian Gao,Xin Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion–language systems first coupled the two directions through a shared discrete motion codebook, but quantization limits generation quality. The strongest generators buy quality back at growing cost: stacked residual codebooks enlarge the representation; masked decoding stages, long autoregressive rollouts, and denoising chains of tens to hundreds of steps stretch inference; even the continuous-latent designs among them reach their latent only through an iterative diffusion head; and none of this decoding machinery serves understanding. We therefore propose MUGEN, a unified motion–language framework that pays neither cost: no codebook, one draw. A single adaptive-length autoencoder compresses any-length motion into a few continuous latent slots, the system’s only motion representation: the language model generates them for text-to-motion and reads them back for motion understanding. Depth-routed hidden states let each slot read from the transformer depth it needs, and a calibrated head predicts a joint distribution over the full latent set, so a single draw carries the text-conditional, cross-slot variation a description permits. At a decoding cost of K language-model steps, one draw, and one decoder pass, MUGEN leads language-model baselines on FID on HumanML3D while raising retrieval precision above the real-motion reference under the standard evaluator, achieves the best CIDEr and BLEU@4 scores, and surpasses the discrete-token state of the art on every retrieval and alignment metric on SnapMoGen.
[LG-60] Strategies for Milestone-driven Start-ups in Multi-activity Settings
链接: https://arxiv.org/abs/2607.27563
作者: Zhengli Wang
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR); Applications (stat.AP); Machine Learning (stat.ML)
*备注:
Abstract:New venture start-ups need to ``survive’’ through multiple stages of reaching milestone targets. We investigate the strategies for start-ups in a milestone-oriented setting. We examine a model of an entrepreneurial start-up firm, where its state is captured by a diffusion process. The entrepreneur can choose between multiple activities (or controls), which incur different cost and determine the drift and the variance of the process. Depending on whether the process reaches a fixed upper boundary or a lower one, the start-up firm succeeds or fails. Continuous-time stochastic models with multiple ( \ge 3 ) controls are typically very challenging to deal with. In this work, we are able to completely solve for the optimal policy and provide an explicit characterization of its structure. In particular, the optimal policy only uses controls from a set characterized by a so-called efficient frontier curve that orders the controls by two intuitive measures: riskiness (drift-to-volatility ratio) and cost-effectiveness (drift-to-cost ratio). A unique feature of our model is that depending on the model parameters, the efficient frontier curves can be of different types, resulting in qualitatively different structures of the optimal policy. As far as we know, this is the first study that analyzes a stochastic control model which admits efficient frontier curves of different types. Our work provides start-up firms with intuitive measures to evaluate their activities and offers valuable insights on how the optimal strategies in a milestone-oriented setting change qualitatively contingent upon the specific scenario. We believe the results provide a foundational block in the study of entrepreneurial decision-making.
[LG-61] Memory Efficient Tabular Foundation Models ICML2026
链接: https://arxiv.org/abs/2607.27546
作者: Shuting Luo,Monika Mikhail Kanaan,Cameron Gordon,Anna Leontjeva,Simon Lucey
类目: Machine Learning (cs.LG)
*备注: 12 pages, 3 figures Accepted at FMSD @ ICML 2026
Abstract:Tabular Foundation Models, such as TabPFN, have received a large amount of recent attention due to their performance on in-context tabular machine learning tasks, which often exceeds classical baselines. However, practical deployment considerations of these models has received less attention. In this paper we investigate the memory requirements for these models. We demonstrate that employing model compression approaches can enable memory reductions of up to 7.6 with similar levels of performance, reducing deployment requirements by nearly 87%. Our work provides insight to practitioners seeking efficient deployment of these models in practical settings.
[LG-62] When Does Explicit View Routing Work? A Controlled Study of Multi-View Graph-Text Alignment
链接: https://arxiv.org/abs/2607.27530
作者: Xiao Yue,Guangzhi Qu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Graph-text retrieval typically maps a graph and its description to a single embedding, even when a query concerns only one semantic aspect, such as a class label or molecular property. Multiple heads can separate these aspects, but a change in the query head may alter retrieval even when the wrong text is sent to that head. Such behavior demonstrates architectural channelization, not necessarily semantic routing. We examine the conditions under which this distinction can be resolved. Our controlled version of MV-GTA uses deterministic, verifiable text segments; isolated text encoders; view-specific graph heads; and relevance derived from external labels or RDKit descriptors. Correct routing and per-sample derangements form a causal test of whether retrieval depends on content. On BBBP and BACE, correct routing improves label and property nDCG by 0.305 to 0.685 over deranged training. The expected graph head exceeds the best wrong head by 0.303 to 0.453. Topology does not specialize consistently across the two datasets. In a matched three-seed comparison, one joint model obtains mean topology, label, and property nDCG of 0.720/1.000/0.877; three separately trained Single specialists obtain 0.633/0.976/0.859. Property paraphrase augmentation also improves unseen-template nDCG by 0.140 and 0.147 over a matched-exposure canonical control. Consistency and hard-template extensions, however, reduce canonical retrieval in some settings. The evidence is therefore limited to explicit, externally grounded label and property routing and observed multi-interface consolidation. It does not establish free-form routing, consistent three-view specialization, statistical equivalence to specialists, or superior downstream prediction.
[LG-63] Latent-Kernel Discrete Flow Maps for Few-Step Generation
链接: https://arxiv.org/abs/2607.27529
作者: Mansoor Ahmed,Yue-Tsz Fan,Hemanth Venkateswara,Murray Patterson
类目: Machine Learning (cs.LG)
*备注:
Abstract:Discrete diffusion and flow-matching models denoise a sequence over many steps, but to keep each step cheap, they factorize the transition across positions and decide every token independently. This makes few-step generation challenging for text when the target couples two positions, such as a subject and a verb that must agree. An independent update commits to them separately, and many function evaluations are spent repairing the mismatch. Existing few-step methods buy back the lost correlation by distilling or rectifying a slow teacher, and so inherit the teacher’s quality ceiling. We ask instead whether a model can express correlated steps natively, and answer with Latent-Kernel Discrete Flow Maps (LKF), a from-scratch flow-map kernel that is a mixture of M factorized components tied by a single shared latent. Conditioned on the latent, each component is cheap, and the mixture is summed over the latent in closed form for small M. We show that a single step places mass on correlated completions with the same sampling time complexity as a factorized model, since one latent is drawn per sequence and reused across the entire denoising trajectory. We also show that the Masked Diffusion Language Model (MDLM) is a special case of our LKF model at M=1. The experiments for unconditional text generation on the One-Billion-Word (LM1B) and WikiText-103 benchmarks show that our LKF model learns strongly heterogeneous components and improves generative perplexity by 2.1x to 3.3x over the likelihood baselines without losing diversity. The gain grows with M, and at M=8, it surpasses distilled and rectified few-step samplers. The source code is available at: this https URL
[LG-64] Sparsity Induced Identifiability in Matrix Tri-Factorisation
链接: https://arxiv.org/abs/2607.27507
作者: Tingting Mu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Matrix factorisation is a fundamental tool for exploiting low-dimensional structure in high-dimensional data, with applications such as data compression, denoising, structure discovery, interpretable representation learning, and dimensionality reduction. Compared to conventional two-factor models, matrix tri-factorisation provides greater modelling flexibility, while sparsity constraints often improve both interpretability and recovery performance. Although the role of sparsity has been extensively studied for two-factor matrix factorisation, rigorous theoretical guarantees for general real-valued matrix tri-factorisation remain largely unexplored. To address this gap, we establish, to the best of our knowledge, the first rigorous theoretical study for sparsity-induced identifiability in general real-valued matrix tri-factorisation. Our analysis is enabled by a novel decomposition strategy that transforms the original problem into two coupled auxiliary factorisation problems, while preserving the structural information necessary to the recovery of the original factor matrices from the observations. Building upon this decomposition, we derive recovery guarantees and structural consistency results that characterise how coefficient sparsity influences the sufficient recovery conditions, convergence behaviour, spectral approximation error, high-probability bounds, and structure preservation. Comprehensive Monte Carlo experiments validate the proposed theory and demonstrate close agreement between the theoretical results and empirical observations.
[LG-65] A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation
链接: https://arxiv.org/abs/2607.27501
作者: Liangyu Wu,Qibin Liu,Alexander Yue,Julia Gonski
类目: Machine Learning (cs.LG); High Energy Physics - Phenomenology (hep-ph)
*备注: 18 pages, 5 figures, 1 table
Abstract:We present a lightweight approach to foundation modeling (\textbfNEXUS) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder model with approximately 3 million parameters. The model pre-trains with no supervision over a large-scale collision dataset from the Large Hadron Collider modeled by charged particle track features. Downstream tasks for collider analyses, such as kinematic regression and event classification, are developed on pre-trained model weights and achieve improved accuracy with only small labeled datasets when compared to equivalent architectures trained from scratch. The benefits of pre-training are additionally investigated through latent space interpretation and application to other domains, including gravitational waves, flood forecasting, and neural activity. Furthermore, the relative computational simplicity of NEXUS is demonstrated compared to transformer approaches at comparable scale, opening the door to power-efficient inference and real-time or edge applications of foundation models in scientific experiments.
[LG-66] Schreier-Coset Graph Rewiring
链接: https://arxiv.org/abs/2607.27479
作者: Aryan Mishra,Randy Martinez,Lizhen Lin
类目: Machine Learning (cs.LG)
*备注: 26, 3
Abstract:The information flow in the graph neural networks (GNNs) is fundamentally constrained by over-squashing, where structural bottlenecks impede long range information propagation. Graph-rewiring methods, which modify graph topology, have been extensively used to alleviate this. However, existing approaches often introduce prohibitive structural and computational bottlenecks, fail to preserve the critical properties of original graphs, and increase the edge counts massively. We introduce a novel method Schreier-Coset Graph Rewiring , a group-theoretic rewiring method that augments the input graph with a Schreier-Coset graph derived from a special linear group. Our method provides theoretical guarantees, a graph that exhibits spectral gap and a bounded effective resistance, creating a low-resistance bypass for long-range communication. Empirical evaluations demonstrate that SCGR reduces effective resistance by 5-40% across various learning tasks, effectively mitigating connectivity bottlenecks while maintaining competitive accuracy.
[LG-67] Neural Network-Assisted CLEAN for Channel Modeling in Low-SNR Regimes
链接: https://arxiv.org/abs/2607.27450
作者: Chaofan Deng,Linyu Sun,Jaeho Lee,Arijit Raychowdhury
类目: Machine Learning (cs.LG)
*备注:
Abstract:Accurate multipath parameter estimation is critical for modern wireless communication systems, particularly in challenging low-SNR environments. Traditional Maximum Likelihood Estimation algorithms, such as CLEAN, provide high-resolution parameter extraction but suffer from prohibitive computational complexity due to exhaustive grid search. Conversely, purely data-driven deep learning approaches lack physical grounding and struggle to generalize across variable multipath densities and off-grid parameters. To address these limitations, this paper proposes Neural Network-Assisted CLEAN (NN-CLEAN), a hybrid framework that embeds a multi-head residual network directly into the iterative CLEAN extraction loop. By replacing the exhaustive grid search with rapid, parallelizable forward passes while delegating residual subtraction to exact mathematical models, NN-CLEAN isolates physical multipath parameters without accumulating non- physical errors. Extensive Monte Carlo simulations demonstrate that NN-CLEAN achieves estimation accuracy exceeding 96% at 5 dB SNR, matching the traditional Grid-Search CLEAN (GS- CLEAN) baseline, while providing a massive reduction in computational complexity and substantially outperforming subspace methods and standalone one-shot neural networks. Crucially, NN-CLEAN exhibits a near-flat scaling in execution runtime and memory consumption as batch sizes increase. This highly efficient parallelization establishes NN-CLEAN as a robust, real- time solution for channel estimation in MIMO systems.
[LG-68] Comparison of a Parametric Physics-Informed Neural Network and a Tensorial Reduced-Order Model for the Shallow-Water Dam-Break Problem
链接: https://arxiv.org/abs/2607.27433
作者: Anton Myshak,Md Rezwan Bin Mizan,Ilya Timofeyev
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG); Fluid Dynamics (physics.flu-dyn)
*备注:
Abstract:We develop two parametric data-driven reduced models: a physics-informed neural network (PINN) and a non-intrusive tensorial reduced-order model (TROM), and apply both approaches to the parametrized one-dimensional shallow-water dam-break problem. Both reduced models do not require time integration and learn a direct solution map from space, time, and dam-break parameters to the physical state. We present a detailed comparison for out-of-sample and extrapolated parameter values. In addition, we demonstrate that it is essential to introduce shock-aware collocation to improve the robustness of the PINN model.
[LG-69] Good Rankers Bad Objectives: Bilinear Contrastive Critics under Expressive Policy Search
链接: https://arxiv.org/abs/2607.27422
作者: Ayushman Singh,Siddharth Aphale
类目: Machine Learning (cs.LG)
*备注:
Abstract:Good action rankings do not make a contrastive critic safe to maximize. These critics increasingly act as value-like objectives for best-of- K selection, planning, and critic-guided generation. Unbounded bilinear scores can let large embedding norms inflate off-support values, but cosine bounding does not remove the failure. A controlled support decomposition attributes most raw bilinear regret to norm drift. Cosine and hybrid critics nevertheless select off-support actions from most pools and incur comparable regret. Contrastive scores are weakly calibrated or inverted in the top score decile across four OGBench navigation tasks, and they fail to order fixed-query actions by value. Bellman-trained TD-Q succeeds, including in a parameter-matched function-class control. Realized costs depend on the task: simulator rollouts reveal single-step selection costs on PointMaze and the exact- Q^* toy but well-powered nulls on AntMaze and HumanoidMaze, where the controller can self-correct. A training/readout decomposition traces the lost ordering to the cosine training objective; raw-trained embeddings retain weak ordering after inference-time normalization. Candidate maximization can therefore exploit false positives caused by norm drift, score saturation, or in-support misranking. Contrastive critics remain useful compatibility rankers on navigation and manipulation tasks, but action selection requires a value-calibrated scalar.
[LG-70] Context-Informed Ship Trajectory Prediction via Conditional Attention
链接: https://arxiv.org/abs/2607.27418
作者: Yuan Guan,Chandler Squires,Timothy Hu,Pradeep Ravikumar
类目: Machine Learning (cs.LG)
*备注: Accepted for publication at the 2026 29th International Conference on Information Fusion (IEEE FUSION)
Abstract:Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architectures have improved forecasting horizons, they predominantly rely on historical kinematic states, treating vessel motion as an isolated system. In reality, maritime navigation is profoundly modulated by extrinsic factors like weather and constrained by static vessel characteristics. Existing multimodal approaches fundamentally model the joint distribution over states and contexts, treating environmental variables as peer features rather than encoding the directional physical dependence of vessel dynamics on environmental conditions. In this work, we propose the Conditional Informer, a novel encoder-decoder architecture that formulates trajectory prediction as a conditional generation task. We employ a dedicated Conditional Attention mechanism where the vessel state explicitly queries environmental contexts through cross-attention, encoding the physical prior that weather modulates - but is not generated by - vessel dynamics. Furthermore, to address the intermittency of real-world data, we introduce a Modality Masking training strategy to prevent catastrophic degradation during sensor fallback. Extensive experiments on AIS and ERA5 data demonstrate that our approach outperforms kinematic and concatenation-based baselines by 15.4% in prediction accuracy when context is available. Crucially, Modality Masking prevents shortcut learning, reducing fallback error by nearly an order of magnitude compared to unconstrained models. Comments: Accepted for publication at the 2026 29th International Conference on Information Fusion (IEEE FUSION) Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.27418 [cs.LG] (or arXiv:2607.27418v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27418 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-71] he Convergence Behavior of Adam under Heavy-Tailed Noise UAI2026
链接: https://arxiv.org/abs/2607.27383
作者: Yijiang Pang
类目: Machine Learning (cs.LG)
*备注: Accepted at UAI 2026
Abstract:We establish the first convergence guarantees for the plain vector-form \emphAdam optimizer under heavy-tailed stochastic noise. While several Adam variants are known to achieve optimal iteration complexity in bounded-variance nonconvex optimization, little is understood about their behavior when stochastic gradients admit only a bounded p -th central moment for some p \in (1,2] , a setting increasingly observed in modern deep learning. To address this gap, we generalize the recent online-to-nonconvex conversion framework to accommodate heavy-tailed martingale-difference noise. Building on this generalized framework, we develop a discounted regret analysis for Adam, without restrictive parameter coupling. Our results show that Adam converges to (\rho,\epsilon) -stationary points under heavy-tailed noise. However, it exhibits a suboptimal iteration complexity and p -dependent convergence, a suboptimality that persists even in the bounded-variance case ( p=2 ). When the domain radius is known and used to control the online-learner output, a standard setup in related literature, the convergence rate improves to match the optimal complexity. These findings provide new theoretical insight into the robustness and limitations of Adam in heavy-tailed regimes.
[LG-72] Compression-Based Behavioral Similarity for Open-World Sybil Discovery on Ethereum
链接: https://arxiv.org/abs/2607.27370
作者: Michał Bartnicki,Jarosław A. Chudziak
类目: Machine Learning (cs.LG)
*备注: Accepted as a full paper and scheduled for presentation at the European Conference on Advances in Databases and Information Systems (ADBIS 2026)
Abstract:Sybil attackers are Blockchain actors that adopt the characteristics of regular users to exploit airdrops or influence governance. Current methods of Sybil actor detection include constructing graphs, which requires token transfers between examined wallets. Machine learning algorithms have been employed as well, but they treat the task as a closed-set classification problem, making them vulnerable to frequent changes in attack strategies or evasion tactics. We address the following questions: can compression-based similarity differentiate Sybil bots, organic users, and arbitrage bot wallets without direct financial links? What is the effect of high-signal contracts on the discovery of Sybils, and how robust are behavioral graphs under temporal drift and adversarial perturbations? Our approach synthesizes a symbolic Transaction Grammar from EVM (Ethereum Virtual Machine) traces, capturing separately transaction rhythm, execution structure, and functional intent. The high-signal contracts are filtered with our own protocol, called the Blind-Spot Protocol. Gzip-based NCD is used to construct a behavioral graph for Sybil discovery. We validate this framework against supervised machine learning baselines, a temporal split, and synthetic camouflage stress tests. Ultimately, we contribute a leakage-aware behavioral framework for Sybil candidate discovery. Its core NCD primitive requires no supervised training and can expand suspicious seed wallets without explicit funding links. We position the method as a training-free local discovery primitive for open-world blockchain audits, rather than as a formal open-set recognition system.
[LG-73] Modeling Decisions in Blockchain Analytics: A Leakage-Aware Evaluation of Tree-Based vs. Sequential Models
链接: https://arxiv.org/abs/2607.27350
作者: Michał Bartnicki,Jarosław A. Chudziak
类目: Machine Learning (cs.LG)
*备注: Accepted for presentation at the 23rd International Conference on Modeling Decisions for Artificial Intelligence (MDAI 2026)
Abstract:Sybil bots are Ethereum actors that imitate legitimate users to extract airdrop rewards or influence governance. Recent Sybil detection methods increasingly use deep learning and treat blockchain activity as a quasi-linguistic sequence. However, complex sequence models are computationally expensive for real-time monitoring, and their reported performance may be inflated by label leakage from high-signal smart contracts. We ask whether and how organic users, Sybil bots, and MEV bots differ in the structural complexity of their transaction histories; whether sequential models outperform tree-based tabular models once leakage is reduced; whether transaction order or timing provides the stronger behavioral signal; and whether the resulting models are practical for low-latency deployment. Our approach to leakage-aware Sybil bot detection consists of a Blind-Spot protocol and a Transaction Grammar representation of wallet behavior. The former eliminates shortcuts associated with high-signal contracts, whereas the latter models wallets using rhythm, EVM execution structure, and intent. We evaluate this approach on Ethereum actor classification by comparing Transformer and BiLSTM sequence models against XGBoost and SVM baselines. We contribute a framework for leakage-aware Ethereum actor classification and a Transaction Grammar representation of wallet behavior. Our results demonstrate that, under leakage-aware evaluation, XGBoost outperforms Transformer-based sequence models while providing lower latency and estimated energy use.
[LG-74] ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
链接: https://arxiv.org/abs/2607.27308
作者: Christopher Warner,Jonas Mago,JR Huml,Beren Millidge
类目: Machine Learning (cs.LG); Neurons and Cognition (q-bio.NC)
*备注:
Abstract:We introduce ZUNA1.1, a 380M-parameter diffusion autoencoder for flexible EEG signal reconstruction. ZUNA1.1 is capable of reconstructing variable length sequences of up to 30s, with an arbitrary number of EEG channels at arbitrary scalp locations, and can reconstruct arbitrary temporal intervals within channels in addition to reconstructing entire channels. We demonstrate that ZUNA1.1 performs at least on par with our earlier ZUNA1 model, while being far more flexible and capable of handling a wide range of reconstruction tasks. ZUNA1.1 continues to substantially outperform standard EEG denoising and reconstruction methods such as spherical spline interpolation, which is ubiquitously deployed in the MNE package. The ZUNA1.1 model is released open source under the permissive Apache 2.0 license.
[LG-75] EvoCause: LLM -Guided Evolution of Causal Graphs for Root Cause Analysis
链接: https://arxiv.org/abs/2607.27290
作者: Lei Zan,Keli Zhang,Shifeng Xie,Jiale Zheng,Zehao Xiao,Zhiwei Dong,Ke Zhang,Ruichu Cai,Malik Tiomoko,Lujia Pan
类目: Machine Learning (cs.LG)
*备注:
Abstract:Modern telecommunication, cloud, and microservice systems emit correlated alarm cascades when components fail. Root cause analysis (RCA) aims to identify the small set of alarms that initiate each cascade. A common approach learns a causal graph from observational logs and predicts all zero-in-degree alarms in each incident-induced subgraph. However, the learned graph remains fixed and cannot benefit from expert diagnoses of historical incidents. We close this loop with EvoCause. Expert labels constrain which alarms should be source nodes but do not specify the edge edits needed to satisfy those constraints. EvoCause uses a large language model (LLM) to propose semantically plausible graph edits, while deterministic code validates node identities and acyclicity and retains the best graph on a labeled alignment set. At test time, the refined graph alone produces transparent predictions without an LLM call. We also release TeleRCA, an expert-annotated benchmark from a production telecommunication network containing 485,681 alarm events spanning 194 alarm types over 5,621 resources. On synthetic data, EvoCause initialized with the PC causal discovery algorithm outperforms the unrefined PC baseline, raising Node F1, Case EM, and Graph F1 by 11.59 , 9.40 , and 4.59 percentage points, respectively, while reducing nSHD by 0.2379 . On TeleRCA, replacing human-readable alarm titles with anonymous identifiers lowers Node F1 and Case EM by 6.12 and 8.04 percentage points, respectively, indicating that alarm-name information contributes to graph refinement.
[LG-76] IER-MoE: Trust-Informed Expert Routing via Conditional Modality Risk for Multimodal Fusion in Biomedical Classification
链接: https://arxiv.org/abs/2607.27289
作者: Yu Chang,Anzhe Cheng,Chenwei Wu,Zhuoran Wang,Jiahao Chen,Tamoghna Chattopadhyay,Sophia I. Thomopoulos,Paul M. Thompson,Liyue Shen,Paul Bogdan
类目: Machine Learning (cs.LG)
*备注:
Abstract:The promise of multimodal fusion lies in combining complementary sources of evidence, yet more evidence does not always yield a better prediction. Recent multimodal models have advanced fusion through richer cross-modal interaction and sample-adaptive fusion. However, the influence assigned to a modality during fusion does not reveal whether that source is unreliable, redundant, or poorly matched to a specialized expert. To address this limitation, we introduce TIER-MoE, a risk-guided subspace mixture-of-experts model that defines sample-specific modality reliability as the prediction loss its unimodal predictor is expected to incur. This risk is learned from out-of-fold predictions generated by models that were not trained on the corresponding sample. TIER-MoE combines the estimated risk with expert-specific subspace compatibility for sparse modality-expert routing, while an always-active shared path preserves multimodal complementarity. We evaluate TIER-MoE on four public multimodal biomedical datasets spanning Alzheimer’s disease status, skin-lesion malignancy, and retinal classification. Results demonstrate its superiority over state-of-the-art methods in predictive performance and probability calibration, with consistent improvements in Macro-F1 and Brier score and strong zero-shot generalization to an external cohort.
[LG-77] he Kinetics of Training: A Driven-Nucleation Rate Law for Emergence Plasticity Loss and Circuit Control in Language Models
链接: https://arxiv.org/abs/2607.27281
作者: Lei Dong
类目: Machine Learning (cs.LG)
*备注: 40 pages, 12 figures
Abstract:A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step of capability formation. Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait counts missing parts, not size; and on Pythia across seven capabilities and three scales, ablating one part leaves a median 17% of the capability in 32 of 32 discriminating cells, where partial credit predicts 50-83% (p = 2e-10), while a random non-part head leaves 100%. One rare event whose barrier grows with missing parts yields a rate equation – sites x attempts x drive x exp(-beta*K), minus destruction – read three ways, each preregistered with frozen constants. Forward: a capability flat at baseline ignites at a step of our choosing once the mix passes a concentration floor (10/10 above, 0/12 below), and while still flat its arrival is datable from its precursor to 5% median error on six held-out models. Backward: the delay to learn a withheld capability grows with waiting until, past a critical step, it never ignites – yet validation loss falls smoothly throughout, so standard monitors are blind to it. We locate the damage (heads commit to the base data) and isolate the cure: re-initializing only the query-key slices restores learnability (6/6) while the value slices do nothing (0/6). We prove the mechanism in a controlled gated-attention model: occupation forces a deadline whose consequences need no mixing assumption. Completed: SGD’s noise fails the fluctuation-dissipation test, so we install one and anneal, melt and pin circuits on schedule. Scope: conjunction circuits in transformers to 1.4B. Comments: 40 pages, 12 figures Subjects: Machine Learning (cs.LG) Cite as: arXiv:2607.27281 [cs.LG] (or arXiv:2607.27281v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2607.27281 Focus to learn more arXiv-issued DOI via DataCite
[LG-78] Rethinking EEG-Based Disease Diagnosis: Decoupling Instance Representation Learning from Subject-Level Supervision
链接: https://arxiv.org/abs/2607.27274
作者: Zhiyuan Ma,Zeyuan Li,Zhiyi Lu,Jiacheng Hao,Youlang Du,Zhen Jiang,Xinche Zhang,Yuhao Sun,Sen Song
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:EEG-based disease diagnosis requires one prediction per subject, yet common pipelines segment recordings into short instances, inherit the subject label for every instance, and train instance-level classifiers. This assumes that all instances provide equally reliable diagnostic evidence. Multiple instance learning (MIL) avoids inherited labels by treating each subject as a bag. However, EEG datasets contain far fewer subjects than instances, which can limit the quality of the representations learned by end-to-end MIL. We propose BridgeMIL, a two-stage framework that decouples instance representation learning from subject-level supervision. Stage 1 pretrains the encoder without inherited instance labels by aligning temporally nearby windows and independently sampled within-subject sub-bags. Variance and covariance regularization prevent collapse and reduce redundancy without negative pairs. Stage 2 transfers the encoder to an attention-based MIL aggregator, applies supervision only to subject predictions, and limits representation drift through feature retention. Across three EEG disease datasets and five representative backbones, BridgeMIL attains the highest mean accuracy in 14 of 15 dataset-backbone settings and an overall mean accuracy of 76.57%, 4.28 percentage points higher than the strongest baseline. Further analyses reveal substantial variation in inherited-label reliability across instances, greater performance sensitivity to subject scarcity than to instance scarcity, and a more structured representation space with distinct subject-wise clusters and improved separation between diagnostic classes. Together, these findings underscore the importance of aligning supervision with the subject-level prediction objective while learning from abundant EEG instances without assigning disease labels to individual instances.
[LG-79] SDO: Structure-Aware Data Organization for Efficient LLM Post-Training
链接: https://arxiv.org/abs/2607.27273
作者: Jinliang Gao,Ning Yang,Hai Wang,Baili Xiao,Pin Lyu
类目: Machine Learning (cs.LG)
*备注: 9 pages, 5 figures, 5 tables
Abstract:Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules. However, data organization itself is usually treated as a static preprocessing step: embedding-based grouping methods construct fixed partitions before training and cannot adapt to the evolving sample exposure during optimization. As a result, all samples receive similar exposure despite their different optimization needs, leading to redundant updates for some samples while leaving others under-optimized. To address this problem, we propose SDO (Structure-Aware Data Organization), a plug-and-play data organization framework with an exposure-driven feedback mechanism that organizes mini-batch composition and sample exposure according to representation-space structure. SDO operates epoch by epoch on frozen external embeddings, avoiding model warm-up training overhead: within each epoch, locality-aware batching forms coherent mini-batches via KNN neighborhood traversal; across epochs, exposure-balanced scheduling records per-sample participation and reduces the sampling probability of over-exposed samples to preserve long-term coverage. Across SFT, DPO, and GRPO, SDO accelerates convergence, with the largest gains observed in the early-to-mid phase, producing more coherent gradients and more balanced accuracy across question types without permanently excluding training samples.
[LG-80] RLPF: Reinforcement Learning from Performance Feedback for Code Generation
链接: https://arxiv.org/abs/2607.27271
作者: Huihao Jing,Haozhe Cui,Wenbin Hu,Shaojin Chen,Haochen Shi,Changxuan Fan,Yuxuan Liu,Hanyu Yang,Sirui Zhang,Ziyi Chen,Haoran Li,Yangqiu Song
类目: Machine Learning (cs.LG); Software Engineering (cs.SE)
*备注:
Abstract:Code models are increasingly trained with execution feedback, but most training signals still stop at correctness. This leaves an important gap for systems code: two programs can pass the same tests while differing greatly in runtime. We study how to train code agents to prefer faster correct implementations, rather than treating efficiency only as an evaluation metric. The key difficulty is that runtime is a fragile reward. It is meaningful only after a program is correct, varies across tasks, and gives little guidance when most sampled programs fail to compile or run. We propose \textbfRLPF, reinforcement learning from performance feedback, which turns execution outcomes into a staged reward. Failed programs are ordered by execution progress, while correct programs are ranked by their relative improvement from the baseline toward the expert reference. This gives useful feedback before correctness and performance-sensitive feedback after correctness. Fine-tuning Qwen3-32B with RLPF on PerfCodeBench raises correct-and-runnable solutions from 11.1% to 54.6% and improves relative efficiency from 8.1% to 38.6% . The trained model becomes competitive with stronger open-weight systems, and its optimization behavior transfers modestly to EffiBench-X. Additional studies show that model-generated references provide useful but weaker supervision, and that the full composite reward is more reliable than correctness-only or runtime-only baselines. These results suggest that code agents can be trained not only to pass tests, but also to optimize the programs they write.
[LG-81] Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
链接: https://arxiv.org/abs/2607.27269
作者: Weiye Shi,Fanxu Meng,Muhan Zhang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic. Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA’s cache efficiency without retraining from scratch. Speculative decoding offers complementary acceleration, but its speedup depends on agreement between draft proposals and target verification. We find that direct MHA/GQA-to-MLA conversion can sharply reduce this agreement: low-rank factorization and RoPE handling introduce attention-function errors that may be tolerable for standalone generation but substantially lower draft-token acceptance. We therefore formulate MLA draft construction as functional reconstruction rather than cache compression. Our end-to-end (E2E) method optimizes each converted MLA attention module to reproduce the post-output-projection response of its original MHA/GQA counterpart on calibration hidden states. This converter-agnostic post-conversion procedure preserves the converted cache and inference graph and requires neither verifier logits nor verifier supervision. We evaluate 192 model-converter-backend-method-task configurations spanning four Llama/Qwen draft-target pairs, TransMLA and MHA2MLA, HF and vLLM, and four 200-prompt tasks. With a 0.5-percentage-point reporting tolerance, Functional Reconstruction materially improves acceptance in 37 of 64 matched task cells, leaves 26 practically unchanged, and materially decreases one. Code and evaluation artifacts are available at this https URL.
[LG-82] PlatformBid: An Auto-Bidding Benchmark from a Unified Advertising Platforms Perspective
链接: https://arxiv.org/abs/2607.27265
作者: Shengtian Yang,Yewen Li,Peng Jiang,Zhiyi Lyu,Bo An,Peng Jiang,Qingpeng Cai,Lei Feng
类目: Machine Learning (cs.LG)
*备注:
Abstract:Real-time bidding is central to computational advertising, comprising three elements: Supply Side Platform (SSP) selling ad impressions, Demand Side Platform (DSP) bidding for advertisers, and Ad Exchange conducting auctions between them. Traditional auto-bidding algorithms focus solely on the DSP side, maximizing advertiser conversions by adjusting bids against competitors. However, current big ad platforms, such as social media and e-commerce companies, now integrate SSP, DSP, and Ad Exchange functions internally. From such ad platforms’ perspective, the goal of the auto-bidding algorithms is not only to maximize the advertisers’ conversions, but also the total revenue of the platform. Given the lack of platform-centric evaluation frameworks and the pressing need to advance auto-bidding research, we propose PlatformBid - the first comprehensive benchmark designed from a unified ad platform’s perspective. To accurately reflect the real-world auto-bidding scenarios, we define three representative settings: (1) homogeneous competition with identical algorithms across advertisers, (2) heterogeneous competition with diverse algorithmic strategies, and (3) promotional competition where some advertisers surge budgets for boosting sales during promotional events like Black Friday. We systematically evaluate a broad spectrum of existing auto-bidding methods across these settings, encompassing classical control methods, RL-based methods, and recent generative methods. Besides these methods, we further propose a novel auto-bidding method based on flow-matching, termed BidFlow, which leverages the flow-matching method’s expressive policy representation to effectively handle dynamic competitive environments. Online experiments on Kuaishou further show a +0.68% improvement in target cost, providing deployment evidence for the offline-online consistency of PlatformBid.
[LG-83] DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series
链接: https://arxiv.org/abs/2607.27263
作者: Dennis Thumm,Billy Tim Anthony,Ying Chen
类目: Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an); Methodology (stat.ME)
*备注:
Abstract:Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science. We introduce \textbfDoTime, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventions, released as the \codedotime PyPI package together with four frozen evaluation suites. Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emphwindows, counterfactual sampling modes with a positivity guard, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics by construction with switching SCM parameters, and deterministic ramp and sinusoidal intervention profiles that place trends and structural breaks \emphinside the evaluation window. Moreover, it demonstrates the suitability of the generator as a prior for a causal foundation model reference implementation. The released suites span a training-scale snapshot of 100,000 trajectories and eight named identification structures, each with exact ground truth: paired interventional trajectories from the same SCM throughout, and shared-noise counterfactuals in the continuous-time suite. We ship reference baseline implementations with an evaluation harness, and pose a falsifiable claim: interventional training buys a measurable direction-accuracy advantage over an observational model of identical capacity. It is tested across three training seeds per arm. Under structure-matched evaluation on held-out episodes, the interventional prior-fitted network’s (PFN) gap is positive in every structure, trajectory length, and seed tested.
[LG-84] Regularizing modality contribution drift in multimodal continual learning
链接: https://arxiv.org/abs/2607.27260
作者: Zhen Zhang,Jielei Chu,Bin Liu,Tianrui Li
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multimodal continual learning (MMCL) aims to learn emerging knowledge from multimodal data while preserving knowledge. To mitigate forgetting, current MMCL methods usually focus on cross-modal representation alignment or semantic similarity, but they overlook whether the relative contributions of individual modalities and their interactions remain stable across incremental tasks. We term this decision-level shift Modality Contribution Drift (MCD) and quantify it with the MCD score, which combines contribution-strength and relative-reliance changes under controlled interventions on modality subsets. Theoretical and empirical analyses further explain why current MMCL methods cannot reliably mitigate this drift. To this end, we propose Continual Modality Contribution Drift Regularization (CMCDR), which preserves the modality contribution structure of previously learned tasks. Since MMCL settings differ in whether old exemplars are available, CMCDR includes both replay-based and replay-free versions. The replay-based version uses modality-subset interventions as diagnostic probes on stored old samples, compares their contribution profiles between the current model and a frozen previous model, and constrains changes in old-sample modality-specific and interaction contributions. The replay-free version uses current-task samples as probes and distills the frozen model’s old-task contribution responses, thereby regularizing the observed contribution profile without exemplars. Experiments on multimodal class-incremental learning and continual visual question answering validate the generality and effectiveness of CMCDR.
[LG-85] Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories
链接: https://arxiv.org/abs/2607.28567
作者: Mengfei Ran,Yifeng Shen,Ruijie Guan
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Computation (stat.CO); Methodology (stat.ME)
*备注:
Abstract:Longitudinal causal studies often record histories as irregular functional fragments: laboratory values, physiologic signals, sensor streams, and image-derived summaries measured at unequal and informative times. Standard doubly robust estimators usually require scalar summaries, whereas sequence learners optimize prediction losses that need not stabilize the efficient influence function. We propose Doubly Robust Functional Representation Learning (DR-FRL), a cross-fitted workflow that turns irregular histories into estimand-targeted states for observed-history regimes. Functional and temporal encoders map point clouds and prior histories into states; nuisance heads estimate outcome, treatment, and censoring functions; and EIF-targeted validation, calibration, overlap, tail, and ablation diagnostics assess whether the state supports the estimating equation. If the selected state preserves the nuisance information needed by the EIF, representation error enters the same second-order product remainder as ordinary nuisance error, and the mean estimator is asymptotically linear under explicit rate, overlap, calibration, and stability conditions. Catoni aggregation is treated separately as a bounded-influence point estimator, not a replacement for Wald inference. Simulations show gains when functional confounding is high-dimensional, measurement is informative, support is weak, or pseudo-outcomes are heavy-tailed. A VitalDB audit shows that DR-FRL can use irregular laboratory point clouds and deliver a useful negative finding: for this ICU-disposition endpoint, scalar laboratory summaries already carry much endpoint-relevant information.
[LG-86] Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets
链接: https://arxiv.org/abs/2607.28537
作者: Ali Rayat,Yunhao Fan,Gia-Wei Chern
类目: rongly Correlated Electrons (cond-mat.str-el); Machine Learning (cs.LG); Computational Physics (physics.comp-ph)
*备注: 15 pages, 5 figures
Abstract:Metallic magnets exhibit complex spin dynamics governed by electronically generated interactions. Predictive simulations of such dynamics typically require repeated solutions of an underlying electronic problem throughout the time evolution, creating a major computational bottleneck. Here we introduce a graph neural network (GNN) magnetic force-field framework that learns the effective magnetic energy functional governing itinerant spin dynamics directly from electronic calculations. Conceptually analogous to machine-learned interatomic potentials, the proposed framework enables efficient evaluation of spin torques while capturing the nonlinear and spatially extended interactions generated by itinerant electrons. We benchmark the method on representative metallic magnetic systems exhibiting collinear, noncollinear, and noncoplanar magnetic order. The learned force fields accurately reproduce electronically generated spin torques and yield nonequilibrium spin dynamics in excellent agreement with direct electronic simulations. Our results establish graph neural networks as a powerful framework for machine-learned magnetic force fields, providing a pathway toward predictive large-scale simulations of nonequilibrium magnetism across multiple length and time scales.
[LG-87] Reflected diffusion no-flux continuity equations and confined Lagrangian flows in bounded domains
链接: https://arxiv.org/abs/2607.28344
作者: Rama Cont
类目: Classical Analysis and ODEs (math.CA); Machine Learning (cs.LG); Analysis of PDEs (math.AP); Probability (math.PR)
*备注: 31 pages, 1 figure
Abstract:Motivated by marginal distribution flows of reflected diffusions in bounded domains, we investigate when a density/flux pair solving a no-flux continuity equation admits a regular Lagrangian flow that remains in the closed domain and generates the prescribed density flow. We give sufficient conditions in terms of interior bounded-variation regularity, bounded-variation control on a boundary collar, a one-sided bound on an absolutely continuous divergence, and vanishing normal trace of the velocity. The proof uses the fact that tangency removes the singular boundary contribution to the divergence of the zero extension, thereby making the extended velocity admissible for the Ambrosio-DiPerna-Lions theory. We show that these boundary assumptions cannot be jointly relaxed so as to admit a boundary current mechanism. We construct an explicit smooth density/flux pair carrying a boundary current. Its density evolution is unique in a weighted class and its characteristics are unique, confined and transport the marginals, yet it admits no regular Lagrangian flow because the compressibility bound fails arbitrarily close to the initial time. We also establish two uniqueness results for no-flux Fokker-Planck equations: a duality result for bounded measurable drifts and a weighted energy result for entrance-type drifts singular at the boundary. Our results provide a rigorous mathematical justification for using the ODE-based sampling of reflected diffusion models under minimal regularity assumptions on the coefficients, and also indicate when such ODE-based samplers may fail.
[LG-88] A Distributed Acoustic Sensing Dataset for Vessel Detection and Localization in Submarine Cable Protection
链接: https://arxiv.org/abs/2607.28306
作者: Erick Eduardo Ramirez-Torres,Javier Macias-Guarasa,Daniel Pizarro,Javier Tejedor,Sira Elena Palazuelos-Cagigas,Pedro J. Vidal-Moreno,María R. Fernández-Ruiz,Sonia Martin-Lopez,Miguel Gonzalez-Herraez,Roel Vanthillo
类目: Geophysics (physics.geo-ph); Machine Learning (cs.LG)
*备注: 19 pages, 8 figures, 3 tables. Submitted to be considered for publication as a Data Descriptor in the Scientific Data Journal
Abstract:Recent incidents of accidental damage and suspected sabotage to submarine telecommunication and power cables, particularly in the Baltic Sea, have underscored their vulnerability and the need for continuous monitoring solutions. Distributed acoustic sensing (DAS) applied to submarine optical-fiber cables enables wide-area monitoring of underwater acoustic activity. We present the Marlinks-NS DAS dataset, comprising processed submarine DAS measurements and AIS-derived vessel information curated for cable-protection research. The dataset defines two machine-learning tasks (vessel detection and vessel-to-cable distance estimation) allowing reproducible research under realistic marine conditions. The dataset contains 74,771 labeled data instances from ten days of continuous recording along a 2,554 m segment in a 28 km buried fiber-optic cable in the North Sea. Each instance includes spectral-energy features from 250 sensing channels, together with anonymized distance measurements and metadata from AIS information. The released HDF5 data, documentation, processing description, and example code support reproducible development and evaluation of DAS-based vessel-monitoring methods for submarine cable protection. Comments: 19 pages, 8 figures, 3 tables. Submitted to be considered for publication as a Data Descriptor in the Scientific Data Journal Subjects: Geophysics (physics.geo-ph); Machine Learning (cs.LG) Cite as: arXiv:2607.28306 [physics.geo-ph] (or arXiv:2607.28306v1 [physics.geo-ph] for this version) https://doi.org/10.48550/arXiv.2607.28306 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-89] Uncertainty quantification for trustworthy deep learning: Methods and measures
链接: https://arxiv.org/abs/2607.28248
作者: H. Martin Gillis,Thomas Trappenberg
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:The deployment of deep neural networks in safety-critical domains demands reliable estimates of predictive confidence, yet conventional architectures lack principled uncertainty quantification. This survey provides a structured, critical review of methods for Uncertainty Quantification (UQ) in deep learning, scoped to ensemble-based and approximate Bayesian approaches and the measures used to summarize their outputs. Relative to existing UQ surveys, our contribution is depth on efficient ensemble approximations and single-pass methods, and a unified treatment that separates the method producing a predictive distribution from the measure that summarizes its uncertainty. We organize methods into five families: Bayesian neural networks, Monte Carlo Dropout, deep ensembles, efficient ensemble approximations, and last-layer or single-pass approaches. We situate adjacent work on evidential and prior networks, conformal prediction, and post-hoc calibration, together with the decision-time tasks of out-of-distribution detection and selective prediction. For each, we examine theoretical motivation, implementation, empirical performance, and limitations. We then review ensemble diversity theory and uncertainty measures and their decompositions, contrasting the entropy decomposition with pairwise divergence measures, and consolidate evaluation methodology so that our qualitative comparisons share a common basis. We close with a brief treatment of uncertainty in large language models and open research directions, including efficient epistemic measures for classification, last-layer diversity, diversity and calibration under shift, and hybrid architectures.
[LG-90] Weather Emulators at the Frontier of Heat Extremes Predictability
链接: https://arxiv.org/abs/2607.28220
作者: Cas Decancq,Thomas Mortier,Jessica Keune,Diego G. Miralles
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Machine Learning (cs.LG)
*备注: Main: 34 pages, 5 figures. Supplementary: 29 pages, 21 figures, 11 tables
Abstract:Atmospheric predictability declines rapidly beyond the next ten days, such that forecasts at longer lead times primarily convey large-scale trends rather than specific states. Yet in a warming world, improving early warnings of extreme heat is an increasingly critical challenge. Here we evaluate six state-of-the-art deep learning weather emulators - Pangu-Weather, FuXi, ArchesWeather, AIFS, GraphCast and Aurora - alongside leading dynamical systems and statistical baselines in forecasting global near-surface temperature and extreme heat at lead times of 10-15 days. We find that several emulators rival or even surpass physics-based forecasts in deterministic temperature skill, but do so at the cost of reduced spectral fidelity, in a process widely known as blurring. While all models show some degree of predictive skill for extreme heat, most emulators under-represent peak intensities, and IFS recall is greater than that of any of the emulators. These results highlight both the emerging potential of AI to enhance extended range temperature prediction, and the remaining challenges in delivering reliable, actionable early warnings in a changing climate.
[LG-91] Meteosat Third Generation imagery improves CNN-based SSI retrieval
链接: https://arxiv.org/abs/2607.28093
作者: Gordei Pribõtkin,Piia Post,Velle Toll
类目: Atmospheric and Oceanic Physics (physics.ao-ph); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
*备注:
Abstract:Accurate Surface Solar Irradiance (SSI) estimation is increasingly important for photovoltaic energy monitoring and forecasting. The recently introduced Meteosat Third Generation (MTG) satellite constellation provides imaging data with higher spatial resolution compared to the Meteosat Second Generation (MSG) satellite constellation, but its benefits for machine-learning-based SSI retrieval have not been well established. In this work, we introduce a multi-imager and multi-resolution convolutional neural network architecture for 10-minute SSI retrieval over Northern Europe (Estonia) using MSG/SEVIRI and MTG/FCI satellite imagery together with solar-geometry and clear-sky irradiance features. Model performance is evaluated against ground-based pyranometer measurements from eight Estonian meteorological stations using site-based cross-validation and multiple training seeds. Model performance is also compared with the SARAH-3 physics-based satellite SSI product. The hybrid SEVIRI-FCI model significantly outperformed the SEVIRI-only model under overcast and cloudy conditions, reducing RMSE by 8.2 W m ^-2 and 5.7 W m ^-2 , respectively. However, under partly cloudy or clear skies, no statistically significant difference in RMSE was observed between the SEVIRI-FCI hybrid and the SEVIRI-only models. Compared with physics-based SARAH-3, the hybrid model yielded skill scores of 35 % under overcast conditions, 21 % under cloudy conditions, and 20 % overall. Furthermore, both models underperformed SARAH-3 in clear-sky conditions. These results show that higher-resolution MTG/FCI imagery improves CNN-based SSI retrieval when clouds dominate irradiance variability, but also indicate that higher spatial resolution alone is insufficient to address clear-sky limitations in machine-learning-based SSI retrieval.
[LG-92] Generalization and Trade-off in Adversarial Training: An RKHS Perspective via Kernel Integral Operators
链接: https://arxiv.org/abs/2607.27995
作者: Yiling Xie,Xiaoming Huo
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Adversarial training has emerged as a powerful approach for protecting models against adversarial attacks in a broad range of real-world applications. In this paper, we study adversarial training in the reproducing kernel Hilbert space (RKHS) framework through the associated kernel integral operator. We first derive source-uniform generalization error bounds for the RKHS adversarial training estimator in terms of the robustness level, sample size, source smoothness, and kernel spectrum. On a fixed polynomial-spectrum model, we further establish a matching lower bound showing that the optimally balanced generalization rate can be slower than the minimax prediction benchmark. This result reveals a loss of statistical accuracy in adversarial training. Our analysis shows that this loss arises from the interaction between adversarial robustness and observation noise: the noise contribution in the mixed robustness term slows the approximation rate, although the same term reduces the estimation complexity. To address this limitation, we propose a two-stage noise-debiased procedure that estimates and removes the noise contribution from the mixed term. The resulting estimator improves the generalization rate and attains the minimax polynomial rate, up to a logarithmic factor, when the robustness level is selected at the stated sample-dependent order. Our results characterize the generalization behavior of adversarial training in a nonparametric framework and provide a new interpretation and a principled solution for the trade-off between adversarial robustness and generalization. Numerical experiments support the theoretical findings and demonstrate the effectiveness of the proposed method.
[LG-93] ZAPs: A Reward Attribution Framework for DeFi Ecosystems with Adversarial-Robust Scoring via Parallel Anomaly Ensemble Detection
链接: https://arxiv.org/abs/2607.27859
作者: Girish G N,Ashutosh Sahoo,Ajay Bhat,Akshay SP,Gurukiran S,Parag Paul,Dhanashekar Kandaswamy
类目: General Finance (q-fin.GN); Machine Learning (cs.LG)
*备注: 19 pages, 5 figures, 7 tables
Abstract:Incentive programs are central to user acquisition in decentralized finance, but many reward systems rely on raw volume, transaction count, and wallet count, making them vulnerable to bots and sybil operations. We present ZAPs, a reward attribution framework that combines economic contribution scoring with adversarial robustness. A composite activity score uses protocol-specific percentile normalization to limit whale dominance while preserving differentiation among users. A two-layer weighting mechanism combines protocol share within sector and sector share within the ecosystem, which reduces the profitability of farming small protocols. We show that the maximum reward obtainable from any protocol is bounded by that protocol’s global volume share. ZAPs also introduces a four-layer defense stack consisting of transaction-level integrity checks, a parallel anomaly ensemble, post-distribution behavioral memory, and graph-based sybil clustering. The anomaly ensemble combines a one-class reconstruction model with an isolation forest and applies graduated rather than binary penalties. On 1,073 labeled malicious wallets covering 124,638 transactions, the ensemble achieves 0.923 +/- 0.013 ROC-AUC, compared with 0.891 +/- 0.016 for the reconstruction model alone, when the isolation forest is trained on benign wallets. Training it on the pooled population reverses its polarity and removes the ensemble gain. Controlled simulations reduce adversarial reward capture by 30-90 percent while legitimate-user scenarios change by 1-8 percent. Live campaigns recorded a 56 percent reduction in sybil allocation, a 49 percent increase in quality-wallet participation, and a 50 percent reduction in sell pressure. Comments: 19 pages, 5 figures, 7 tables Subjects: General Finance (q-fin.GN); Machine Learning (cs.LG) Cite as: arXiv:2607.27859 [q-fin.GN] (or arXiv:2607.27859v1 [q-fin.GN] for this version) https://doi.org/10.48550/arXiv.2607.27859 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-94] Robust Estimation of Sparse Numerical Vectors under Local Differential Privacy
链接: https://arxiv.org/abs/2607.27815
作者: Puning Zhao,Zhikun Zhang,Shaowei Wang,Sheng Yue,Bangzhou Xin,Tianhang Zheng,Pengfei Zhang,Xiaochun Cao
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Local differential privacy (LDP) protocols are vulnerable to poisoning attacks. Existing research have proposed efficient defense strategies for single-item users. However, in practice, a user may possess multiple items. The defense against poisoning attacks for multi-item users is challenging, because due to larger output spaces, the adversary can conduct more powerful attacks without being detected. In this paper, we address the robust sparse vector mean estimation problem, in which each user has a vector with m nonzero coordinates. We propose Randomized Projection with Clipping (RPC). Firstly, the server sends a random binary vector to each user. The user then projects its local data on the vector, and clip the value to restrict the attacker’s capability. To handle clipping bias, we propose a correction method based on a careful analysis that gives an exact expression of the bias. As a result, bias-variance tradeoff is no longer needed, thus the clipping threshold can be further reduced to shrink the output space and enhance robustness. We provide a rigorous theoretical guarantee of the estimation error under all possible attacks. Numerical experiments show that under trusted environments, our new method achieves comparable or better performance than existing methods, indicating that our method is already an efficient estimator in its own right. Under untrusted environments, our method is also significantly more robust to poisoning attacks.
[LG-95] Neural Network Approximation of Solutions to Fractional Parabolic Partial Differential Equations
链接: https://arxiv.org/abs/2607.27781
作者: Jae-Hwan Choi,Hyojae Lim,Jinsol Seo,Young-Jin Sim,Changhoon Song
类目: Analysis of PDEs (math.AP); Machine Learning (cs.LG)
*备注: 29 pages
Abstract:We establish a dimension-efficient neural network approximation theory for solutions to fractional parabolic equations with lower-order drift and potential terms. By introducing anisotropic spectral Barron spaces, which measure temporal and spatial regularity separately in frequency space, we first develop a dimension-independent maximal regularity theory for these equations, using dimension-independent multiplication estimates and the method of continuity to incorporate the lower-order terms. A key technical novelty is the application of the Vandermonde matrix to the global-in-time extension of the finite-time fractional heat semigroup with sufficient regularity at the initial time, thereby enabling analysis of the forward-in-time evolution via the global space-time Fourier structure of anisotropic Barron norms. We also show that a corresponding uniform-in-time estimate of the spectral Barron regularity generally fails. Finally, we derive n^-1/2 two-layer approximation bounds in mixed Sobolev norms for non-constant periodic activations and, under additional anisotropic Barron regularity, for non-periodic activations satisfying a polynomial-decay condition.
[LG-96] Error Analysis of Neural-Network-Based Engression
链接: https://arxiv.org/abs/2607.27723
作者: Juntong Chen,Zijian Guo,Xinwei Shen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注: 37 pages, 1 figure
Abstract:Engression (Shen and Meinshausen, 2024) learns a conditional distribution by fitting a generative model Y = f(X,\varepsilon) under the energy score, a strictly proper scoring rule. We provide a theoretical error analysis of engression implemented with deep neural networks. We decompose the excess risk into three components: the approximation error, the stochastic error, and the Monte Carlo error. Based on this decomposition, we establish convergence rates under the assumption that the target conditional generator admits a compositional smoothness structure.
[LG-97] Robust Wavelength Selection for Partial Least Squares Sugar Content Estimation Using Combinatorial Bayesian Optimization
链接: https://arxiv.org/abs/2607.27645
作者: Mitsunobu Kanebako,Ami S. Koshikawa,Masaru Hitomi,Takuro Tanaka,Mahito Chiba,Maiko Mori,Masayuki Ohzeki
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Wavelength selection is one of the important preprocessing methods in near-infrared spectroscopy to improve prediction accuracy and interpretability of spectral data. We formulate wavelength-region selection for sugar content estimation as a binary black-box optimization problem and propose a method based on Bayesian optimization. The proposed method constructs a sparse quadratic surrogate model and sequentially extracts interested wavelength regions by Thompson sampling. Minimizing an acquisition function is performed as a quadratic unconstrained binary optimization problem by simulated or quantum annealing. Experiments show that the proposed method improves the prediction accuracy of partial least squares regression and yields more consistent wavelength regions than genetic-algorithm-based selection and simulated annealing. Under one-bit local perturbations, the selected wavelength regions show minimal fluctuations in root mean square errors between observed and predicted values of a validation set. This local stability suggests that our method converges to a smoother error landscape and avoids isolated overfitted solutions. These results indicate that combinatorial Bayesian optimization is a useful framework for robust feature selection in spectroscopic prediction tasks.
[LG-98] HOMER: Huber-of-Means for Efficient and Robust Estimation in Hilbert Spaces
链接: https://arxiv.org/abs/2607.27532
作者: Kisung You,Boram Cho
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Heavy tails weaken high-confidence control for the empirical mean. Geometric median-of-means (MOM) also lacks a threshold that moves toward mean efficiency. We propose \emphHOMER, or Huber-of-Means for Efficient and Robust Estimation. HOMER aggregates block means through a radial Huber center. Its canonical and pseudo-Huber forms bound each block score and interpolate between median-like robustness and the empirical mean. We establish a Hilbert-space majority theorem and a MOM-order deviation bound under a finite second moment. Canonical HOMER recovers the sample mean inside its quadratic region. Pseudo-HOMER approaches the sample mean as the threshold grows. It also admits asymptotic linearity and consistent sandwich covariance estimation around the population block-Huber target. Under a finite third moment, fixed finite-dimensional projections support mean inference at the usual parametric rate. This result requires growing block sizes and counts, with block sizes increasing faster. Heavy-tailed simulations show that HOMER remains stable when a minority of block summaries is displaced. On clean Gaussian data, both versions closely approach the empirical mean’s efficiency. Finite-block sandwich intervals undercovered, especially for skewed functional data. Further studies show failure when contamination affects most blocks or compromises ordinary within-block means.
[LG-99] Emulating Cosmic Structure Formation with a Lagrangian Neural Cellular Automaton
链接: https://arxiv.org/abs/2607.27320
作者: Cooper Jacobus,Beatriz Tucci,Oliver Philcox
类目: Instrumentation and Methods for Astrophysics (astro-ph.IM); Cosmology and Nongalactic Astrophysics (astro-ph.CO); Machine Learning (cs.LG)
*备注:
Abstract:Field-level inference of cosmological initial conditions from galaxy surveys requires a forward model that is simultaneously accurate in the non-linear regime, computationally efficient, and fully differentiable. Traditional N-body simulations are accurate but computationally prohibitive for iterative inference, while approximate solvers like Lagrangian Perturbation Theory (LPT) fail to capture the knotty halo-forming dynamics of the cosmic web at late times. We introduce the \textitLagrangian Neural Cellular Automaton (LNCA), a hybrid deep learning framework that can be applied to emulate structure formation as a local, iterative dynamical process on a comoving lattice. Unlike standard Eulerian Convolutional Neural Networks (CNNs) which map fixed density fields, the LNCA operates in the Lagrangian frame, advecting the computational graph itself to follow the flow of mass. By training the network to learn only the \textitresidual displacement corrections to the Zeldovich approximation, we achieve high-fidelity emulation of the non-linear physics while guaranteeing accuracy at large scales. We further constrain our model to produce complete trajectories, not just final states, by adopting an equivariant cellular automaton architecture, which recurrently iterates on its internal states to yield a dynamic history. The resulting model is strictly local, translationally and rotationally equivariant, and naturally supports continuous time integration, making it a reliable differentiable forward model for reconstructing the initial conditions of the universe from lightcone data. Our trained model supports percent-level precision in the power and cross spectra well into the non-linear regime ( k \lesssim 0.5 , h \textMpc^-1 ), while requiring \sim10^4 times fewer learned parameters than comparable models which take the form of an interpretable internal dynamic rule set.
[LG-100] An analysis of binary isotonic regression: degrees of freedom and implications for calibration
链接: https://arxiv.org/abs/2607.27301
作者: Raphael Rossellini,Rina Foygel Barber,Zhimei Ren,Jake A. Soloff
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:Isotonic regression is a canonical tool for estimating monotone functions and calibrating probabilistic predictors. We provide a fully sharp finite-sample characterization of its worst-case degrees of freedom on binary samples. Specifically, we identify the binary sequences that maximize the number of distinct fitted values produced by isotonic regression. We develop a sharp bound on the degrees of freedom with a leading term of \frac3(4\pi^2)^1/3 n^2/3 using analytic number theory, improving on previous bounds. We then apply this result to calibration. Calibration is a central requirement for probabilistic prediction, and isotonic regression is a widely used post-processing method for improving calibration. Building on deterministic degrees-of-freedom bounds, we derive, to our knowledge, the first nontrivial distribution-free guarantee on the Expected Calibration Error (ECE) of isotonic regression. This ECE bound is fully model-free and distribution-free, only assuming Y \in \0,1\ . Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG) Cite as: arXiv:2607.27301 [stat.ML] (or arXiv:2607.27301v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2607.27301 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-101] MatCreatioNN: Machine learning-guided computational discovery of photocatalysts for environmental applications
链接: https://arxiv.org/abs/2607.27295
作者: Satya Kokonda
类目: Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
*备注: 8 figures, 0 tables. Supplementary information attached to file and data (Zenodo, Github) available
Abstract:The rational design of photocatalysts for environmental remediation and CO2 conversion remains limited by the high computational cost and sparse experimental data describing multi-parameter photocatalytic behavior. This work presents an integrated machine-learning framework that couples reinforcement learning-based metal-organic framework (MOF) generation with a multi-stage Crystal Graph Convolutional Neural Network (CGCNN) prediction funnel to identify photocatalysts optimized across multiple electronic and structural features. 120,000 MOF candidates were generated and screened using 13 key descriptors, including band-gap suitability, CO2/H2O selectivity, adsorption energy, and structural stability. The funnel approach reduced computational cost by 4.13-fold while maintaining predictive robustness. Two top candidates, a Cr-based and a Zn-based MOF, exhibited predicted photocatalytic fitness values of 1.70 +/- 0.25 and 1.20 +/- 0.05 fold higher respectively than benchmark materials such as PCN-224(Zr), demonstrating simultaneous improvements in light absorption, redox energetics, and framework durability. Simulated X-ray diffraction patterns confirmed strong structural agreement with experimentally synthesized MOFs, indicating high synthesizability. Post-hoc analysis revealed recurring structural motifs, such as the N262 metal cluster, that correlated strongly with high predicted photocatalytic activity. These results highlight the potential of data-driven methods to accelerate discovery of efficient and durable photocatalysts for environmental and energy-related transformations, providing a foundation for experimental realization and large-scale implementation of computationally designed MOFs.
[LG-102] Expected Survival-Time Bounds for Robust Optimization Over Time under Isotropic Gaussian Dynamics
链接: https://arxiv.org/abs/2607.27280
作者: Pavel Novoa-Hernández
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR)
*备注:
Abstract:Robust Optimization Over Time (ROOT) is a recent branch of evolutionary dynamic optimization that seeks solutions capable of remaining effective across multiple consecutive environments. Unlike the traditional track-the-moving-optimum (TMO) paradigm, which reoptimizes after every environmental change, ROOT explicitly values persistence. Although the field has grown considerably, most contributions remain algorithmic and empirical, leaving several fundamental properties poorly understood from a theoretical perspective. One such property is survival time, defined as the number of future environments in which a deployed solution continues to satisfy a prescribed quality threshold. While survival time is widely used as a measure of temporal robustness, little is known about how its expected value depends on environmental dynamics, deployment quality, or problem characteristics. This paper studies expected survival time for a fixed deployed solution under isotropic Gaussian environmental dynamics. Modeling survival as a discrete first-exit problem, we derive a rigorous lower bound and a computable multi-step upper bound. The analysis shows that expected survival scales as \Theta(\sigma^-2) in slowly varying environments and approaches its minimum value of one future change in high dimensions. A comprehensive Monte Carlo study validates the theoretical predictions, examines sensitivity to modeling assumptions and parameter uncertainty, and illustrates how the bounds can support deployment decisions after optimization. The resulting framework provides an analytical characterization of deployment lifetime and identifies when a required deployment horizon can be guaranteed, ruled out, or remains analytically unresolved.
[LG-103] PlantBGC: Transformer for Plant BGC Discovery via Label-Free Domain Adaptation and Weak Supervision
链接: https://arxiv.org/abs/2607.27258
作者: Yuhan Zhao,Nidhi Grover,Zhishan Guo,Ning Sui
类目: Genomics (q-bio.GN); Machine Learning (cs.LG)
*备注:
Abstract:Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale. Existing plant BGC mining tools are largely signature- and rule-driven and do not fully leverage recent advances in contextual representation learning for modeling long-range domain context and controlling false positives under strong domain shift. We seek an AI-assisted workflow that narrows experimental search space by transferring supervision from well-annotated microbial BGCs to plant genomes. We present PlantBGC, representing genomes as ordered Pfam-domain sequences and learning BGC-likeness with an encoder-only Transformer trained on MIBiG microbial BGCs and adapted to plants via label-free masked language modeling. On microbial benchmarks, PlantBGC achieves token-level AUC = 0.988 (10-fold CV) and 0.979 (leave-class-out). On plants, adaptation improves known-BGC recovery on n = 34 curated loci under strict 100% coverage, increasing recovery from 29.4% to 67.6% and indicating more complete boundaries. GO/KEGG-derived weak supervision reduces proxy primary-like ratio by 48.40% (GO) and 45.20% (KEGG), with consistent per-species reductions (paired Wilcoxon p = 1.53e-5). Compared to plantiSMASH, PlantBGC yields more compact loci on matched regions (median length ratio = 0.278; 93.8% of pairs are shorter).
[LG-104] More Data Worse Decisions? Preference Reversals in Neural Networks under Gram Incompatibility
链接: https://arxiv.org/abs/2607.27255
作者: Yanli Yan,Yuanzheng Li,Yong Zhao,Hongbo Guo,Shoudong Han
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 15 pages
Abstract:Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization. This raises a reliability question: whether a model refitted on pooled data preserves an action ordering supported by both sources. Case-Based Decision Theory (CBDT) formalizes this requirement through its composition axiom, which requires source-supported preferences to survive their union. We study when this property holds for fixed-representation neural networks with ordinary least squares (OLS) output heads. First, we show that pooled refitting recomputes the inverse-Gram geometry used to weight source evidence, which can reverse shared preferences, and derive exact and approximate preservation conditions. Next, we introduce a scale-invariant Gram mismatch measure for prioritizing candidate pools and geometry-oriented regularization for shaping source geometry during training. Finally, we develop a three-stage audit that traces strict pairwise reversals through decision changes to task-defined utility loss. Experiments spanning a load-based bidding proxy and medical and financial decision proxies reveal stable and reversal-prone pooling regimes: the load audit identifies a measurable nonzero class of source-consensus-relative harmful decisions under the proxy utility, while cross-domain audits show that comparable mismatch can correspond to sharply different preservation rates. Geometry-oriented objectives occupy distinct descriptive accuracy-consistency-geometry-harm operating points. Together, the framework makes compositional reliability measurable and operational through screening, analytic certification, geometry-oriented training, and decision-consequence auditing.
[LG-105] Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry
链接: https://arxiv.org/abs/2607.27224
作者: Aakash Bhagat,Shashank Choudhary
类目: Applications (stat.AP); Machine Learning (cs.LG)
*备注:
Abstract:External and synthetic control arms (ECAs) are entering psychiatric drug development, but the field lacks a benchmark that evaluates the properties regulators care about: not only how accurately a method reconstructs untreated trajectories, but whether its uncertainty is calibrated, whether it is robust to the informative observation times common in mental-health records (sicker patients are seen more often), and what false-positive rate it induces in go/no-go trial decisions. Real psychiatric trial data (e.g. STAR-D and registry cohorts) require credentialed access and lack ground-truth counterfactuals, so, following established semi-synthetic benchmarks in causal inference (IHDP, ACIC, and the PK-PD tumor-growth simulator), we release Psych-ECA, a fully reproducible generator of longitudinal symptom trajectories for depression (PHQ-9), anxiety (HAM-A), and psychosis (PANSS) with known counterfactual control arms, informative visits, and validated-scale measurement noise. We benchmark eight estimators spanning carry-forward, pooled real-world-data averages, nearest-neighbour matching, linear mixed models, gradient boosting, and the Scribe trajectory-bridge method. Three findings emerge. First, trajectory and flexible machine learning methods achieve the best counterfactual accuracy (about 2.3 PHQ-9 RMSE), outperforming cross-sectional baselines. Second, only Scribe is both accurate and calibrated, achieving 93-96% empirical coverage of nominal 90% prediction intervals, compared with 87-88% for gradient boosting and 62-75% for uncalibrated SDE models. Third, inverse-intensity correction reduces bias under informative sampling, while Scribe’s calibrated intervals are the only trajectory method that maintains nominal false-positive rates as informativeness increases. We release all code, data-generation scripts, and random seeds to enable fully reproducible evaluation.
[LG-106] Foundation-Model Earth Representations Enable Regional-Scale Forest Aboveground Biomass Monitoring Across the Northeastern United States
链接: https://arxiv.org/abs/2607.27217
作者: Shashika Lamahewage,Chandi Witharana
类目: Applications (stat.AP); Machine Learning (cs.LG)
*备注:
Abstract:Forest aboveground biomass (AGB) is a critical indicator of ecosystem productivity and terrestrial carbon storage, yet regional carbon monitoring remains constrained by the sparse spatial and temporal availability of field inventories and airborne structural measurements. Recent Earth observation foundation models provide globally consistent geospatial representations derived from diverse multimodal datasets, offering a potential pathway toward scalable biomass monitoring. Here, we evaluate Google Satellite Embeddings (GSE), generated by the AplphaEarth Foundation Model, for regional-scale AGB estimation across diverse temperate forest ecosystems in the northeastern United States. We integrated annual GSE observations, airborne LiDAR, and continuous forest inventory measurements from the Northeastern Forest Inventory Network (NEFIN) within a machine-learning framework. Combined LiDAR-GSE models achieved an R^2 of 0.79 for AGB estimation. Capitalizing on annual GSE observations expanded the training dataset by more than tenfold through temporal growth adjustment, increasing predictive performance to R^2 = 0.82 while reducing model bias by over 70%. Spatial autocorrelation analyses showed that integrating foundation-model representations and structural predictors substantially reduced residual spatial dependence. Monte Carlo simulations demonstrated that hyperparameter optimization reduced model-performance variability by 27.9%. Our findings demonstrate that foundation-model Earth representations capture ecologically meaningful information relevant to forest biomass and provide a scalable framework for annual carbon monitoring in regions with incomplete airborne LiDAR coverage. Our fundings establish a pathway toward next-generation forest carbon assessment based on globally available foundation-model Earth observations.
附件下载


