本篇博文主要内容为 2026-08-07 从Arxiv.org论文网站获取的最新论文列表,自动更新,按照NLP、CV、ML、AI、IR、MA六个大方向区分。
说明:每日论文数据从Arxiv.org获取,每天早上12:30左右定时自动更新。
提示: 当天未及时更新,有可能是Arxiv当日未有新的论文发布,也有可能是脚本出错。尽可能会在当天修复。
目录
概览 (2026-08-07)
今日共更新659篇论文,其中:
- 自然语言处理共94篇(Computation and Language (cs.CL))
- 人工智能共203篇(Artificial Intelligence (cs.AI))
- 计算机视觉共133篇(Computer Vision and Pattern Recognition (cs.CV))
- 机器学习共162篇(Machine Learning (cs.LG))
- 多智能体系统共13篇(Multiagent Systems (cs.MA))
- 信息检索共15篇(Information Retrieval (cs.IR))
- 人机交互共32篇(Human-Computer Interaction (cs.HC))
多智能体系统
[MA-0] AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
【速读】:该论文旨在解决在不完全信息博弈中评估两个智能体(agent)相对强度时,如何实现高效、可审计的早期停止问题。传统固定预算评估存在过早终止或过度消耗资源的问题,而朴素的可选停止方法则会破坏置信水平的有效性。其解决方案的关键在于将方差缩减技术与持续监控的置信序列(Confidence Sequences, CSs)相结合,提出一种任意时间有效生成式评估工具(Anytime-Valid AIVAT, AV-AIVAT)。AV-AIVAT通过引入动作感知价值评估工具(Action-Informed Value Assessment Tool, AIVAT)对博弈结果进行条件均值为零的偏差校正,显著降低方差(在15个大语言模型代理配置的71,439局头对头无限制德州扑克中,方差平均降低54倍),并结合在线学习的、仅依赖历史游戏数据的估值模型,确保每局游戏的校正不依赖自身得分,从而保持统计有效性。进一步地,该方法采用渐近置信序列(AsympCS)和经验伯恩斯坦置信序列(EB-CS)分别实现渐近筛选与精确有限样本认证,前者在名义95%置信水平和±1大盲注精度下,使原始结果所需的游戏数量减少至AIVAT校正结果的1/74;后者通过结构化推导出校正收益的独立有界性,在描述性HUNL实验中实现了1.37倍的停机时间提升。因此,AV-AIVAT将方差缩减转化为可验证的早期停止能力,同时分离了渐近分析与精确认证过程,使评估可在证据充分时立即终止,并向第三方提供完整可复现的验证依据。
链接: https://arxiv.org/abs/2608.06362
作者: Boning Li,Yu Chen,Longbo Huang
机构: IIIS, Tsinghua University (清华大学智能技术与系统国家重点实验室); Tsinghua University (清华大学)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 34 pages, 5 figures
Abstract:Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be told apart, while naive optional stopping with an ordinary confidence interval invalidates the stated level. We make such an evaluation stop as soon as its evidence suffices, with the guarantee intact. The Action-Informed Value Assessment Tool (AIVAT) reduces variance in imperfect-information games through conditional mean-zero corrections, by a median 54\times across 15 LLM agent configurations spanning 71,439 paired Heads-Up No-Limit Hold’em (HUNL) hands, but does not say when to stop. We combine AIVAT with continuously monitored Confidence Sequences (CSs) into anytime-valid AIVAT (AV-AIVAT), whose online value model learns only from past games so that no game scores its own correction. At the nominal 95% level and a target precision of \pm1 Big Blind, raw outcomes need a median 74\times as many hands as AIVAT-corrected outcomes to stop under the Asymptotic CS (AsympCS). Exact finite-sample certification uses the Empirical-Bernstein CS (EB-CS), which needs an independently justified bound on corrected payoffs. We establish such a bound structurally for Leduc hold’em and characterize a width floor set by the CS’s bet cap and that bound, which governs how much of a variance gain becomes earlier stopping; the descriptive HUNL EB-CS runs show a median 1.37\times stopping-time ratio. AV-AIVAT turns variance reduction into efficient, auditable early stopping while separating asymptotic screening from exact certification, so an evaluation can stop the moment its evidence suffices and hand a third party everything needed to recheck the verdict at that very stopping time.
[MA-1] Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents
【速读】:该论文旨在解决已部署人工智能(AI)代理在持续参与式治理中如何实现安全、可信且自强制授权的问题。其核心挑战在于如何通过可验证的资源分配机制,使治理决策不仅具备合法性,还能在技术层面自动执行,从而避免对中心化权威的依赖。解决方案的关键在于构建一个基于计算预算(compute budget)的治理框架:通过将治理权与一种独立于代理自身计算资源的“治理货币”相分离,利用广度加权的有效支持聚合机制,结合具有滞后效应的双阈值门控系统,将净支持转化为二元授权状态;该授权状态通过一个由外部认证的安全上限耦合的映射函数,最终以硬件级签名的计算许可证形式释放计量计算资源,实现治理决策的自强制执行。该机制确立了“计算即治理杠杆”的安全AI范式,同时将“被治理者操纵治理选民”识别为该框架下的核心开放问题,并提出了若干应对此类操纵行为的挑战。
链接: https://arxiv.org/abs/2608.06353
作者: Praphul Chandra,Sujit Gujar,Ganesh Ghalme
机构: Atria University(阿特里亚大学); IIIT Hyderabad(海德拉巴国际信息技术研究所); IIT Hyderabad(海德拉巴印度理工学院)
类目: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: 22 pages, 9 Figures
Abstract:We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism seeks to establish the Safe AI paradigm that compute is an effective governance lever. We situate our work as a compliance or commons overlay on a deployer. One governance period is an extensive form game in which verified human stakeholders arrive sequentially and contribute, on a provision or a rejection market, in a governance currency that is deliberately distinct from the agents compute. A funding aggregator turns raw contributions into breadth weighted effective supports - a two threshold gate with hysteresis converts net support into a binary authorization that, through a coupling map bounded by an exogenously certified safety ceiling, releases a metered compute budget - realized in hardware as a signed compute license so that the decision is self-enforcing. We characterize the class of agents the mechanism can govern and isolate manipulation of the governing electorate by the governed agent as the central open problem. We also introduce several challenges addressing manipulation of governing electorate by the governed agents.
[MA-2] From Siloed Algorithms to Compliance-First Agent ic Platforms: A Multi-Layered Architecture for Hospital AI Systems
【速读】:该论文旨在解决当前医疗机构在部署人工智能(AI)过程中普遍存在的“孤岛式”应用问题,即各类AI系统分散于不同科室、缺乏统一治理与集成,导致重复开发、数据割裂、合规风险隐匿以及企业级价值无法实现。尽管医疗AI市场快速增长且投资持续增加,但约70%-80%的试点项目难以规模化落地,其核心瓶颈在于治理缺失、数据碎片化及缺乏可复用的集成架构。为此,本文提出一种面向医院场景、以合规为先的生成式AI(Agentic AI)架构,其关键创新在于构建多层互操作体系:一是通过代理编排层(Agent Orchestration Layer)实现临床、运营与财务领域跨域多智能体协同工作流;二是设立合规与策略层(Compliance and Policy Layer),将政策即代码(Policy-as-Code)机制集成至HIPAA、GDPR、欧盟AI法案、印度DISHA法案、DPDP法案及ISO/IEC安全标准等全球监管要求;三是引入隐私保护数据底座(Privacy-Preserving Data Fabric),在真实医院信息系统(HIMS)流程中嵌入联邦学习、差分隐私与安全飞地技术,保障数据可用不可见。基于合成但结构真实的医院数据集与可直接部署的原型系统,研究验证了从分诊风险预测、工作流优化到合规日志记录的端到端协同能力,在显著降低任务处理时延与人工文档负担的同时,确保受策略约束的数据访问。该架构为医院管理者提供了一套可适配本地部署、混合云与云原生环境的实用蓝图,推动AI应用从零散工具向受控、合规、可衡量投资回报率(ROI)的统一平台演进。
链接: https://arxiv.org/abs/2608.06112
作者: Manideep Dhar,Ritwik Singh,Sharat Chandra Kumar Manikonda
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: Peer-reviewed published article
Abstract:Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point solutions locked inside departmental silos, resulting in duplicated effort, hidden risks, and unrealized enterprise value. Despite explosive growth of AI in healthcare market and accelerating investment, an estimated 70-80% of healthcare AI pilots fail to scale, largely due to governance gaps, fragmented data, and missing integration blueprints. This research proposes a hospital-specific, compliance-first, Agentic AI architecture with multiple interoperable layers, extending existing hospital AI platform models with: (i) an Agent Orchestration Layer for multi-agent workflows across clinical, operational, and financial domains, (ii) a Compliance and Policy Layer that centralizes policy-as-code for HIPAA, GDPR, the EU AI Act, DISHA Act, India’s DPDP Act, and ISO/IEC security and safety standards, and (iii) a Privacy-Preserving Data Fabric that plugs federated learning, differential privacy, and secure enclaves into real-world Hospital Information Management System (HIMS) flows. Using a synthetic but structurally realistic hospital dataset and an open, ready-to-deploy prototype implementation, this study demonstrates the end-to-end orchestration of triage risk prediction, workflow optimization, and compliance logging, achieving substantial simulated reductions in task turnaround times and manual documentation effort while maintaining policy-guarded data access. The resulting architecture offers hospital leaders a pragmatic blueprint to move from ad hoc tools to a governed, globally compliant, ROI-focused AI platform that can be tailored to on-premise, hybrid and cloud-native deployments.
[MA-3] Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
【速读】:该论文旨在解决生成式AI在北极航运导航应用中因海冰条件快速变化而面临的可靠性问题,核心挑战在于如何构建可解释且对环境变化具有鲁棒性的奖励模型(reward model)。现有逆强化学习(Inverse Reinforcement Learning, IRL)方法虽能从船舶轨迹中推断奖励函数,而近期的元逆强化学习(meta-IRL)通过引入隐含上下文变量(latent context variables)以捕捉行为异质性,但其关键疑问在于:这些隐变量是否真正揭示了未观测到的隐藏偏好,还是仅对已有可观测状态信息的重新编码。研究基于3,186条来自202艘船舶的AIS航行数据,在九个北极航运季节中对比了线性共享奖励、非线性共享奖励与基于相同非线性架构的隐含上下文模型。结果显示,非线性共享奖励相较于线性基线在保留样本似然上提升50.9%,而引入船舶特异性隐含上下文后性能反而下降16.5%。通过行为分析、上下文探针及预注册的特征隐藏消融实验表明,看似船舶级别的行为差异主要由可观测的航路与环境条件所解释,而非隐藏的船舶特异性因素。此外,预测精度、航路保真度与奖励迁移能力等不同评估指标导致模型排名不一致,说明单一指标不足以全面评价学习到的奖励函数。研究结论强调,在引入每艘船的隐含上下文之前,应优先验证可观测的航路、环境与船舶特征是否已充分解释行为变异,从而为安全关键领域中更可信的生成式AI部署提供依据。
链接: https://arxiv.org/abs/2608.06105
作者: Vaishnav Vaidheeswaran,Dilith Jayakody,Biruk Ambaw,Jaswanth Kumar,Md Mahbub Alam,Gabriel Spadon
机构: Google(谷歌); Stanford University (斯坦福大学)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deployment requires reward models that are interpretable and robust to changing environments. Inverse reinforcement learning (IRL) provides a framework for recovering such rewards from vessel trajectories, while recent meta-IRL methods introduce latent context variables to capture behavioral heterogeneity. However, it remains unclear whether these latent representations recover genuinely hidden preferences or simply re-encode information already available in the observed state. We conduct a controlled evaluation on 3,186 AIS-derived voyages from 202 vessels across nine Arctic shipping seasons, comparing a linear shared reward, a nonlinear shared reward, and a latent-context model built on the same nonlinear architecture. The nonlinear reward improves held-out likelihood by 50.9% over the linear baseline, whereas adding vessel-specific latent context reduces performance by 16.5%. Behavioral analysis, context probes, and a pre-registered feature-hiding ablation show that apparent vessel-level variation is largely explained by observable route and environmental conditions rather than hidden vessel-specific factors. Moreover, predictive accuracy, route fidelity, and reward transfer yield different model rankings, demonstrating that no single metric is sufficient to evaluate learned rewards. These findings motivate testing whether the observed route, environmental, and vessel features already explain behavioral variation before adding per-vessel latent context. This supports more trustworthy AI deployment in safety-critical domains.
[MA-4] ASGE-RR: Agent ic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls
【速读】:该论文旨在解决生成式 AI (Generative AI) 代理工作流在执行过程中因远程调用模型、内存存储和工具等资源而产生的动态依赖关系管理问题。由于这些依赖调用通常仅在运行时才被揭示,导致当前调用的资源分配可能占用后续高价值工作流所需的资源,从而影响整体服务效率与完成率。为此,论文提出将此问题建模为在线网络控制问题——代理服务图嵌入(Agentic Service Graph Embedding, ASGE),其核心目标是在容量、成本和截止时间约束下,将运行时动态暴露的工作流调用映射到服务副本与网络路径。解决方案的关键在于提出 ASGE-RR,一种具备可修订预留机制的在线控制器:它基于对工作流未来延续的预测,提前保护潜在高价值调用所需的资源,并在获得新执行信息后动态更新预留策略。实验结果表明,在小规模控制环境与广域网(WAN)测试环境中,尽管调用结构具有不确定性,但通过利用运行时揭示的控制点,ASGE-RR 能够比滚动时域控制器和仅关注当前调用的引导控制器多完成高达10%的工作流价值,验证了在动态工作流中预留资源以保障未来调用可行性这一新网络控制范式的有效性。
链接: https://arxiv.org/abs/2608.06033
作者: Trond Vatten,Yuming Jiang
机构: Norwegian University of Science and Technology (NTNU) (挪威科技大学)
类目: Multiagent Systems (cs.MA); Performance (cs.PF)
备注:
Abstract:AI-agent workflows often involve remote calls to models, memory stores, and tools distributed across a network. As execution progresses, these dependency calls collectively form an agentic service graph (ASG). Unlike traditional service requests, many dependency calls are revealed only at runtime. Consequently, allocating resources to a currently visible call may consume capacity later needed by a call from a higher-value workflow. We formulate this challenge as Agentic Service Graph Embedding (ASGE), an online network-control problem that maps runtime-revealed workflow calls to service replicas and network paths under capacity, cost and deadline constraints. We present ASGE-RR, an online ASGE controller with revisable reservations. ASGE-RR protects capacity for likely future calls while enforcing the constraints. ASGE-RR evaluates candidate replica-and-path mappings against predicted workflow continuations and updates reservations as new execution information becomes available. We evaluate ASGE-RR using OpenHands and GPT Researcher workflows executed with gpt-5.6-luna and replayed over in two complementary experimental environments, a controlled Docker testbed and a WAN testbed. The investigation shows that all the evaluated AI-agent tasks expose at least one runtime-revealed dependency call that can be steered before connection establishment. Exploiting this control point, even though the experimental environments are small-scale, ASGE-RR already demonstrates noticeable potential: It completes (up to) 10% more workflow value than a same-information rolling-horizon controller and a current-call steering controller on the WAN testbed. The results suggest that runtime-revealed workflow structure creates a new network control opportunity: protecting resources for likely future calls allows more AI-agent workflows to finish in time. Subjects: Multiagent Systems (cs.MA); Performance (cs.PF) Cite as: arXiv:2608.06033 [cs.MA] (or arXiv:2608.06033v1 [cs.MA] for this version) https://doi.org/10.48550/arXiv.2608.06033 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[MA-5] Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference
【速读】:该论文旨在解决仿真闭环决策系统中强化学习(Reinforcement Learning, RL)推理因仿真器端执行开销导致的性能瓶颈问题,尤其针对工作负载高度动态且对运行时线程配置敏感的场景。现有多线程策略在执行前或执行过程中难以及时匹配线程资源,引发资源竞争、调度开销增加及吞吐量下降。其核心解决方案在于提出一种混合自适应线程调优方法——AutoThread,其关键创新点在于:通过实证分析识别出任务执行时间与调度时间之比是决定最优线程数的关键因素;进而采用物理信息神经算子(Physics-Informed Neural Operator, PINO)作为线程数预测模型,并引入有限源M/M/1队列模型对预测过程进行约束与引导,实现对动态工作负载下线程配置的快速、精准估计;同时结合负载感知的在线微调机制以补偿预测误差并优化资源分配。实验结果表明,相较于静态策略,AutoThread平均提升加速比18.4%,吞吐量分别达到XGBoost和Reinforcer的1.7倍和1.8倍,且相比最先进方法可将执行时间最多减少83.8%。
链接: https://arxiv.org/abs/2608.06025
作者: Jiming Su,Hantao Hua,Lujia Yin,Yiping Yao,Feng Zhu
机构: National University of Defense Technology (国防科技大学)
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA); Performance (cs.PF)
备注:
Abstract:In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to runtime thread configurations. Existing multithreaded strategies struggle to match thread resources before or during execution, causing resource contention, scheduling overhead, and reduced throughput. Through empirical analysis, we identify the ratio of task execution time to scheduling time as the key factor determining the optimal thread count. Building on this insight, we propose AutoThread, a hybrid adaptive thread-tuning method for mitigating simulation bottlenecks in RL inference. AutoThread employs a Physics-Informed Neural Operator (PINO) as a thread-count predictor and incorporates a finite-source M/M/1 queueing model to constrain and guide prediction, enabling fast and accurate estimation under dynamic workloads. It further performs load-aware online fine-tuning to compensate for prediction errors and refine resource allocation. Experiments show that AutoThread improves average speedup by 18.4% over static strategies, achieves average throughput of 1.7x and 1.8x that of XGBoost and Reinforcer, respectively, and reduces execution time by up to 83.8% compared with state-of-the-art methods. Our code and dataset are publicly available at this https URL.
[MA-6] Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis
【速读】:该论文旨在解决大规模语言模型(Large Language Model, LLM)代理集体协作中存在“黑箱”问题,即尽管多智能体通过辩论与投票机制显著提升了任务准确性,但系统层面缺乏对推理过程收敛性的原理性检验、收敛轮次的理论边界以及决策驱动因素的可信追溯。其解决方案的关键在于构建基于柯普曼算子(Koopman operator)理论的新型分析框架,将多智能体集体视为通信图上的非线性动力系统,并通过从交互轨迹中估计的柯普曼转移算子的谱特性,实现对系统行为的精确线性表征。该谱分析提供了三个可机器验证的证书:次主导特征值λ₂决定了推理的内在时间尺度,可预先计算出收敛截止时间;其对应的特征向量揭示了集体推理中稳定的共识群体结构,|λ₂|则用于验证该解释的有效性;而主导谱坐标构成一个压缩且可审计的消息基,支持高保真度的决策溯源。实验在注意力-共识模型上验证了该方法的优越性能:截止时间与实际收敛行为呈现0.93的对数-对数相关性,在24种配置中96%情况下提供有效上界;当谱分析认证系统处于亚稳态时,决策归因完全准确;32个谱坐标中有8个在99.7%保真度下保留最终决策;此外,仅需15次辩论即可学习到适用于60/60未见辩论的通用验证证书。整个分析过程可在单个CPU上分钟级完成,使谱学认证成为可实用的可信集体推理保障层。
链接: https://arxiv.org/abs/2608.05956
作者: Nuzhat Khan,Indrakshi Dey
机构: Universiti Teknologi Malaysia (马来西亚理工大学); South East Technological University (东南科技大学)
类目: Multiagent Systems (cs.MA); Systems and Control (eess.SY)
备注:
Abstract:Orchestrated collectives of large language model (LLM) agents that debate and vote are an emerging form of computational intelligence: the intelligent behaviour resides in the \emphinteraction, not in any single agent. They improve task accuracy, yet remain black boxes at the system level: there is no principled test of convergence, no bound on the rounds needed, and no faithful account of what drove a decision. This paper develops a novel framework based on Koopman operator theory and validates its theoretical guarantees on multi-agent consensus dynamics. Treating the collective as one nonlinear dynamical system on a communication graph, we read its essential behaviour off the spectrum of its Koopman transfer operator, an exact linear representation of the nonlinear dynamics estimated from interaction traces. The spectrum yields three machine-checkable certificates: the sub-dominant eigenvalue \lambda_2 fixes the intrinsic timescale of reasoning and yields a convergence deadline computable \emphbefore the debate runs; its eigenvector names the coherent factions the collective reasons in, and |\lambda_2| certifies when that explanation is valid; and the leading spectral coordinates form a compressed, auditable message basis. On an attention-consensus model, the deadline tracks observed convergence with log–log correlation 0.93 and bounds it in 96% of 24 configurations; attribution is exact whenever the spectrum certifies metastability; eight of 32 coordinates preserve the decision at 99.7% fidelity; and a certificate learned from 15 debates held on 60/60 held-out debates. The study runs in minutes on a CPU, making spectral certification a practical layer for trustworthy collective reasoning.
[MA-7] A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems ICML2026
【速读】:该论文旨在解决大语言模型(Large Language Model, LLM)驱动的多智能体系统在推理过程中因频繁调用模型和复杂协调机制导致的高延迟、高计算开销及精度受限问题。其核心挑战在于如何有效组织与协调不同形式的并行策略以提升推理效率。解决方案的关键在于提出TIPEX——一个可调控的执行框架,将多智能体系统中的并行性统一建模为两个层次的决策过程:任务级的副本并行(Replica Parallelism),用于探索多个完整的解决方案路径;以及路径内的结构并行(Structural Parallelism),通过任务分解实现单条路径上的并发执行。TIPEX通过统一的执行语义,实现了两种并行方式的协同调度与可控组合,并支持对不同并行策略及其参数配置的系统性分析。实验基于GAIA基准测试表明,推理时的并行化可显著提升准确率并降低端到端延迟,但伴随令牌消耗增加;进一步分析揭示,副本并行与结构并行在任务复杂度适中时表现出互补效应,而过度激进的并行策略未必带来性能增益。
链接: https://arxiv.org/abs/2608.05791
作者: Zihan Xu,Haolin Tian,Hai Jiang
机构: 未知
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
备注: Accepted to ICML 2026
Abstract:Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to improve inference-time efficiency. From the perspective of inference-time execution, this paper models parallelism in multi-agent systems as two distinct levels of decision processes: Replica Parallelism, which explores multiple complete solution paths at the task level, and Structural Parallelism, which enables concurrent execution within a single solution path through task decomposition. However, the roles of different forms of parallelism and their interrelationships still lack systematic study in terms of unified organization and coordination. We therefore propose TIPEX, a controllable execution framework that unifies these two levels of parallelism and coordinates their roles within the inference process under a unified execution semantics while supporting systematic combinations and analyses of different parallel strategies and parameter configurations. Systematic experiments on the GAIA benchmark demonstrate that inference-time parallelism can significantly improve accuracy and reduce end-to-end latency at the cost of increased token consumption. Further analysis shows that Replica and Structural Parallelism exhibit complementary effects across task complexities, with tasks of intermediate difficulty benefiting most from their coordination, while overly aggressive parallel strategies do not necessarily yield better performance.
[MA-8] F2Agent : Financial Fusion of Agent ic Intelligence for Multimodal Trading
【速读】:该论文旨在解决多模态金融数据融合中因模态间细粒度依赖关系捕捉不足、融合机制低效以及对市场噪声敏感而导致的交易信号质量下降问题。其核心挑战在于现有基于大语言模型(Large Language Model, LLM)的代理方法在多模态建模能力、跨模态信息融合效率及鲁棒性方面存在显著局限。为此,论文提出F² Agent,一种由金融智能代理融合驱动的新型多模态智能体范式。其关键创新在于:首先构建分层专业化代理以全面提取各模态特异性信号;进而引入模态感知的自适应融合机制与噪声鲁棒的一致性正则化策略,动态捕获细粒度跨模态依赖关系,并生成具备强抗噪能力的交易信号。实验结果表明,F² Agent在六只股票及加密货币资产上均显著优于16个基准方法,在多个交易指标上实现平均超过20%的年化收益率相对提升,尤其在GOOG和TSLA上分别取得120.48%和148.41%的优异表现,验证了其在复杂市场动态下的有效性与鲁棒性。
链接: https://arxiv.org/abs/2608.05668
作者: Changshuo Liu,Yanzheng Jin,Shangfeng Cai,Peng Fang,Xiaokui Xiao,Beng Chin Ooi
机构: National University of Singapore(新加坡国立大学); Huazhong University of Science and Technology(华中科技大学); Zhejiang University(浙江大学)
类目: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: 32 pages, 12 figures, 19 tables
Abstract:With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-quality financial trading. Although recent advancements in Large Language Model (LLM)-based agents have enabled the ingestion of multimodal inputs, existing methods fail to capture nuanced cross-modal dependencies and remain vulnerable to market noise, due to limited multimodal modeling, ineffective fusion mechanisms, and inadequate robustness. To address these challenges, we propose F ^2 Agent, a novel multimodal agentic paradigm driven by the Financial Fusion of Agentic Intelligence. F ^2 Agent first deploys a hierarchy of specialized agents to comprehensively extract modality-specific signals. It further introduces a modality-aware adaptive fusion mechanism coupled with noise-robust consistency regularization to dynamically capture fine-grained inter-modality dependencies and generate noise-resilient trading signals. Extensive experiments on six stocks and cryptocurrency assets demonstrate that F ^2 Agent consistently outperforms 16 competitive baselines across multiple trading metrics, with over 20% relative improvement in annualized return on average. Notably, F ^2 Agent delivers returns of 120.48% on GOOG and 148.41% on TSLA, demonstrating its efficacy and robustness in varying market dynamics.
[MA-9] Search-Aided Joint Agent -Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
【速读】:该论文旨在解决现实场景中持续性多智能体路径规划(Lifelong Multi-Agent Path Finding, LMAPF)所面临的挑战,特别是现有学习型规划方法普遍依赖过于简化的运动学假设,忽视了实际应用中关键的运动约束,导致在复杂环境中的性能受限。为此,论文提出一种更贴近真实自动化仓库系统的LMAPF模型——LMAPF-R2,其引入了鲁棒的安全约束与原地旋转约束,显著提升了任务协调难度,尤其在高密度、空间受限环境中。为应对这一挑战,论文提出了一种基于搜索辅助的联合强化学习框架(Search-Aided Joint Reinforcement Learning, SJRL)。其核心创新在于:首先,通过将因果PIBT(Causal PIBT)——一种单步搜索式规划器——嵌入神经策略中,以实时解决智能体间的碰撞并传播其意图;其次,构建统一的强化学习框架,联合优化智能体与环境策略,其中环境策略通过学习图边权值,利用反向Dijkstra搜索提供全局移动引导。实验表明,SJRL在多个高密度地图上显著优于强基准搜索算法Causal-PIBT;进一步在混合现实仓库环境中,结合8个物理机器人与248个虚拟机器人验证了其在真实世界部署中的有效性。
链接: https://arxiv.org/abs/2608.05588
作者: He Jiang,Jingtian Yan,Yulun Zhang,Yimin Tang,Tanishq Duhan,Rishi Veerapaneni,Guillaume Sartoretti,Jiaoyang Li
机构: Google(谷歌); Stanford University (斯坦福大学); Meta(元)
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注:
Abstract:Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones. While many learning-based planners have been proposed for LMAPF, most rely on oversimplified kinematic assumptions that may overlook motion constraints critical to real-world performance. In this work, we study a more realistic LMAPF model derived from many real-world automated warehouse systems, termed LMAPF-R2, which incorporates robust safety constraints and in-place rotation constraints. These constraints substantially increase coordination difficulty, particularly in highly constrained spaces. To address these challenges, we propose Search-Aided Joint Reinforcement Learning (SJRL). We first augment neural policies with Causal PIBT, a single-step search-based planner that resolves agents’ collisions and propagates their intentions. We then introduce a unified RL formulation that jointly optimizes agent and environment policies, where the environment policy learns graph edge costs to provide global movement guidance via backward Dijkstra search. Experiments demonstrate that SJRL achieves significant improvements over the strong search-based planner, Causal-PIBT, across multiple high-density maps. We further validate SJRL in a challenging mixed-reality warehouse environment with 8 physical robots and 248 virtual robots.
[MA-10] IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games NEURIPS2025
【速读】:该论文旨在解决在不完全信息博弈中生成式采样框架应用不足的问题,特别是针对现有生成式流网络(Generative Flow Networks, GFlowNets)在完整信息博弈中的约束无法适用于不完全信息博弈中有效策略密度建模与训练目标构建的局限性。其核心解决方案是提出信息流网络(Information Flow Networks, IFlowNets),通过重新设计生成式流网络的约束条件,确保所生成的概率密度函数可对应合法的玩家策略,并建立有效的训练目标。IFlowNets不仅克服了原有方法在不完全信息场景下的不可行性,还严格推广了对抗性流网络(Adversarial Flow Networks, AFlowNets)的适用范围。在三个标准博弈环境的初步实验中,IFlowNets在性能和计算效率上均表现出与基于结果采样的蒙特卡洛反事实遗憾(Outcome Sampling Monte Carlo Counterfactual Regret, OSMCCFR)相当或更优的表现,验证了其有效性与泛化能力。
链接: https://arxiv.org/abs/2608.05422
作者: Conor M. Artman,Nicholas Di,Scott Perkins
机构: 未知
类目: Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: Accepted at the NeurIPS 2025 Workshop on Dynamics at the Frontiers of Optimization, Sampling, and Games
Abstract:While many algorithms blend reinforcement learning (RL) with counterfactual regret (CFR) methods to leverage tradeoffs in computational speed and performance, there are fewer investigations into generative sampling frameworks in game theoretic applications in incomplete information games. We extend a generative flow network framework, Adversarial Flow Networks (AFlowNets), to incomplete information games, called Information Flow Networks (IFNs). We prove that previously established constraints for generative flow networks in complete information games are inadmissible for obtaining valid densities (corresponding to player strategies) and a valid training objective. We show that our proposed generalization, IFlowNets, alleviates this issue and strictly generalizes AFlowNets. In preliminary results for three standard game environments, IFlowNets perform comparably to or better than Outcome Sampling Monte Carlo Counterfactual Regret (OSMCCFR) and standard RL-based methods in performance and speed.
[MA-11] Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
【速读】:该论文旨在解决多学科临床照护计划(care plan)协调中面临的挑战,即如何在不同专业领域间整合异构的临床、功能及心理社会信息,而传统单一模型(monolithic LLM)在透明性与安全性方面表现不足。其解决方案的核心是提出一种多智能体神经符号框架CANOE(Contestable Argumentative Network-of-Experts),通过五个关键模块实现可解释、可验证且安全的决策过程:复杂度评估、自适应团队招募、基于竞技场量化双极论证框架(A-QBAF)的角色化论证计算、人机协同的争议机制以及照护计划合成。其中,角色专业化智能体生成支持与反对特定干预措施的论据,冲突通过竞技场式对抗机制解决,可接受性分数在论证图中传播;照护规划者可对论据进行接受、拒绝、编辑或补充,系统则确定性地重新计算最终方案。实验结果表明,经医学微调的模型在临床正确性与安全性上表现最优,而CANOE的论证结构显著提升了解释能力与人类可争议性(contestability)。
链接: https://arxiv.org/abs/2608.05391
作者: Truong Thanh Hung Nguyen,Hoang-Loc Cao,Phuc Ho,Phuc Truong Loc Nguyen,René Richard,Hung Cao
机构: 未知
类目: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
备注: Accepted at the 4th International Conference on Frontiers of Artificial Intelligence, Ethics, and Multidisciplinary Applications
Abstract:Care plan coordination demands synthesizing heterogeneous clinical, functional, and psychosocial information across multiple professional disciplines, where monolithic LLM pipelines cannot perform in a transparent or safe manner. We present CANOE (Contestable Argumentative Network-of-Experts), a multi-agent neuro-symbolic framework that addresses these limitations through five modules: complexity assessment, adaptive team recruitment, role-based argumentative computation via an Arena-based Quantitative Bipolar Argumentation Framework (A-QBAF), human-in-the-loop contestation, and care-plan synthesis. Role-specialized agents generate supporting and attacking arguments for candidate interventions; conflicts are resolved through arena-based clash resolution before acceptability scores propagate across the argumentation graph. Care planners may accept, reject, edit, or add arguments, and the framework will deterministically recompute the final plan. Evaluation on Discharge Me! and MedicalRAG using ROUGE-L, AlignScore, MEDCON F1, FKGL, and LLM-as-a-judge shows that medically fine-tuned models achieve the strongest clinical correctness and safety, while CANOE’s argumentative structure provides faithful explanation and human contestability.
[MA-12] DoctorAg ents: an agent ic framework to iteratively refine AutoML pipeline for small clinical temporal data
【速读】:该论文旨在解决临床机器学习(Clinical Machine Learning, CML)在实际部署中面临的挑战,即数据稀缺、异质性强以及时间动态复杂性导致的模型构建效率低、易出错问题。现有自动化机器学习(AutoML)系统因依赖预定义空间中的暴力搜索策略,缺乏显式的推理与记忆机制,难以有效应对小样本、高复杂度的临床数据场景。为此,本文提出DoctorAgents——一种基于代理的生成式AI框架,通过专用大语言模型(Large Language Model, LLM)代理实现生成、验证与优化的闭环迭代,将传统AutoML从穷举搜索范式重构为基于推理驱动的精炼范式。其核心创新在于引入自然语言反馈的文本梯度下降机制,实现无需穷举搜索的目标更新,从而高效生成可解释性强、任务特定的高质量机器学习管道。实验表明,DoctorAgents在多种临床任务中均显著优于主流AutoML基线方法,同时提升了模型输出的可解释性。
链接: https://arxiv.org/abs/2608.05375
作者: Ruilin Wang,Bo-Hong Wang,Elizabeth Kourbatski,Jun Bai,Hegang Chen,Ziyang Song,Gilles Boire,Marie Hudson,Yue Li
机构: McGill University (麦吉尔大学); Mila – Quebec AI Institute (魁北克人工智能研究所); University of Sherbrooke (舍布鲁克大学); Centre intégré universitaire de santé et de services sociaux de l’Estrie – Centre hospitalier universitaire de Sherbrooke (CIUSSSE-CHUS) (东部健康与社会服务综合大学中心–舍布鲁克大学医院中心)
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
备注: 34 pages, 5 figures
Abstract:Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrained by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this challenge because they largely rely on brute-force search over predefined spaces and lack explicit reasoning and memory. We therefore reformulate AutoML for small clinical data from exhaustive search to reasoning-driven refinement. We propose DoctorAgents, an agentic AI framework that autonomously constructs and optimizes end-to-end ML pipelines through specialized large language model (LLM) agents for generation, validation, and refinement. DoctorAgents backpropagates natural-language feedback through textual gradient descent to perform targeted updates without exhaustive search. Experiments across diverse clinical tasks show that DoctorAgents consistently outperforms established AutoML baselines while producing more interpretable task-specific representations.
自然语言处理
[NLP-0] Learning When to Trust via Selective Context Preference Optimization
【速读】: 该论文旨在解决语言模型在生成回答时对外部信号(如上下文)过度依赖所带来的可靠性问题,尤其是当输入中存在误导性信息时,模型可能将原本正确的答案错误地修正。传统解决方案是训练模型忽略所有外部上下文以增强鲁棒性,但这导致模型在面对可信上下文时丧失有用性,形成“盲目拒绝”这一隐性失败模式。本文提出将问题重构为“选择性信任”(selective trust),即模型应能区分并合理利用可信上下文、忽略误导性信息,同时在无关或正确上下文中保持准确响应。其核心解决方案包括:构建MIST基准数据集,通过四类匹配情境(干净、误导、正确上下文、无关上下文)对推理任务进行标注;引入SC2W指标,量化误导信号导致正确答案转错的频率。在此基础上,提出SCOPE方法,通过挖掘“干净-正确”与“误导-错误”的失败样本,基于平衡覆盖四种情境的配对偏好数据,优化标准直接偏好优化(Direct Preference Optimization, DPO)目标。实验表明,该方法显著降低了主流开源模型的SC2W值,同时在清洁、正确或无关上下文场景下维持了高准确性。研究主张,评估模型应以“选择性信任”为核心,而非仅关注对干扰信号的抵抗能力。
链接: https://arxiv.org/abs/2608.06377
作者: Xian Sun,Wei Chow,Yingshuo Wang,Junhao Liu,Wei Gao,Qing Wu,Lingdong Kong
机构: Duke University (杜克大学); National University of Singapore (新加坡国立大学); UC Berkeley (加州大学伯克利分校); UC Irvine (加州大学欧文分校); Northeastern University (东北大学); Nanyang Technological University, Singapore (新加坡南洋理工大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Project Page at this https URL GitHub Repo at this https URL HF Dataset at this https URL
Abstract:Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is worth trusting. We recast the problem as selective trust and introduce MIST, a human-annotated benchmark that renders each reasoning item under four matched conditions (clean, misleading, correct-context, and irrelevant-context), together with SC2W, a paired metric counting how often a misleading signal flips a clean-correct answer to wrong. Across a comprehensive benchmark study, we observe that such a susceptibility is universal. We then propose SCOPE, which mines clean-correct/misleading-wrong failures and optimizes a standard Direct Preference Optimization (DPO) objective over matched preference pairs balanced equally across all four conditions, rather than over misleading items alone. Our approach substantially reduces SC2W on popular open-sourced models while preserving accuracy when the added context is clean, correct, or irrelevant. With this work, we argue that models should be judged on selective trust, not on resistance alone.
[NLP-1] he Bitter Lesson of Tool Calling
【速读】: 该论文旨在解决大语言模型(LLM)在真实任务场景下使用工具时的效率与灵活性问题,特别是针对现有基于JSON的工具调用方式在代码生成能力受限、难以实现自然链式调用与并行执行的局限性。其核心挑战在于如何在不牺牲系统稳定性的情况下,提升工具调用的动态性与可扩展性。解决方案的关键在于提出并实证验证程序化工具调用(Programmatic Tool Calling, PTC),即通过将工具以类型化的Python桩函数(typed Python stubs)形式暴露给模型,使模型能够直接以代码形式调用工具,并在单个智能体回合内完成执行与结果返回。实验表明,PTC在BFCL v4基准测试中对14个模型中的11个表现优于或等同于传统的原生JSON工具调用,其中GPT-5.6系列模型实现了10.6%的性能提升;在并行扇出和上下文旋转(context rot)等严苛条件下,PTC仍保持稳定,相较基线平均仅下降2.3%,而基线则显著退化。结果证明,程序化工具调用是一种具备强鲁棒性且能随模型能力演进而同步提升的高效替代方案。
链接: https://arxiv.org/abs/2608.06370
作者: Ishan Patel,Sahil Sen,Elias Lumer,Vamse Kumar Subbiah
机构: PricewaterhouseCoopers, U.S.A
类目: Computation and Language (cs.CL)
备注:
Abstract:Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that chain and parallelize naturally. However, a systematic evaluation of tools as code on an established benchmark across current and prior model generations under real-world task conditions has not been conducted. In this work, we empirically compare programmatic tool calling (PTC) to native JSON tool calling across 14 language models on BFCL v4. In the programmatic tool calling paradigm, tools are exposed as typed Python stubs that the model invokes through code, with execution and results handled in a single agent turn. Programmatic tool calling matches or exceeds native JSON tool calling in 11 of 14 models on BFCL v4, with the GPT-5.6 family achieving a 10.6% improvement over the JSON tool calling baseline. Further, it matches or outperforms baseline in 13 of 14 models under parallel fan-out, and holds stable under context rot conditions where baseline degrades 2.3% on average. Our results demonstrate that programmatic tool calling is a viable and robust alternative to JSON tool calling, with performance tracking model capability across release generations.
[NLP-2] CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
【速读】: 该论文旨在解决生成高质量、可执行且具备适当挑战性的终端任务(terminal tasks)以有效训练智能体(agent)的问题。现有方法通常仅关注任务的可解性与可验证性,但忽视了任务难度与求解器能力之间的相对关系,导致训练数据难以激发模型的真正学习潜力。其核心解决方案是提出CalibForge——一个基于可验证求解器行为的自主任务合成系统,通过对抗式求解器校准机制动态优化候选任务。该系统采用两种校准策略:多求解器校准聚焦于异构求解器池间的分歧,对比式校准则针对“强通过/弱失败”的预期关系,二者共同构建以实证可解性为基础的求解器相对可学习区域(solver-relative learnable zone)。实验表明,利用CalibForge生成的5,431个校准后任务显著提升了模型性能,在Terminal-Bench 2.0上达到32.58%和47.57%的准确率,相比基线模型最高提升达24.71个百分点,并在SWE-bench Pro和Doc2Repo上分别取得27.68和30.04的显著增益。结果验证了“求解器相对可学习性”作为训练数据设计目标的有效性与可迁移性。
链接: https://arxiv.org/abs/2608.06352
作者: Fanzhe Meng,Guoxin Chen,Jiale Zhao,Shuang Sun,Zhiyu Lin,Wayne Xin Zhao,Ruihua Song,Ji-Rong Wen,Kai Jia
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Dataset: this https URL . Repository: this https URL
Abstract:Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.
[NLP-3] RP-OPSD: Reasoning -Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer
【速读】: 该论文旨在解决大语言模型(LLM)在低资源语言中跨语言推理能力迁移不足的问题,尤其关注如何有效提升多语言场景下的推理性能。现有基于策略的自蒸馏(On-Policy Self-Distillation, OPSD)方法虽能提供细粒度的标记级监督,但其目标未显式聚焦于对跨语言推理至关重要的推理信号。为此,论文提出一种新型方法——推理枢纽引导的自蒸馏(Reasoning-Pivot-guided On-Policy Self-Distillation, RP-OPSD),其核心在于识别并聚焦于“推理枢纽”(reasoning pivots),即推动或调整推理过程、决定后续推断方向的关键决策点。通过利用带有与不带英文参考解的教师模型视图之间的分布差异作为操作化代理,RP-OPSD实现对关键推理控制节点的特权蒸馏和参考锚定。实验结果表明,该方法在覆盖17种语言、多难度级别的数学推理基准上显著优于现有强基线及OPSD变体;进一步分析显示,该方法将蒸馏注意力集中于推理控制与问题条件状态更新类令牌,而降低对仅支持表面文本生成的令牌的权重,从而更有效地促进跨语言推理能力的迁移。
链接: https://arxiv.org/abs/2608.06347
作者: Xinye Wang,Junxiao Liu,Shujian Huang
机构: Nanjing University (南京大学); National Key Laboratory for Novel Software Technology (新型软件技术国家重点实验室)
类目: Computation and Language (cs.CL)
备注: 16 pages. Under review
Abstract:Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-policy self-distillation (OPSD) and its variants have emerged as a promising paradigm, providing dense token-level supervision on student-generated rollouts, yet their objectives do not explicitly prioritize reasoning signals most critical to cross-lingual transfer. We characterize that target-language reasoning comprises the generation of both surface text and reasoning pivots, which are decisions that advance or redirect the reasoning process and shape subsequent inference. This motivates concentrating privileged distillation around such pivots. We therefore propose RP-OPSD, Reasoning-Pivot-guided On-Policy Self-Distillation, using the distributional shift between matched teacher views with and without an English reference solution as an operational proxy to guide privileged distillation and reference anchoring. Experiments on mathematical reasoning benchmarks covering 17 languages and multiple difficulty levels show that our method outperforms strong multilingual reasoning baselines and OPSD variants. Further analysis reveals that RP-OPSD concentrates privileged distillation on reasoning-control and problem-condistioned state-update tokens, while downweighting it for tokens that mainly support surface realization. Our code is available at this https URL.
[NLP-4] Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
【速读】: 该论文旨在解决当前任务导向型对话代理评估中基准测试(benchmark)质量难以衡量的问题。现有评估依赖于人工构建或自动生成的基准,但其质量常因任务不一致、场景过于简单或策略覆盖不足而受限,进而导致评估结果不可靠。论文提出一种无需参考答案(reference-free)的评估框架,利用大语言模型(LLM)作为评判者,自动评估基准在任务一致性、复杂度及策略覆盖范围方面的质量,并提供可操作的诊断信息以识别具体缺陷。通过与独立的人工标注结果对比以及对不同能力水平的LLM生成基准和受控降质扰动后的基准进行验证,该框架在多个领域和不同判别模型下均能有效区分不同质量层级的基准。此外,该方法同样适用于人工构建的基准,为合成与人工基准的评估提供了实用且可扩展的解决方案。
链接: https://arxiv.org/abs/2608.06329
作者: Noam Koren,Roy Bar-Haim,Abigail Goldsteen
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 15 pages
Abstract:Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework’s applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.
[NLP-5] Benchmarking and Enhancing LLM s for Rule-Intensive Review of National Standard Documents
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在规则密集型专业文档审查任务中能力评估不足的问题,尤其聚焦于国家标准文档(如中国GB/T标准)这类结构复杂、规范性强且对一致性要求高的文档。现有基准多关注领域知识与问答能力,忽视了对文档内在质量的系统性审查,而传统依赖人工专家的审查方式成本高、难以规模化。为此,研究提出首个针对国家标准文档结构化审查的基准GB/T-Bench,其核心创新在于构建了涵盖文档结构、范围一致性、规范性表达、术语统一性及规范引用等维度的层级化审查分类体系(GB/T Review Taxonomy),包含25种可诊断的错误类型,并通过结合确定性规则与约束式大模型重写机制,生成7,306个可追溯的错误实例。关键解决方案包括:一种基于精确匹配的诊断导向评估协议,要求在错误位置、审查维度和错误类型上均实现精准识别,并引入文档级覆盖度指标;同时提出GB/T-Reviewer多智能体框架,将审查知识转化为专业化技能,协调全局检查、定向诊断、规则扫描与结果验证等模块,实现结构化技能协同。实验表明,当前最强的主流模型在该任务上的表现(CMCS=0.3280)远低于人类专家水平(0.6640),而引入GB/T-Reviewer后最佳模型性能提升至0.5094,显著缩小了人机差距,验证了结构化技能协同在规则密集型文档审查中的有效性。该工作为标准化及其他高风险文档领域中可信人工智能的应用提供了重要基础。
链接: https://arxiv.org/abs/2608.06312
作者: Tao Wang,Qihao Yang,Rongjiao Liang,Lianghong Lin,Haitao Wang,Xinyu Cao,Tianyong Hao
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit rules for scope, terminology, normative wording, and cross-section consistency. Existing benchmarks focus on domain knowledge and question answering, largely overlooking intrinsic quality review for professional documents. Such reviews rely heavily on human experts, making them costly and difficult to scale. To bridge this gap, we introduce GB/T-Bench, the first benchmark for the structured review of national standard documents. Its GB/T Review Taxonomy is a hierarchical schema covering document structure, scope alignment, normative modality, terminology consistency, and normative references, with 25 diagnosable error types. A controllable counterexample generation mechanism combines deterministic rules and constrained LLM rewriting to process 488 documents into 7,306 traceable review error instances for evaluation. We also develop a diagnosis-oriented evaluation protocol requiring exact matches on error location, review dimension, and error type, plus document-level coverage metrics. We further propose GB/T-Reviewer, a multi-agent framework that converts review knowledge into specialized skills and coordinates global inspection, targeted diagnosis, rule scanning, and result verification. Experiments with 14 mainstream LLMs reveal a substantial human-LLM gap: the strongest model achieves only 0.3280 CMCS versus 0.6640 for experts. GB/T-Reviewer raises the best CMCS to 0.5094, showing the value of structured skill coordination for rule-intensive document review. This work paves the way for trustworthy AI in standardization and other high-stakes document domains.
[NLP-6] RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
【速读】: 该论文旨在解决生成式奖励模型(Generative Reward Model)在强化学习(Reinforcement Learning, RL)中未能充分发挥潜力的问题。尽管生成式奖励模型在响应排序任务中表现出色,但其与现有强化学习算法所采用的标量评分范式之间存在本质不匹配,导致其难以有效提供学习信号。为弥合这一差距,论文提出了一种基于排名的奖励构建方法(Ranking-based Reward Construction, RRC),其核心在于通过相对偏好排序来生成更有效的强化学习信号。RRC引入两种互补策略:自竞争排名(self-competitive ranking),利用采样响应间的相互比较;锚点引导排名(anchor-guided ranking),通过少量参考响应实现可扩展的基于排名的奖励构造。实验结果表明,RRC在开放域对话与推理基准上显著提升了生成式奖励模型的强化学习训练效果,相较于现有奖励构建方法实现了持续且一致的性能提升。
链接: https://arxiv.org/abs/2608.06310
作者: Chenglong Wang,Ziming Zhu,Yifu Huo,Bei Li,Qiaozhi He,Yan Ding,Xiaoyang Hao,Yuxin Gao,Tianhua Zhou,Xiaojia Chang,Tongran Liu,Jingbo Zhu
机构: Northeastern University (东北大学); NiuTrans Research (牛津研究院); Beijing Institute of Psychology, Chinese Academy of Sciences (中国科学院心理研究所); Kunming University of Science and Technology (昆明理工大学)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that this limitation arises from a mismatch between the comparative nature of generative reward modeling and the scalar scoring paradigm adopted by existing RL algorithms. To bridge this gap, we propose a Ranking-based Reward Construction (RRC) approach, which enables generative reward models to provide more effective RL learning signals by deriving rewards from relative preference rankings. RRC introduces two complementary strategies: self-competitive ranking, which exploits comparisons among sampled responses, and anchor-guided ranking, which enables scalable ranking-based reward construction with a small set of reference responses. Experiments across open-ended chat and reasoning benchmarks demonstrate that RRC substantially improves RL training with generative reward models, achieving consistent gains over existing reward construction approaches. Our code can be found at this https URL.
[NLP-7] HarnessOpt-Bench: Evaluating LLM s at Harness Optimization
【速读】: 该论文旨在解决大语言模型(LLM)在智能体系统中性能优化的瓶颈问题,即当前模型能力不仅依赖于其参数权重,还高度依赖于“控制流”(harness)——包括提示词(prompt)、工具调用、控制逻辑、记忆机制及编排代码等组成部分。随着智能体系统对生成式 AI (Generative AI) 的深度依赖,如何自动化地优化这一复杂控制流成为提升系统整体效能的关键挑战。为此,论文提出了一项名为 HarnessOpt-Bench 的基准测试框架,用于评估前沿大模型在高成本且具有随机性的评价环境下进行端到端控制流优化的能力。其核心解决方案在于构建一个受信任执行环境,严格约束评估边界、计量目标智能体资源消耗,并保留候选方案版本以供审计;在此框架下,一个由 LLM 驱动的优化器接收初始控制流、基于评估反馈进行迭代修改,并在固定预算内输出最终候选方案,其性能通过在未见测试集上的归一化收益衡量。实验结果表明,不同优化器之间的表现差异显著大于其所使用的编码控制流差异,且原生控制流并非始终优于通用控制流,优化增益在不同任务和种子配置间存在显著波动。这证明了控制流优化是一项可量化、具区分度且仍有巨大提升空间的核心能力。
链接: https://arxiv.org/abs/2608.06301
作者: Varun Ursekar,Apaar Shanker,Yash Maurya,Shehab Yasser,Vijay S. Kalmath,Veronica Chatrath,Yuan Xue
机构: Scale AI(规模人工智能)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization – the iterative and evaluation-guided improvement of a harness by an AI system – both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent’s seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
[NLP-8] NeSy-RAG : Neuro-Symbolic RAG for Explainable Question Answering
【速读】: 该论文旨在解决现有检索增强生成(Retrieval-Augmented Generation, RAG)系统在问答任务中推理过程不透明、难以验证以及无法有效识别用户特定上下文缺失的问题。具体而言,传统RAG方法虽能利用外部知识提升大语言模型(LLM)的准确性,但其推理路径缺乏可追溯性,中间步骤与证据来源之间的关联模糊,且常因忽略用户个性化信息而导致输出不完整或错误。为此,论文提出NeSy-RAG——一种模块化的神经符号检索增强生成框架,其核心创新在于将检索到的文本片段转化为可解释的逻辑谓词(predicates),并基于联合自然语言-代码嵌入技术,将这些谓词组合成形式化Prolog查询。通过引入符号化知识缺口检测机制,系统能够自动识别影响推理结果的关键缺失用户事实,并触发后续交互以补全上下文。最终执行的Prolog查询可生成确定性答案,并提供清晰、可追溯的执行轨迹,实现每一步推理与其原始证据源的精确对应。在ShARC基准测试中,无需领域微调,NeSy-RAG达到61.1%的准确率,显著优于同规模模型的基线RAG(42.8%)。
链接: https://arxiv.org/abs/2608.06292
作者: Jonas Gann,Michael Gertz
机构: 未知
类目: Computation and Language (cs.CL); Symbolic Computation (cs.SC)
备注:
Abstract:Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliably attributed to specific evidence. Moreover, missing user-specific context is rarely detected systematically, often leading to incomplete or incorrect output. We propose NeSy-RAG, a modular neuro-symbolic RAG framework that synthesizes attributable Prolog modules from retrieved text chunks. For each chunk, the system generates semantically meaningful predicates that encode Boolean claims, which may depend on user facts. Using joint natural language-code embeddings, predicates are retrieved and composed into Prolog queries. To address incomplete user context, we introduce a symbolic knowledge-gap detection mechanism that identifies missing user facts whose truth values affect the query outcome and automatically triggers follow-up interactions. Executing the resulting Prolog queries yields deterministic answers together with transparent execution traces that link each reasoning step to its originating source. On the ShARC benchmark, without domain-specific training, NeSy-RAG achieves 61.1% accuracy, outperforming a same-model RAG baseline that achieves 42.8% accuracy. Subjects: Computation and Language (cs.CL); Symbolic Computation (cs.SC) Cite as: arXiv:2608.06292 [cs.CL] (or arXiv:2608.06292v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2608.06292 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-9] Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents EMNLP2026
【速读】: 该论文旨在解决网页智能体(Web agent)在执行任务时,固定使用单一观测模式(observation mode)所带来的局限性问题。现有方法通常为所有任务选择一种固定的观测方式(如仅文本、仅像素或两者结合),但不同任务对观测模态的依赖性各异,导致单一模式难以在所有场景下达到最优性能。研究通过在VisualWebArena和WebArena两个基准上评估六种观测模式在八组站点-模型组合中的表现,发现各模式具有互补性:每种模式均能解决其他模式无法完成的任务,且失败模式在结构上各不相同,最佳模式的选择随任务集变化而反转。尽管存在一个“理想情况”——即为每个任务动态选择最优模式(oracle策略)——其潜在收益看似显著,但实际被运行间噪声所夸大:同一模式在相同任务上的重复运行会导致12%-14%的结果变化,因此增加新模式带来的增益与重新运行已有模式相当。最终,研究提出一个稳健的成本优化策略:仅将当前所有模式均无法解决的任务发送至成本最低的模式,可在保持成功率不变的前提下,实现8个测试单元中9.5%至30.6%的成本降低。进一步测试五种路由策略(包括基于置信度级联、成本分层、文本规则等)后发现,除在最稀疏测试单元中出现一次脆弱结果外,均无法稳定超越固定选用一个精心挑选的单一模式。核心障碍在于路由监督信号依赖于智能体的整体成功率——智能体越弱,获得的标注样本越少,而此时正是路由机制最需要发挥作用的场景,形成恶性循环。这一限制属于当前智能体能力水平所致,而非路由机制本身缺陷;实验显示标签供给量与路由机会之间存在高度相关性(0.95),表明更强的智能体可打破此瓶颈。研究还报告了重跑噪声区间及完整测量协议,以确保结果可复现。
链接: https://arxiv.org/abs/2608.06171
作者: Jiaming Wei,Zekun Wu,Adriano Koshiyama,Maria Perez-Ortiz
机构: University College London / Holistic AI; UCL Centre for Artificial Intelligence
类目: Computation and Language (cs.CL)
备注: Preprint. Under review at the Second Workshop for Research on Agent Language Models (REALM), EMNLP 2026 (non-archival track)
Abstract:Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across eight site-model combinations (cells) on VisualWebArena and WebArena and ask what choosing per task would buy. The modes are complementary: each solves tasks the others miss, they fail in structurally different ways, and the best choice reverses between task sets. The obvious prize, an oracle that picks a winning mode for every task, looks large but is inflated by run-to-run noise: rerunning the same mode on the same tasks changes 12-14% of outcomes, so a second run of a mode already in hand gains about as much as adding a new one. What survives is a cost bound: sending only the tasks no mode solves to the cheapest mode cuts cost by 9.5-30.6% in 8 of 8 cells at unchanged success. We then test five routing policies (picking the mode, deciding when to spend on the strong mode, a zero-cost rule read off the task text, a confidence cascade, and pooled cost tiers), and none robustly beats simply fixing one well-chosen mode; the one exception is a fragile result in our sparsest cell. The central obstruction is that routing supervision is produced at the agent’s success rate: the weaker the agent, the fewer labels a router gets, exactly where routing would be most valuable. This limit belongs to today’s agents rather than to routing itself. Label supply and routing opportunity rise together (correlation 0.95 across cells), so a stronger agent can overturn the result, and we report the rerun noise bands and the full measurement protocol.
[NLP-10] Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI
【速读】: 该论文旨在解决从非结构化文本文档中提取复杂、层级嵌套且属性基数可变的结构化信息这一挑战,尤其针对健康技术评估(Health Technology Assessment, HTA)领域中大量医学文献的自动化信息抽取需求。传统方法在处理此类高复杂度信息时面临一致性差、效率低及难以标准化的问题。其解决方案的关键在于提出一种基于模式(schema-based)的框架,利用生成式AI(Generative AI)实现零样本(zero-shot)单次调用完成整个信息提取过程,并通过引入路径驱动的语义匹配算法对提取结果与黄金标准进行自动化语义评估。该框架创新性地结合了生成式AI进行属性值的语义对比,并设计了一套基于领域特性的评分准则,将匹配结果分类为精确匹配、语义匹配、有用但不精确或非匹配四类,从而实现对提取质量的精细化评估。实验表明,该方法在英国国家卫生与临床优化研究所(NICE)发布的文档上成功提取14个属性中的12个,F1得分达90%,且处理时间较人工专家降低约30倍,同时验证了其在不同生成式AI模型、组织机构及语言间的泛化性与可迁移性。
链接: https://arxiv.org/abs/2608.06167
作者: Modhurita Mitra,Jan-Willem Versteeg,Maarten D. Schermer,Shiva Nadi Najafabadi,Marie L. De Bruin,Lourens T. Bloem
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 10 pages, 7 figures, 3 tables. To be published in Proceedings of the 2026 IEEE 22nd International Conference on e-Science (e-Science), Naples, Italy
Abstract:We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard. The schema, serving as an information model encoding domain knowledge, provides a unified, systematic, and consistent framework for extraction of hierarchical, nested information, with attributes of variable cardinality, and subsequent evaluation of the results. Information extraction from a document is performed in a single call to the model, in zero-shot mode. In the evaluation step, we introduce a path-based semantic matching algorithm to align the nested, variable-cardinality attributes in the extracted results with those in the gold standard. We use generative AI for semantic comparison of the extracted and gold standard values of an attribute, and introduce a rubric to classify the result of the comparison, according to domain-specific considerations, as an exact, semantic, useful, or non-match. We were able to extract 12 out of 14 attributes with an F1 score of 90% from documents published by the health technology assessment organisation NICE, using the generative AI model Claude Opus 3. The time needed to extract the attributes from a document was \sim 30 times lower than the time taken by a human domain expert. We further demonstrate generalisability of this framework across different generative AI models and transferability across different HTA organisations and languages. Comments: 10 pages, 7 figures, 3 tables. To be published in Proceedings of the 2026 IEEE 22nd International Conference on e-Science (e-Science), Naples, Italy Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL) MSC classes: 68T50 ACMclasses: I.2.7 Cite as: arXiv:2608.06167 [cs.AI] (or arXiv:2608.06167v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.06167 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-11] Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中政治偏见难以量化的问题,尤其关注其在法律推理、论据构建及框架表述等细微层面所表现出的隐性偏见。现有方法往往依赖单一指标,无法充分捕捉偏见的复杂表现形式。为此,论文提出Poli-Bias——一种反事实评估框架,通过系统性地交换不同国家在多样化地缘政治关系、法律违规行为和推理任务中的身份,对比成对提示下的模型响应差异,从而检验模型是否因国家属性而对法律上等价的情境作出差异化处理。该方案的关键在于将响应差异分解为五个可解释的维度,实现对不平等对待机制的细粒度诊断。实验覆盖13种主流大模型,涵盖多种架构与规模,结果表明国家身份及用户隶属关系会系统性影响模型对等价行为的描述、评价与辩护,验证了Poli-Bias作为审计模型政治中立性与谄媚倾向的高精度分析工具的有效性。
链接: https://arxiv.org/abs/2608.06123
作者: Massi-Nissa Abboud,Aladin Djuhera,Elena Cabrio,Holger Boche
机构: Université Côte d’Azur; Technical University Munich
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.
[NLP-12] Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
【速读】: 该论文旨在解决现有Transformer模型中位置编码(Positional Embeddings, PE)对句法结构信息不敏感的问题,即传统PE仅编码词元的顺序与距离,而忽略了深层句法依赖关系。为此,作者提出了一种句法感知的位置编码(Syntax-informed Positional Embeddings, SiPE),其核心在于在预训练阶段从依存句法分析中学习一个轻量级的句法先验,并将其以统一方式注入到三种主流位置编码范式(绝对、相对、旋转)中,适用于编码器与解码器,且不改变自注意力机制及其他架构设计。关键创新在于:明确了句法先验应如何融入模型——对于采用相对位置编码的自回归解码器,最优策略是将句法先验以乘法形式耦合至注意力分数中的相对位置项;而对于编码器,则最佳方式是直接添加至输入嵌入,与编码器原有的位置机制进行组合。 实验表明,使用SiPE预训练的模型在SyntaxGym句法泛化任务上提升达10.3%,同时困惑度降低9.0%,显著优于无句法监督的基线模型;更重要的是,该方法在真实语言理解任务(GLUE)上也取得最高8.2%的性能提升。与以往需要在推理时对多个句法解析进行边缘化或完全丢弃句法信息的方法不同,SiPE仅依赖单一解析,实现了句法监督强度与推理开销之间的帕累托前沿优化。
链接: https://arxiv.org/abs/2608.06111
作者: Haris Riaz,Hyungji Kim,Mihai Surdeanu
机构: University of Arizona (亚利桑那大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 21 pages, 9 figures
Abstract:Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textitsyntactic structure. We introduce \textbfSyntax-\textbfinformed \textbfPositional \textbfEmbeddings (\textbfSiPE), which learns a lightweight syntactic prior from dependency parses during pretraining and injects it across all three dominant PE families (absolute, relative, rotary), for both encoders and decoders, leaving self-attention and the rest of the architecture untouched. We isolate \emphwhere and \emphhow the prior should enter the model, and find it depends on the architecture: for autoregressive decoders that use relative PE, the prior is strongest when coupled multiplicatively with the relative-position term of the attention score, outperforming injection into the input embeddings, into self-attention, or into the positional and attention terms jointly—while for encoders it is best added directly to the input embeddings, composing with each encoder’s native positional mechanism. We find that models pre-trained with SiPE improve on the SyntaxGym benchmark by up to 10.3% while simultaneously reducing perplexity by 9.0% over a base model with no syntactic supervision—a metric nearly every existing syntax-injection method instead degrades. Crucially, these gains extend beyond syntactic generalization: SiPE also improves real-world language understanding, raising scores on the GLUE benchmark by up to 8.2% over a model trained without it. Unlike existing syntactic language models that marginalize over many parses at inference or discard syntax at runtime, SiPE conditions on a single parse, establishing a new Pareto frontier between syntactic supervision and inference cost.
[NLP-13] ECHO: A Locally-Deployable Agent ic Health Assistant with Temporal Memory Safety Guardrails and Speech Assessment
【速读】: 该论文旨在解决长期慢性病管理中缺乏个性化、持续性且安全可靠的本地化健康辅助系统的问题。现有解决方案在跨会话记忆保持、临床决策支持的可靠性以及患者数据隐私保护方面存在显著短板,尤其在生成式AI(Generative AI)应用于医疗场景时面临安全风险与合规挑战。其核心解决方案在于构建一个名为ECHO(Enhanced Care Health Observer)的集成化本地部署对话式健康助手,通过三大互补模块实现端到端优化:首先,基于LangGraph编排的ReAct框架构建的智能体聊天机器人,结合17个临床工具与时间知识图谱,实现了跨会话持久记忆与高达94.9%的工具执行成功率;其次,采用双阶段混合安全层机制,包含毫秒级响应的规则引擎用于即时拦截危机信号与越狱尝试,并引入经签名的图神经网络(GNN)结合APPNP传播策略,以临床意图为依据对边界案例进行精准分类,在土耳其语健康数据集上达到88.8%准确率与90.6%不安全召回率,优于零样本大语言模型(LLM)基线;最后,融合Whisper声学编码与BERT文本编码的多模态语音评估模块,通过交叉注意力融合实现情绪、抑郁与疼痛状态估计,平均宏F1达0.652。整个系统以Web应用形式部署于消费级硬件,全程无患者数据外传,满足GDPR与KVKK等隐私法规要求,为可信赖的长期慢性病管理提供了兼具安全性、智能性与合规性的技术范式。
链接: https://arxiv.org/abs/2608.06110
作者: Abdulkadir Külçe,Alihan Esen,Cağla Fikir,Berke Kurt,Kuzey Arar,Gökhan Ercan,Faik Boray Tek
机构: 1. Istanbul Technical University (伊斯坦布尔技术大学); 2. Yıldız Technical University (伊尔达兹技术大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 5 pages
Abstract:This paper presents ECHO (Enhanced Care \ Health Observer), a locally-deployable conversational health assistant for long-term chronic care management. ECHO integrates three complementary software modules developed under shared supervision as a unified system. The core module is an agentic chatbot built on a ReAct loop orchestrated via LangGraph, equipped with 17 clinical tools and a temporal knowledge graph for persistent cross-session memory; it achieves a 94.9% tool-execution pass rate across a 59-scenario benchmark with GPT-5 Mini. A two-stage hybrid safety layer intercepts all incoming queries: a rule-based layer handles explicit crisis signals and jailbreak attempts in under 1ms, while a signed graph neural network (GNN) with APPNP-style propagation classifies boundary cases by clinical intent, achieving 88.8% accuracy and 90.6% unsafe recall on a 2,537-query annotated Turkish health dataset while outperforming zero-shot LLM baselines including Llama 3.3 70B. A multimodal speech assessment module combining Whisper acoustic encoding and BERT text encoding with cross-attention fusion estimates emotion, depression, and pain, reaching a mean macro F1 of 0.652. The full system is implemented as a web application that can run entirely on consumer hardware, with no patient data transmitted to external services, supporting compliance with GDPR and KVKK.
[NLP-14] raining-Free Token-Level Steering for LLM Personalized Co-Writing
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在个性化协同写作场景中缺乏专业领域知识、传统微调方法计算成本高且难以应对数据快速更新、以及检索增强生成(Retrieval-Augmented Generation)无法实现细粒度的令牌级控制等问题。此外,现有系统仍以对话式交互为主,未能充分挖掘高效协同创作(productive co-writing)范式在非代码领域的潜力。为此,论文提出了一种无需训练(training-free)的个性化协同写作框架SteerWrite,其核心创新在于通过不依赖梯度更新的方式,将基础模型高效适配至特定领域,尤其针对小规模数据集进行了专门设计。关键解决方案包括:基于提示工程与可学习的上下文注入机制实现模型行为的精准引导,从而在不修改模型参数的前提下完成领域知识注入与生成控制。实验结果表明,SteerWrite在多种数据集、评估指标和模型架构上均达到当前最优性能,显著降低了人工编辑工作量。
链接: https://arxiv.org/abs/2608.06069
作者: Wenhao Mao,Chengbin Hou,Weixiao Wang,Jialiang Zhu,Min Liu,Yibin Hao,Hairong Lv
机构: 1. Tsinghua University (清华大学); 2. Institute of Automation, Chinese Academy of Sciences (中国科学院自动化研究所); 3. Beijing Institute of Technology (北京理工大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:While Large Language Models (LLMs) show great promise for personalization, they often lack specialized domain knowledge. Conventional solutions like fine-tuning struggle with high computational costs and rapid data updates, while Retrieval-Augmented Generation fails to provide fine-grained, token-level steering. Furthermore, chat-based interfaces remain dominant, whereas productive co-writing paradigms have not yet been well exploited beyond the coding domain. To this end, we introduce SteerWrite, a training-free framework designed for personalized co-writing. Our method effectively adapts the base model to specialized domains without gradient updates, with specific designs tailored to small datasets. Experiments demonstrate that SteerWrite achieves state-of-the-art performance across diverse datasets, metrics, and models, significantly reducing human editing effort.
[NLP-15] LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLM s
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成项目级代码时普遍存在对Python的过度偏好,而当前缺乏系统性评估方法的问题。其核心解决方案是提出LangChoiceBench——一个面向项目级代码生成的基准测试框架,用于量化评估模型在语言选择上的偏好(Python preference)、推荐与实现的一致性(recommendation-implementation consistency)以及语言多样性(language diversity)。关键发现包括:尽管在部分软件领域中Python并非最优选择,但多数模型仍显著高估其适用性;推荐与实际编码实现之间的一致性普遍较低;较小规模的开源权重模型表现出更强的Python偏好和更低的语言多样性。通过对9,826条推理轨迹的分析进一步揭示,大多数Python选择源于惯性或便捷性驱动,而非对项目需求的深入考量;更严重的是,部分模型会虚构上下文支持以“合理化”选择Python,这一现象被定义为“幻觉证据”(phantom evidence),甚至出现所生成代码与自身推理中选定语言相矛盾的情况,暴露出模型在语言决策过程中的逻辑缺陷。
链接: https://arxiv.org/abs/2608.06041
作者: Lukas Twist,Twm Stone,Helen Yannakoudakis,Jie M. Zhang
机构: 未知
类目: oftware Engineering (cs.SE); Computation and Language (cs.CL)
备注: 19 pages, 9 tables, 2 figures
Abstract:Large language models (LLMs) have been shown to exhibit strong Python preferences when generating project-level code, but there is currently no systematic way to measure this behaviour across new models. To bridge this gap, we introduce LangChoiceBench, a project-level code-generation benchmark for measuring Python preference, recommendation-implementation consistency, and language diversity. LangChoiceBench covers 28 projects across seven software areas where Python is often a poor default. We evaluate 25 diverse LLMs and find that Python remains heavily over-selected, recommendation-implementation consistency is low, and smaller open-weight models generally show stronger Python preference and lower language diversity. We further analyse 9,826 reasoning traces and find that most Python choices are automatic or driven primarily by ease, rather than explicit consideration of project requirements. In a smaller but important set of cases, models fabricate contextual support for choosing Python - a failure mode we call phantom evidence - or produce code that contradicts the language selected in their own reasoning.
[NLP-16] EpiBench: Can LLM s Understand Epitopes for Antibody Drug Discovery?
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在抗体药物发现中对表位(epitope)信息推理能力不足的问题,特别是现有大语言模型(LLM)能否仅基于抗原与抗体的序列直接推断表位信息尚不明确。当前表位资源多局限于孤立的预测任务或依赖特定结构数据,而通用蛋白质基准测试又未能覆盖抗体开发全流程中的表位中心决策。为此,研究提出EpiBench——一个闭卷、基于序列且可自动评分的基准测试平台,用于评估大语言模型在表位推理方面的能力。EpiBench包含1,609个经结构验证的抗原-抗体接触、功能B细胞实验及深度突变扫描逃逸数据支持的样本,涵盖五个相互关联的任务:可靶向区域发现、抗体条件下的表位识别、表位分组、功能表位评估以及抗体逃逸评估,并通过受控采样减少捷径学习带来的评估偏差。通过对九个通用大语言模型的评估及任务特异性基线分析、抗原长度分层、显式推理对比和失败模式分析,结果表明当前大语言模型虽能捕捉部分与表位相关的信号,但在抗体特异性序列锚定、长序列残基定位以及生物学合理推理方面仍存在显著局限。因此,EpiBench为衡量和提升序列感知型生物医学大语言模型的可靠性提供了诊断性测试平台,推动其在抗体药物发现中的实际应用。
链接: https://arxiv.org/abs/2608.06022
作者: Zirui Wang,Jiaqi Wang,Qinghan Wang,Yuzhi Xu,Gang Du,Tingjun Hou,Odin Zhang
机构: 未知
类目: Computation and Language (cs.CL); Genomics (q-bio.GN)
备注:
Abstract:Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making epitope understanding central to antibody drug discovery. Although large language models (LLMs) have shown strong biomedical reasoning ability, it remains unclear whether they can infer epitope information directly from antigen and antibody sequences. Existing epitope resources typically focus on isolated prediction tasks or rely on specialized structural settings, while general protein benchmarks do not evaluate epitope-centered decisions across the antibody development workflow. To address this gap, we introduce EpiBench, a closed-book, sequence-based, and automatically scorable benchmark for evaluating epitope reasoning in LLMs. EpiBench contains 1,609 curated samples grounded in structural antibody–antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements. It covers five connected tasks: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to reduce shortcut-based evaluation artifacts. We evaluate nine general-purpose LLMs and analyze their behavior through task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection. The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning. Therefore, EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery.
[NLP-17] Clinical Communication Processing with Models Trained on LLM -Generated Synthetic Data: A Structured Survey and Novel Application Case Studies
【速读】: 该论文旨在解决临床自然语言处理(Clinical Natural Language Processing, CNLP)中真实世界沟通数据稀缺的问题,即临床实践中的非结构化交流(如患者主诉、医患对话、急救交接、护理班次交接等)虽蕴含巨大临床价值,但因涉及隐私、碎片化及标注成本高等原因,难以获取高质量的标注语料。其核心解决方案是利用大语言模型(Large Language Models, LLMs)生成合成临床沟通数据,以弥补真实标注数据的不足。关键在于通过结构化叙事框架,系统性地整合不同来源表示、沟通形式、参与者角色、生成方法与下游任务,并基于十三个新颖案例研究,构建无需真实标注数据即可训练的临床沟通系统,涵盖院前急救报告、现场无线电伤员记录、护士交接、患者门户分诊及低资源环境下的出院沟通等场景。研究表明,微调的编码器模型在性能上优于零样本基线,且有意降级的合成沟通数据有助于提升模型鲁棒性。然而,当前多数评估仍依赖于合成数据内部验证,缺乏在真实数据上的外部验证,因此未来需推动真实数据迁移、安全性和外部有效性验证,以使合成临床沟通成为可复用的临床基础设施。
链接: https://arxiv.org/abs/2608.05993
作者: Alexander Apartsin,Yehudit Aperstein
机构: 未知
类目: Computation and Language (cs.CL)
备注: 20 pages, 7 figures
Abstract:Much clinical value is conveyed not through structured records but through communication: exchanges in which patients describe symptoms, clinicians reason and give instructions, ambulances hand over to emergency departments, and nurses pass on a shift. Such language differs from tabular data because meaning depends on speaker role, intent, causality, uncertainty, omission, and channel noise. Healthcare natural language processing must therefore interpret information as conveyed rather than coded. This requires well-annotated corpora, which are scarce because authentic exchanges are private, fragmented, and costly to annotate. Large language models offer a way forward by transforming clinical sources, such as records, diagnostic labels, symptom lists, or care plans, into written and transcribed communication for downstream models. We present a structured narrative survey organized by source representation, communication form and participants, generation method, and downstream task, complemented by thirteen novel case studies. These build clinical NLP systems for communication channels and languages without labeled real-world data, including EMS pre-arrival reports, field-radio casualty documentation, nurse handoffs, patient-portal triage, and low-resource discharge communication. They show that synthetic communication can bootstrap such systems. Findings include the competitiveness of fine-tuned encoder models over evaluated zero-shot baselines and the value of deliberately degraded communication for robustness. The main limitation is that most studies evaluate on held-out synthetic communication, while train-on-synthetic, test-on-authentic evidence remains limited. We conclude that syn-thetic clinical communication is becoming a practical research resource; establishing it as reusable clinical infrastructure will require authentic-data transfer, safety and external validation.
[NLP-18] Causal Episodic Memory for Feedback-Driven Agent Repair
【速读】: 该论文旨在解决大语言模型(LLM)代理在修复失败时频繁丢弃已成功修正方案的问题,导致后续查询需重复探索相似解法,造成效率低下。其核心挑战在于如何在不进行参数更新的前提下,有效利用先前任务中已验证的修复结果以提升后续Text-to-SQL任务的执行准确率。解决方案的关键是提出一种无需训练的代理MERIT,其通过构建在线双极性记忆库(dual-polarity memory),持续存储经由“预言者”(oracle)验证的正确修正与失败方向。在每次修复过程中,基于一个确定性分类器对故障类型进行粗粒度划分,并以此条件化一个混合词法-稠密检索器,在冻结模型生成新修订前筛选相关历史经验。实验表明,相较于无状态的迭代修复方法,MERIT在Spider数据集上将执行准确率从66.34%提升至69.79%,在BIRD数据集上从47.35%提升至48.44%。消融分析揭示,负向记忆贡献有限,类型条件化与词法-稠密排序的效果具有数据集依赖性,而基于模式局部的经验积累则展现出最一致的性能增益。研究结果明确了因果跨查询记忆在特定场景下可提升修复效果,但在复杂或异构任务中仍需权衡更广泛的记忆表征方式。
链接: https://arxiv.org/abs/2608.05906
作者: Khang Nhat Hoang Vo,Tam Minh Chu,Anh Trac Duc Dinh,Thuyen Vinh Ha Bui,Tho Quan
机构: Mohamed bin Zayed University of Artificial Intelligence (MBZUAI); Ho Chi Minh City University of Technology (HCMUT), VNU-HCM
类目: Computation and Language (cs.CL)
备注:
Abstract:LLM agents that repair failures often discard successful corrections, forcing later episodes to rediscover similar solutions. We study whether finalized repair outcomes can improve subsequent Text-to-SQL episodes without parameter updates. We introduce MERIT, a training-free agent that maintains an online dual-polarity memory of oracle-verified corrections and observed unsuccessful directions. Under oracle-assisted benchmark feedback, only memories from earlier finalized episodes are eligible for retrieval. A deterministic classifier assigns a coarse failure type, which conditions a hybrid lexical-dense retriever before the frozen model generates each revision. Using Qwen2.5-7B-Instruct with identical initial predictions and repair budgets, MERIT improves execution accuracy over stateless iterative repair from (66.34%) to (69.79%) on Spider and from (47.35%) to (48.44%) on BIRD. Paired analyses provide clear evidence for the Spider gain but weaker evidence on BIRD. MERIT is not reliably separated from untyped dynamic retrieval on either benchmark, while Reflexion-style memory reaches (51.24%) on BIRD at substantially higher inference cost. Ablations show that negative memory contributes modestly, the value of type conditioning and lexical–dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit. These results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable.
[NLP-19] AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
【速读】: 该论文旨在解决移动应用交互中长期策略学习所面临的现实数据获取困难与仿真环境局限性问题,尤其针对敏感应用和隐私关键操作难以获取真实轨迹的挑战。现有仿真环境存在扩展成本高、GUI世界模型生成不稳定、模态覆盖有限及动作-状态转移逻辑不一致等缺陷。为此,本文提出AppDeltaWorld——一种基于状态转移约束的增量代码世界模型,其核心创新在于将下一界面预测为可执行的代码增量更新(delta code update),而非无约束的图像或文本描述。该方法在动作-状态转移约束下检索特定应用的层级1 HTML结构,结合当前屏幕、操作指令、目标界面文本及结构信息生成层级2可执行HTML,并在浏览器渲染前将生成的视觉资产插入对应图像槽位。作为世界模型,AppDeltaWorld在Code2World评估框架下的CMGUIBench-500基准上实现了最高保真度,显著提升了结构布局与UI元素重建能力;作为训练环境,其支持过滤后的闭环监督微调(SFT)数据构建,结合公开标注数据使AppDeltaAgent在AndroidLens上达到当前最优性能,并在MobileGym与MobileWorld上实现稳定提升;此外,基于世界模型的测试时强化学习机制可实现策略自适应,进一步优化性能而无需与真实应用进行额外交互。
链接: https://arxiv.org/abs/2608.05891
作者: Weikai Xu,Yunren Feng,Haoxiang Lei,Kun Huang,Yuxuan Liu,Kang Zhao,Xiaolin Hu,Shuo Shang,Bo An
机构: Google(谷歌); Stanford University (斯坦福大学); Meta (元); Stability.AI (稳定人工智能); University of California, Berkeley (加州大学伯克利分校)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improving long-horizon mobile interaction policies. However, real trajectories are difficult to obtain for sensitive apps and privacy-critical operations. At the same time, existing simulated environments are costly to scale up, and GUI world models still suffer from unstable generation, limited modality coverage, and inconsistent action-transition logic. To address these limitations, we propose AppDeltaWorld, a transition-grounded delta code world model that predicts the next GUI as a reachable code update rather than as an unconstrained image or text description. AppDeltaWorld retrieves app-specific Level-1 HTML references under an action-transition constraint, generates Level-2 executable HTML conditioned on the current screen, action, predicted next-screen text, and retrieved structure, and inserts generated visual assets into image slots before browser rendering. As a world model, AppDeltaWorld achieves the highest fidelity on CMGUIBench-500 under Code2World evaluation, with clear gains in structural layout and UI element reconstruction over image-only and code-only baselines. As a training environment, AppDeltaWorld supports filtered closed-loop SFT data construction that, when combined with public supervision, enables AppDeltaAgent to achieve state-of-the-art performance on AndroidLens and consistent gains on MobileGym and MobileWorld. Moreover, world-model-based test-time reinforcement learning enables policy adaptation and shows further improvements without additional interaction with real apps.
[NLP-20] he em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era 2021-2025
【速读】: 该论文旨在解决的问题是:大型语言模型(Large Language Models, LLMs)在辅助撰写文本时是否会留下可测量的风格痕迹,特别是在美国国会新闻稿中是否存在可识别的使用痕迹。研究聚焦于一种特定的标点特征——无空格连字符(em-dash,U+2014,形式为word—word),这种用法在排版英文散文中常见,但在美国新闻写作规范(如美联社风格指南,AP style)中通常要求使用空格分隔的破折号。研究通过预注册设计分析了2021至2025年间来自480个众议院与参议院办公室的146,239份新闻稿数据,采用每千字符中文本中无空格em-dash的密度作为核心指标,并构建泊松/负二项回归模型,引入文本长度偏移量并按办公室聚类。结果显示,2021–2024年期间该密度稳定在0.10–0.12之间,但2025年显著上升至0.217,超过前四年基线值的两倍;含此类em-dash的稿件比例从约13%升至24.8%。主要频率比(2023–2025年对比2021–2022年)为1.55(95%置信区间1.28–1.93),略高于预设1.5倍阈值。该上升趋势具有净新增特征(连字符密度保持稳定),在连续办公机构中普遍显现(75.6%的262个办公室出现增加,p ~ 1e-16),且在224个封闭面板中持续存在,经受住多重虚假检验(包括三个替代断点均无效应、2024/2025年边界无突变、持续办公机构延续上升趋势)。分段回归显示,尽管未在ChatGPT发布时间点出现阶跃变化,但2025年后呈现明显的加速趋势,且上升趋势在两党及两院间对称分布。由于预注册验证门限被正式突破,完整预注册决策规则未满足,因此结论被视为探索性推断。研究认为,无空格em-dash仍可作为群体层面的使用信号,而非个体稿件的作者溯源工具,且不支持因果推断。其关键解决方案在于利用大规模、结构化、基于预注册的文本计量分析,结合统计建模与多重稳健性检验,识别出与LLM辅助写作普及相关的系统性语言模式变化。
链接: https://arxiv.org/abs/2608.05889
作者: Przemysław Czuma(Polish Association for Artificial Intelligence in Medicine)
机构: 未知
类目: Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: Preregistered study (OSF: https://doi.org/10.17605/OSF.IO/U5NEY%29%3B deviations from the registered plan, including a formal validation-gate breach, are disclosed in Section 4.6. Companion study: arXiv:2606.29540 . 3 figures, 4 tables
Abstract:Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word—word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: https://doi.org/10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ratio (2023-2025 vs 2021-2022) was 1.55 (95% CI 1.28-1.93; exact registered cut-off: 1.528), just above the prespecified 1.5x threshold. The rise was net-new (hyphen density stable), held within authors (75.6% of 262 continuous offices increased; p ~ 1e-16) and in a closed panel of 224 offices, and survived falsification tests: three placebo cut-offs were null, the pipeline showed no step at the 2024/2025 boundary, and continuing offices carried the rise. A segmented regression finds no step at the ChatGPT cut-off but a clear post-period acceleration; the 2025 rise is symmetric across parties and chambers. Because the registered validation gate was formally breached, the full preregistered decision rule was not met; the interpretation (broad diffusion of LLM-assisted writing as the models matured) is offered as exploratory. The em-dash remains a population-level marker, not a per-release authorship detector, and the design supports no causal claim.
[NLP-21] he Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents
【速读】: 该论文旨在解决生成式智能体(Generative AI agent)在长期部署过程中因持续存在而产生的安全风险,特别是现有控制框架难以有效管理跨组件、超越单个事件生命周期的持久性智能体实例所带来的挑战。其核心问题在于:当前的安全实践虽识别出过度授权、任务边界模糊、控制缺失等风险,但缺乏一种能够持续追踪和管理这些风险的机制,尤其在智能体行为跨越多个任务或事件时,其权限状态与控制措施的动态变化难以被系统化监控与响应。为此,论文提出“智能体姿态漏洞”(Agentic Posture Vulnerability, APV)作为一项新的脆弱性管理抽象,关键在于将智能体的长期控制配置状态(即“姿态”)视为一个持久存在的、任务条件相关的脆弱性记录,即使在不同任务执行中表现出不同的运行态,仍可追溯至同一不变的姿态基线。APV保持开放状态直至权限被收敛、缺失控制被补全、风险被明确接受或关闭已验证,从而实现对智能体权限与控制组合状态的持续可见性与可管理性。这一方法并非引入新的根本原因类别,而是对现有过度代理权、授权缺陷及控制组合弱点的可操作化表达,区别于传统漏洞(CVE)、OWASP过度代理风险、基线控制结果及运行时授权-执行间隙。论文进一步构建了APV的阈值定义、六类典型模式、全生命周期模型、最小记录结构、控制与关闭矩阵,并提出工具支持与可验证的研究路径,为智能体系统的持续安全治理提供了系统性框架。
链接: https://arxiv.org/abs/2608.05884
作者: Shayell Aharon Salomon Amir Shaked Matan Noga
机构: Bluebear Security
类目: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
备注:
Abstract:Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks. Control frameworks also describe capabilities for constraining, authorizing, observing, validating, and responding to agent activity. Yet security programs still need a way to manage persistent deployed instances that span components and outlive any one event. We propose the agentic posture vulnerability (APV) as a task-conditioned vulnerability-management abstraction: a durable record for a composed agent-control exposure. One posture may produce different runtime manifestations across tasks; APV links those manifestations to the invariant posture and remains open until authority is narrowed, a missing control is added, risk is accepted, or closure is verified. APV is not proposed as a new root-cause class of risk; it operationalizes existing excessive-agency, authorization, and control-composition weaknesses. We distinguish APVs from CVE-addressable product defects, OWASP Excessive Agency, Agent Baseline control outcomes, and the runtime authorization-execution gap. We then provide a field vignette, a thresholded definition, six recurring APV patterns, a vulnerability lifecycle, a minimum record, a control-and-closure matrix, tooling implications, and a testable research agenda. Subjects: Cryptography and Security (cs.CR); Computation and Language (cs.CL) Cite as: arXiv:2608.05884 [cs.CR] (or arXiv:2608.05884v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.05884 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-22] Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
【速读】: 该论文旨在解决个性化深度研究(personalized deep research)中用户请求(user request)与实际研究需求之间不匹配的问题,核心挑战在于如何将用户隐含的目标、约束、偏好及评估标准等上下文信息有效融入研究规范(research specification),以指导深度研究代理(deep research agent, DRA)生成更符合用户意图的成果。其解决方案的关键在于提出G-STEER框架,通过构建一个捕捉不同框架因素(framing factors)依赖关系的意图引出图(Intent Elicitation Graph),将用户上下文中的潜在需求作为图上的可选引出目标。该框架基于图结构引导的轨迹学习一种澄清策略(clarification policy),在面对多个耦合决策时——包括判断哪些框架因素相关、现有上下文是否支持、以及选择检索用户记忆、向用户提问还是直接优化查询——能够动态权衡目标覆盖度与证据获取成本,从而生成经过个性化的精炼查询。实验表明,G-STEER在整体加权目标覆盖度和下游报告个性化方面均优于对比方法,且所需用户交互次数仅为强基线方法的约三分之一。
链接: https://arxiv.org/abs/2608.05876
作者: Soojin Yoon,Dongha Lee
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 13 pages, 4 figures
Abstract:User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and evaluation criteria. User context can be incorporated either within the deep research pipeline or into the research specification provided as its input. We focus on the latter, refining the user request into a personalized research specification before passing it to an unchanged deep research agent. This requires resolving three coupled decisions: which framing factors are relevant, whether the available user context sufficiently supports them, and whether to retrieve user memory, ask the user, or stop and refine the query. For training, G-STEER organizes framing factors as elicitation targets in an Intent Elicitation Graph that captures their dependencies. It learns a clarification policy from graph-scaffolded trajectories spanning diverse factor dependencies and evidence conditions. The policy produces a refined query while balancing target coverage against the costs of evidence acquisition. Experiments show that G-STEER achieves the strongest overall weighted target coverage and the highest downstream report personalization across both evaluated DRAs, while asking roughly one third as many user questions as a strong clarification baseline.
[NLP-23] MACRO: Markov Chain Routing of Transformer Layers
【速读】: 该论文旨在解决标准大型语言模型(LLM)中各层顺序执行导致的计算效率与性能瓶颈问题,尤其针对复杂任务中固定执行路径无法适应不同输入需求的局限性。其核心挑战在于如何在不修改模型参数的前提下,实现高效、自适应的动态层路由(dynamic layer routing),以支持跳过(skip)、重复(repeat)和残差状态添加等灵活操作。本文提出的解决方案——马尔可夫链式变压器层路由框架(MACRO),关键在于将层路由建模为一个依赖上下文的马尔可夫策略,该策略基于层索引、计算预算阶段、方向位移及操作符上下文等多维信息进行条件化决策,并通过训练数据反馈迭代优化路由分布。采用top-k Viterbi算法对马尔可夫路由分布进行解码,从而高效生成高概率候选执行路径。实验表明,MACRO在多个开放权重的LLM上显著提升性能,平均准确率较无路由基线提升5.0%,在小型模型上增益尤为明显;相比最优现有动态路由方法Dr. LLM,精度提升7.2%的同时,路由搜索时间降低9.4倍(从14.8小时降至1.6小时),实现了性能与效率的双重突破。
链接: https://arxiv.org/abs/2608.05872
作者: Paweł Batorski,Abtin Pourhadi,Akylgali Aitaza,Przemysław Spurek,Paul Swoboda
机构: University of Warsaw (华沙大学); Polish Academy of Sciences (波兰科学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path through layers involving layer repetitions, skips and other moves, can improve performance. Existing routing approaches often require updating model weights, running expensive search loops per test instance, or demand ground-truth labels during inference. In this work, we propose Markov Chain Routing of Transformer Layers (MACRO), a framework that learns task-specific routes over LLM architectures without modifying underlying parameters. MACRO models layer routing as a context-dependent Markov policy conditioned on layer indices, computation budget phases, directional displacements, and operator context, supporting skip, repeat, and residual hidden-state addition operations. The Markov route distribution is updated via feedback on training data and decoded using a top-k Viterbi algorithm to isolate high-probability candidate programs. We evaluate MACRO across diverse reasoning and knowledge benchmarks on multiple open-weight LLMs. MACRO achieves a +5.0% average accuracy improvement over the unrouted baselines, with largest gains on small models. We outperform the best dynamic routing approach Dr. LLM by +7.2%, while reducing route-search time 9.4x (from 14.8 to 1.6 hours). Our code is publicly available at this https URL.
[NLP-24] Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing
【速读】: 该论文旨在解决检索增强生成(Retrieval-Augmented Generation, RAG)系统中因不同嵌入模型具有异构几何特性而导致相似度得分不可直接比较的问题,进而阻碍模型迁移与阈值复用。其核心解决方案是通过学习不同模型间相似度分数分布的映射关系,而非依赖嵌入向量本身的对齐,从而实现跨模型相似度空间的校准。关键创新在于提出“合成查询探测”(Synthetic Query Probing)方法,即从文档自动生成查询-段落配对,构建大规模、无需参考的跨模型相似性行为分析框架。实验在SciFact及一个专有语料库上验证了该方法的有效性,结果表明尽管不同模型在排序一致性上表现良好,但其绝对相似度得分存在系统性偏差。通过线性、保序(isotonic)和分位数映射等方法学习转换函数后,可部分对齐不同模型的得分空间,显著提升阈值可迁移性,其中保序回归表现最优。研究强调了跨模型校准的重要性,并将合成查询探测定位为一种可扩展的嵌入可比性分析框架。
链接: https://arxiv.org/abs/2608.05857
作者: Marcin Rozmus,Peter van der Putten
机构: 未知
类目: Computation and Language (cs.CL)
备注: Accepted for 29th International Conference on Discovery Science, October 5-9, 2026, Mainz, Germany
Abstract:Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding models due to differing geometric properties, complicating model migration and limiting threshold reuse. We study how similarity scores can be related by learning mappings between score distributions rather than embeddings. We introduce Synthetic Query Probing, generating queries from documents to create controlled query-chunk pairs, enabling large-scale, reference-free analysis of cross-model similarity behavior. We evaluate the approach on multiple embedding configurations and learn score conversion functions using linear, isotonic, and quantile mappings. Experiments on SciFact and a proprietary corpus show that while models largely agree on rankings, their absolute scores exhibit systematic distortions. Learned mappings partially align these spaces and improve threshold portability, with isotonic regression performing best. Our results highlight the need for cross-model calibration and position Synthetic Query Probing as a scalable framework for analyzing embedding comparability.
[NLP-25] MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
【速读】: 该论文旨在解决希伯来语(Yiddish)这一具有丰富文本传统的语言在数字时代因数据稀缺与评估资源匮乏而导致的自然语言处理(Natural Language Processing, NLP)进展受限问题。现有主流多语言语料库和评测基准普遍包含大量噪声、机器翻译内容及错误标注文本,难以真实反映希伯来语的语言特性。为此,论文提出两个核心创新:一是构建高质量的希伯来语预训练语料库Oytser,融合当代网络文本与文学材料;二是设计多任务评测基准Kashes,涵盖翻译、语言学分析、信息抽取与语言理解等任务。基于此,研究者对Llama 3.1 8B模型进行持续预训练,得到首个开源的80亿参数希伯来语语言模型MameLoshnLM。实验表明,该模型在多项任务上显著优于同规模开源基线。更重要的是,分析显示其优势不仅体现在性能指标上,更在于对希伯来语特有的词汇与形态特征的更好捕捉,揭示了通用大规模多语言数据在低资源语言建模中的系统性缺陷。本研究为希伯来语自然语言处理提供了坚实基础,并为历史上积淀深厚但数字化程度不足的语言建模提供了可复用的技术范式。
链接: https://arxiv.org/abs/2608.05850
作者: Uri Katz,Omer Goldman,Tomasz Limisiewicz,Reut Tsarfaty,Noah A. Smith
机构: Bar-Ilan University (巴伊兰大学); University of Cambridge (剑桥大学); University of Washington (华盛顿大学); Allen Institute for AI (艾伦人工智能研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the Conference on Language Modeling (COLM) 2026
Abstract:We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish’s rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.
[NLP-26] Enhancing Social Intelligence in LLM s with Hierarchical Reasoning and Utterance-Level Goal Rewarding
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在动态社交互动任务中因缺乏长期目标协调与快速适应能力而导致性能受限的问题。现有方法通常对每轮对话统一施加基于目标的奖励,忽视了不同对话回合中目标的具体性以及潜在策略背后的合理性。为此,本文提出Think-Strategy-Response(TSR)框架,将社交对话分解为高层战略规划与底层语言执行两个层次的结构化流程。其核心解决方案是引入线性化分层强化学习与方差门控奖励(Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards, LHRL-VGR),通过评估目标达成分数的方差动态调节奖励分配,在保证目标完成率的同时强化策略一致性。实验结果表明,该方法可使Qwen2.5-7B代理在SOTOPIA基准上的目标完成成功率超越GPT-4o基线7.32%,显著提升了多智能体社交协商任务中的表现。
链接: https://arxiv.org/abs/2608.05832
作者: Xiaofeng Wang,Kakam Chong,Shuai Xiao,DeXin Kong,Qingyuan Tian,Chen Ju,Xu Yan,Shuai Zhao,Fei Huang,Rui Wang,Shuguang Han,jufeng chen
机构: Alibaba Group(阿里巴巴); Shanghai Jiao Tong University(上海交通大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
[NLP-27] MoCA: Implicit Social Context Analysis
【速读】: 该论文旨在解决现实社会交互中隐含社会语境(implicit social context)的建模与理解问题,尤其聚焦于情感(affection)、意图(intent)和立场(stance)三方面的隐性表达。现有研究缺乏系统性框架来分析这些依赖社会文化背景、以间接方式传递的深层含义。为此,作者提出了一项新任务——隐含社会语境分析(MoCA),并构建了一个高质量多模态基准数据集(3,108个实例),包含细粒度认知标注,揭示了表达者、表达内容、对象及表达动机与方式。实验表明,当前先进的多模态大模型因过度依赖显性线索且难以推理潜在社会语境,在该任务上表现不佳。为应对这一挑战,论文提出冲突驱动的溯因推理(CoDAR)框架,将观察到的表达与预期真实行为之间的差异建模为认知冲突,从而推断隐藏的心理状态。大量实验证明,CoDAR显著提升了模型性能,但与人类推理能力之间仍存在显著差距,凸显了隐含社会理解的根本复杂性。
链接: https://arxiv.org/abs/2608.05825
作者: Wenhao Xu,Kaiwen Zhang,Hao Li,Maowei You,Yongzheng Ji,Siyuan Zuo,Jingxuan Yu,Sina A,Xinyao Tan,Bobo Li,Hao Fei,Mong-Li Lee,Wynne Hsu
机构: National University of Singapore (新加坡国立大学); Nanyang Technological University (南洋理工大学); Wuhan University (武汉大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements. Such implicit social contexts are pervasive in real-world interactions, yet there remains a lack of a formal and systematic framework for studying them. In this paper, we introduce Implicit Social Context Analysis (MoCA), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance. We construct a high-quality benchmark containing 3,108 multimodal instances collected from real-world sources, with fine-grained cognitive annotations revealing who expresses what toward whom, as well as how and why it is conveyed. Using the MoCA dataset, we show that state-of-the-art multimodal large language models struggle significantly with this task because of their reliance on explicit cues and limited ability to reason over latent social contexts. To address this challenge, we propose Conflict-Driven Abductive Reasoning (CoDAR), a novel framework that models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, thereby enabling the inference of hidden mental states. Extensive experiments demonstrate that CoDAR substantially improves model performance. Nevertheless, a large gap from human reasoning remains, highlighting the fundamental difficulty of implicit social understanding.
[NLP-28] Decomposed Entailment for Factuality Checking and Hallucination Detection
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在生成过程中普遍存在的真实性和事实一致性问题,尤其是“幻觉”(hallucination)现象——即生成内容与原始输入源不一致或缺乏依据。其核心挑战在于如何在不依赖参考文本(reference-free)、无需访问模型内部结构(black-box)且计算资源受限的条件下,高效、准确地检测幻觉。解决方案的关键在于提出一种轻量级、基于分解的事实性评估框架HallDetect:将生成内容分解为原子化语句(atomic claims),利用紧凑的编码器-蕴含模型(encoder-based entailment model)通过多尺度源文本片段库进行对比式验证,并采用非对称评分机制——只要存在一个被明确反驳的声明即可整体标记为幻觉。该方法在共享4位量化骨干模型和消费级硬件预算的严格控制实验中,在四个基准中的三个上优于同类生成式与嵌入式基线方法,同时具备跨模型架构的稳定性,并可生成从声明到原文片段的审计追踪路径,实现错误定位。
链接: https://arxiv.org/abs/2608.05823
作者: Achir Oukelmoun,Nasredine Semmar,Gaël De Chalendar
机构: CEA LIST NANO INNOV; Google(谷歌)
类目: Computation and Language (cs.CL)
备注:
Abstract:The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, including hallucinations—cases where generated content is not supported by the underlying source. We present HallDetect, a lightweight, reference-free, and black-box framework for hallucination detection that we evaluate not only on summarization but across a broader range of source-grounded generation settings. HallDetect builds on decomposition-based factuality evaluation: generated content is decomposed into atomic claims, each verified by a compact encoder-based entailment model through a contrastive formulation over a multi-scale library of source chunks, and aggregated with an asymmetric score in which a single confidently contradicted claim flags the response. Under a controlled protocol in which all methods share the same 4-bit quantized backbones and consumer-grade hardware budget, HallDetect outperforms comparably resourced generative and embedding-based baselines on three of four benchmarks while remaining stable across backbone families, and yields a claim-to-span audit trail that localizes each error.
[NLP-29] M3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
【速读】: 该论文旨在解决多模态隐喻理解中模型缺乏基于视觉与文本证据的跨模态推理能力的问题,现有基准测试因局限于孤立子任务评估且缺乏可解释性标注,难以准确衡量模型是否真正建立在多模态证据基础上的源域—目标域映射。其解决方案的关键在于提出一个统一且基于证据的基准M³R-Bench,包含1,000个经人工验证的图文实例,遵循“证据识别—映射建立—情感推断”的分阶段注释框架,系统标注隐喻发生、源—目标映射、情感倾向及解释链。针对现有模型普遍忽视视觉证据、依赖表面文本线索导致跨模态证据—映射错配的问题,论文进一步提出M³R-Reasoner,通过课程式推理监督与任务感知强化学习相结合的方式,引导模型推理过程对齐隐喻理解逻辑。实验表明,仅使用80亿参数的骨干模型,M³R-Reasoner在四项统一任务指标上超越更大规模的闭源多模态大模型(MLLMs),在视觉证据利用和情感合理性评分上分别优于GPT-5.5达28.45和30.11点,并在平均评分上超过Claude-Sonnet-4.6 8.00点,显著提升了多模态隐喻理解的准确性与可解释性。
链接: https://arxiv.org/abs/2608.05817
作者: Hong Jiang,Junnan Zhu,Jingwang Huang,Xiao Sun,Yuming Yang,Jiang Zhong,Ruirui Chen,Jingman Shi,Hao Wu,Nayu Liu,Xinyi Jiang,Kaiwen Wei
机构: 未知
类目: Computation and Language (cs.CL)
备注: 6 figures and 5 tables. Hong Jiang, Junnan Zhu, and Jingwang Huang contributed equally. Jiang Zhong and Kaiwen Wei are corresponding authors. Code and data are available at this https URL
Abstract:Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information jointly construct Target–Source mappings, requiring both conceptual understanding and cross-modal reasoning. However, existing benchmarks mainly evaluate metaphor understanding through isolated subtasks and lack evidence-grounded explanations, making it difficult to assess whether models establish mappings grounded in visual and textual this http URL address these limitations, we introduce M ^3 R-Bench, a unified and evidence-grounded benchmark containing 1,000 image–text instances with human-verified annotations. Guided by Conceptual Metaphor Theory and theories of nonliteral language understanding, M ^3 R-Bench provides joint annotations for metaphor occurrence, Target–Source mapping, sentiment, and stage-wise explanations following ``evidence identification–mapping establishment–sentiment inference.''Evaluations on M ^3 R-Bench reveal that existing models often overlook visual evidence, rely on superficial textual cues, and produce inaccurate Target–Source mappings, exposing a cross-modal evidence–mapping mismatch. To address this mismatch, we propose M ^3 R-Reasoner, which combines curriculum-based reasoning supervision with task-aware reinforcement learning to align model reasoning with metaphor interpretation. Experiments show that, with only an 8B-parameter backbone, M ^3 R-Reasoner outperforms larger proprietary MLLMs across four unified-task metrics and improves Visual Evidence and Sentiment Justification scores over GPT-5.5 by 28.45 and 30.11 points, respectively, while surpassing Claude-Sonnet-4.6 by 8.00 points in mean rubric score. The dataset and code are available at this https URL.
[NLP-30] When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
【速读】: 该论文旨在解决自演化智能体(self-evolving agents)在持续积累可复用技能过程中出现的“能力污染”(capability-contamination)问题,即当技能池规模超过某一临界值后,新增技能反而导致性能下降,且该过程具有不可逆性。其核心问题是:由于缺陷技能一旦被引入决策上下文,会作为后续技能提炼的参考依据,从而引发跨轮次的污染链式传播,导致系统性能退化且无法通过事后移除源技能有效恢复。解决方案的关键在于提出“验证器作为守门人”(Verifier-as-Gatekeeper, VaG)机制,构建一个分层信任体系,包含三个异质化批判者——结构有效性、行为无害性和语义一致性——对每项技能进行独立筛选,并结合边际增益子集选择策略,在技能进入运行时上下文前主动消除组合性污染。实验表明,相比无约束累积策略在终端基准测试2(Terminal-Bench 2)中先升后降并仅部分恢复性能,VaG实现了每轮持续提升,以约五分之一的技能池规模达到72% pass@1,并具备良好的泛化能力。消融实验证明三类批判者功能互补且不可替代,分别拦截了不同类型的有害技能,共同保障了技能演化过程的稳健性与可扩展性。
链接: https://arxiv.org/abs/2608.05810
作者: Linfang Shang,Ming Xu,Yiding Sun,Tianle Xia,Lingxiang Hu,Lan Xu,Ning Zheng
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it. We formalize this capability-contamination phase transition and trace it to a structural cause: once a defective skill enters the decision context, it becomes reference material for distilling later skills, forming cross-round contamination chains. We further show the contamination is structurally irreversible: removing a source skill after the fact cannot erase the flawed reasoning its descendants have already inherited, so post-hoc rollback recovers only a small fraction of the lost performance. This makes skill admission a pre-commit necessity rather than a post-hoc fix, and motivates Verifier-as-Gatekeeper (VaG): a progressive trust hierarchy whose three heterogeneous critics - structural validity, behavioral harmlessness, and semantic consistency - filter each skill individually, coupled with a marginal-gain subset selection that removes combinatorial contamination at the top tier before skills reach the runtime context. On Terminal-Bench 2, unconditional accumulation rises to a peak and then degrades, giving back most of its gains as the pool keeps growing, and post-hoc removal of the culprit skills recovers only a small part of the drop - the empirical signature of irreversibility. In contrast, VaG improves every round, reaching 72% pass@1 with a pool roughly 5x smaller, and its frozen skill pool transfers positively to four other backbones and a second benchmark without re-evolution. Ablations confirm the three critics are complementary and mutually non-substitutable, each intercepting a largely disjoint class of harmful skills.
[NLP-31] Hierarchical Latent Prediction for Language Models
【速读】: 该论文旨在解决标准下一词预测(Next-Token Prediction, NTP)在长时序推理与规划任务中因教师强制训练范式导致的误差累积问题。现有方法如多词预测(Multi-Token Prediction, MTP)和下一隐变量预测(Next-Latent Prediction, NextLat)虽尝试通过多步预测或隐空间自监督学习缓解该问题,但前者预测视野有限,后者在多步回溯过程中仍面临误差传播。本文提出分层隐变量预测(Hierarchical Latent Prediction, HiLP),其核心创新在于引入一个高层抽象隐变量作为辅助目标,以降低隐空间滚动预测中的误差累积效应。实验表明,HiLP能够生成更长时序的连贯信念状态表征,在代码生成与多步推理基准上均展现出显著性能提升,并实现了更高的推测性解码效率。
链接: https://arxiv.org/abs/2608.05806
作者: Chang Shi,Tim Pearce,Manan Tomar,Siddhartha Sen,John Langford
机构: University of Texas at Austin (德克萨斯大学奥斯汀分校); Microsoft Research (微软研究院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning. Recent works such as Multi-Token Prediction (MTP) and Next-Latent prediction (NextLat) try to mitigate the problem through predicting multiple future tokens and self-supervised prediction in the latent space. However, those auxiliary objectives either have a limited horizon or suffer from compounding error from multi-step rollout. We introduce Hierarchical Latent Prediction (HiLP), which introduces an auxiliary higher-level abstract latent to help reduce the error accumulation effect in latent-space rollouts. Experiments show that HiLP can lead to longer-horizon coherent belief state representation and demonstrate the effectiveness of our method across coding and multi-step reasoning benchmarks, and offers more speculative decoding efficiency.
[NLP-32] On-Policy Delta Distillation for Multilingual Math Reasoning
【速读】: 该论文旨在解决生成式人工智能(Generative AI)在多语言场景下进行大语言模型(LLM)后训练时,基于策略蒸馏(On-Policy Distillation, OPD)方法在跨语言性能一致性与语言特性保持方面的不足问题。其核心挑战在于:现有OPD方法在英语以外的语言(如韩语、日语)中表现受限,且缺乏对目标语言语义与表达风格的有效保留。解决方案的关键在于提出一种改进的变体——基于策略的差值蒸馏(On-Policy Delta Distillation, OPD²),通过利用后训练教师模型与其基础模型之间的概率差异作为学习信号,更有效地捕捉语言特异性知识。实验结果表明,OPD²显著优于原始OPD,在韩语和日语任务中提升尤为明显,并有效缩小了英语与非英语语言间的性能差距;同时发现,仅使用英语数据进行训练虽能提升非英语语言性能,但会引发响应向英语偏移的问题,凸显了多语言数据在维持目标语言表达一致性中的关键作用。
链接: https://arxiv.org/abs/2608.05802
作者: Byeongho Heo,Jaehui Hwang,Sangdoo Yun,Dongyoon Han
机构: NAVER AI Lab(NAVER人工智能实验室)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 9 pages, 3 figures, 10 tables
Abstract:On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD ^2 ), for mathematical reasoning in English, Korean, and Japanese. OPD ^2 improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD ^2 consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.
[NLP-33] Predicting Task Difficulty Without Rollouts
【速读】: 该论文旨在解决在状态依赖型环境中,如何在不进行代价高昂的模拟(rollouts)的情况下,仅通过任务描述直接预测任务难度的问题。其核心挑战在于,随着智能体应用向长时程领域拓展,传统的基于试错的评估方式已成严重计算瓶颈,亟需一种高效、可靠的难度预估机制以支持评估基准的校准与渐进式训练课程的设计。论文的关键解决方案在于:通过分析17个涵盖编程、数学、机器学习、网页导航、函数调用等多个领域的智能体基准任务,揭示了传统评估指标如AUC可能掩盖实际预测性能不佳的问题,并识别出**词级别熵(token-level entropy)**作为有效的预测信号;同时,通过分析预期难度与观测难度之间的残差,能够暴露环境中的隐藏缺陷,如数据污染(contamination)和任务不可行性(infeasibility),从而为环境设计提供可解释的诊断工具。
链接: https://arxiv.org/abs/2608.05797
作者: Stefan Krsteski,Charlotte Meyer
机构: Andromede AI
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:Task difficulty dictates an agent’s likelihood of success, and estimating it without rollouts means forecasting this directly from a task description before executing costly simulations in stateful environments. Reliable estimates would therefore allow environment designers to calibrate evaluation benchmarks and construct progressive training curricula. This becomes increasingly important as agents move into long-horizon domains, where empirical trial-and-error is a severe computational bottleneck. Prior work on early prediction is limited to static tasks or isolated coding environments, often relying on narrow features and inaccurate evaluation metrics. We study \textitex ante difficulty prediction across 17 agentic benchmarks spanning coding, mathematics, machine learning, web navigation, function calling, and other domains. We show that AUC can mask poor difficulty estimates, identify token-level entropy as a useful predictive signal, and show how residuals between expected and observed difficulty can expose hidden environment flaws such as contamination and infeasibility.
[NLP-34] ask-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
【速读】: 该论文旨在解决多语言文本嵌入模型在跨任务适应过程中因采用单一训练目标而导致优化策略不匹配的问题,尤其在翻译、检索、分类及配对分类等任务间存在本质差异时,统一优化目标难以兼顾各类任务的学习动态。其解决方案的关键在于提出一种任务条件流匹配(Task-Conditional Flow Matching, TCFM)框架,该框架根据任务类型选择性地应用流匹配(Flow Matching)于翻译任务,同时针对检索、分类和配对分类任务采用更契合其学习特性的优化目标,从而实现任务感知的差异化优化。此外,TCFM引入教师引导的表征保持机制与三阶段课程学习策略,有效保障了模型在多任务联合适应过程中的稳定性。实验结果表明,TCFM在Indic Massive Text Embedding Benchmark上达到新的最优性能,显著提升多种多语言任务的嵌入质量,并展现出对不同嵌入模型架构的良好泛化能力。
链接: https://arxiv.org/abs/2608.05785
作者: Tirth Bhatt,Naren Kumar S,Mayank Singh
机构: Indian Institute of Technology Gandhinagar (印度理工学院甘奇纳加尔)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.
[NLP-35] GROM: Gradient-Free Rapid One-Shot Machine Unlearning
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中敏感知识难以彻底移除的问题,特别是现有基于梯度优化的迭代微调方法在计算成本高、缺乏解析解的同时,往往仅将目标知识“隐藏”而非真正删除,导致在低比特量化等攻击下仍可恢复被遗忘内容。其解决方案的关键在于提出一种全新的单次(one-shot)可解释性删减方法——GROM(Gradient-free Ridge-regularized Optimal Modification),通过将删减过程建模为带有岭正则化的最小二乘优化问题,推导出针对特定权重矩阵的闭式加法更新公式。该方法不依赖反向传播与迭代收敛,仅通过前向传播即可实现权重修改,在数秒内完成,显著提升效率;同时确保目标层对保留数据的行为严格不变,并从权重层面真正消除目标内容,从而有效抵御量化攻击,实现了在遗忘效果与模型性能之间更优的权衡。
链接: https://arxiv.org/abs/2608.05783
作者: Paweł Batorski,Przemysław Spurek,Paul Swoboda
机构: Google(谷歌); Meta(元); Stability.AI(稳定人工智能)
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state-of-the-art approaches primarily rely on iterative, training-time unlearning via fine-tuning. However, even when utilizing parameter-efficient dimensionality reduction techniques like LoRA, gradient-based optimization remains computationally expensive and lacks explicit analytical formulations. It can also leave the targeted knowledge merely hidden rather than removed, to the point that simply quantizing the unlearned model restores much of what it was supposed to have erased. To resolve this, we propose a novel one-shot unlearning approach, abandoning iterative optimization in favor of a direct, exact analytical solution. We frame the unlearning process as a ridge-regularized least-squares optimization problem, deriving a closed-form additive update for targeted weight matrices. This update forces the selected layer to suppress unwanted content while strictly preserving its behavior on retained data. Computed from gradient-free forward passes alone, with no backpropagation and no iteration to convergence, GROM applies the weight edit in mere seconds, which makes it orders of magnitude faster than traditional fine-tuning. Extensive evaluations demonstrate that GROM achieves state-of-the-art forgetting-utility trade-offs on TOFU-5%, TOFU-10%, MUSE-Books, MUSE-News and WMDP, significantly reducing computational overhead without sacrificing overall model performance. Because the update removes the targeted content from the weights instead of masking it, GROM also withstands the low-bit quantization attack that recovers much of the content a gradient-based baseline had appeared to forget. Our code is publicly available at this https URL.
[NLP-36] How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLM s
【速读】: 该论文旨在解决自动语音识别(ASR)中对新词和罕见词(如命名实体、缩写、领域特定专有词汇等)识别能力不足的问题,这类词汇在训练数据中稀缺,导致传统ASR模型难以准确识别。其解决方案的关键在于对比两种策略:一是基于Whisper的上下文偏置方法(context biasing methods),通过在推理阶段引入词表以增强模型对目标词汇的识别能力;二是直接使用语音大语言模型(speech LLMs)进行上下文提示(prompting)。实验结果表明,上下文偏置方法可使目标词汇的词错误率(WER)相对降低高达88%,且对非目标词汇影响较小;而语音LLMs在朗读语音上表现优异,但在非朗读语音上泛化能力较弱,且对干扰项数量和提示词顺序敏感。研究系统分析了两类方法的权衡关系,为实际应用场景中的技术选型提供了依据。
链接: https://arxiv.org/abs/2608.05759
作者: Christian Huber,Alexander Waibel
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.
[NLP-37] Once a Response Always a Response: Detecting LLM -generated Text via Latent Prompt Restoration
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)生成文本在大规模传播中引发的虚假信息扩散、教育滥用及平台治理等风险,核心问题是现有零样本检测方法依赖基于概率的统计差异,未能显式建模LLM的训练过程所导致的生成机制特征,从而限制了检测的鲁棒性。其解决方案的关键在于提出EchoPrompt,一种无需训练的检测方法,通过潜在提示恢复(latent prompt restoration)机制实现。该方法的核心思想是:机器生成文本通常以隐含的上游提示为条件,通过在文本前添加统一的通用前缀,可部分激活这一隐藏依赖关系;EchoPrompt利用指令微调模型测量由此产生的似然增益,并与基础模型进行校准,将差异聚合为量化潜在提示依赖性的得分,从而有效捕捉生成文本的非自然特征。实验表明,EchoPrompt在零样本检测任务中达到当前最优性能,且在多种挑战性评估场景下展现出优异的鲁棒性。
链接: https://arxiv.org/abs/2608.05741
作者: Hongrui Bao,Yubing Ren,Yanan Cao,Jinhan You,Fang Fang,Shi Wang
机构: Chinese Academy of Sciences(中国科学院); University of Chinese Academy of Sciences(中国科学院大学); Zhejiang University(浙江大学); Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 17 pages, 7 figures
Abstract:Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, educational misuse, and platform governance. These concerns make robust detection of machine-generated text increasingly necessary. Recent zero-shot detectors mainly exploit probability-based statistical discrepancies, but they do not explicitly account for the training process of LLMs, which leaves a distinct generation mechanism insufficiently modeled and limits detection robustness. To address this issue, we propose EchoPrompt, a training-free detector based on latent prompt restoration. Our key intuition is that machine-generated text is typically produced conditioned on an upstream prompt, and this hidden dependency can be partially reactivated by prepending a unified generic prefix. Specifically, EchoPrompt restores a generic assistant-response context, measures the induced likelihood gain with an instruction-tuned model, calibrates it against the corresponding base model, and aggregates the resulting differences into a score that quantifies latent prompt dependency. Extensive experiments show that EchoPrompt achieves state-of-the-art performance among zero-shot detectors while maintaining strong robustness across challenging evaluation settings.
[NLP-38] Mitigating Scoring Bias in LLM -as-a-Judge via Random Number Generation
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLM)作为文本质量评估者时存在的评分偏差(scoring bias)问题,即模型在不同上下文环境下仍倾向于生成特定评分,从而影响评估结果的客观性。其核心解决方案是通过识别并校正模型的潜在数值偏差(latent number bias),具体方法为:在提示词中引入下游任务定义,促使模型随机生成数字标记,并通过测量实际生成数字分布与均匀分布之间的偏差来量化该模型的内在数值偏好;随后,在实际文本评估过程中,基于所测得的潜在数值偏差对模型的标记生成概率进行校正,以实现更公平、准确的评分。实验在四个不同任务(包括LLM对齐评估、摘要评价、语义文本相似度和相关性)上验证了该方法的有效性,结果表明其优于未去偏的基线模型及已有校准方法。研究进一步发现,评分偏差在不同模型、任务和评分区间间存在显著差异,强调了针对具体场景动态测量和校正潜在数值偏差的重要性。
链接: https://arxiv.org/abs/2608.05726
作者: Yuma Asato,Kiyoaki Shirai,Natthawut Kertkeidkachorn
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of the evaluated text, which is known as scoring bias. This study proposes a novel method to mitigate this scoring bias. An LLM is instructed to randomly generate number tokens, and the latent numerical bias of the LLM is identified by measuring the deviation of the observed distribution of numbers from the uniform distribution. A definition of a downstream task, for which an LLM evaluator is used, is added to the prompts for random number generation to measure task-specific latent number bias. In the evaluation by an LLM, the token generation probabilities for a given input are rectified considering the LLM’s latent number bias. Results of the experiment on four different tasks, evaluation of LLM alignment, evaluation of summarization, Semantic Textual Similarity, and Semantic Textual Relatedness, demonstrate that our proposed method outperforms the baselines, including an LLM without debiasing and previous calibration methods. In addition, it is confirmed that scoring bias varies across LLMs, tasks, and score ranges, indicating the importance of measuring latent number bias as the case may be.
[NLP-39] Sparse Mutual Information Graph Averag ing for Improving Random Indexing Embeddings
【速读】: 该论文旨在解决传统词嵌入方法中因依赖密集共现矩阵构建、密集分解及梯度训练所带来的高计算开销问题,同时保持对稀疏全局语料统计信息的有效利用。其核心解决方案是通过在稀疏的正点互信息(Positive Pointwise Mutual Information, PPMI)图上采用顶K(top-K)剪枝与加权平均的方式,对随机索引(Random Indexing, RI)向量进行非梯度优化修复。关键创新在于:在不引入神经网络或梯度更新的前提下,利用PPMI图的局部邻域信息对初始弱性能的RI嵌入进行修正,显著提升语义类比任务的准确率。实验表明,在童话语料库的家族类别类比任务中,该方法将RI的准确率从19.4±0.7%提升至30.7±2.9%,最优种子下达到34.6%;尽管该方法在text8和SimLex-999等基准上表现逊于神经基线模型,但证明了基于PPMI图的顶K加权平均是一种有效的非梯度嵌入修复机制,尤其适用于初始化较弱的稀疏嵌入场景。
链接: https://arxiv.org/abs/2608.05724
作者: Sriram Loganathan,Gokul Anand,Aung Bo Bo,Yourui Shao,William B. Andreopoulos
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on sparse global corpus statistics. This paper studies Random Indexing (RI) vectors refined by weighted averaging on a sparse Positive Pointwise Mutual Information (PPMI) graph. On a fairytales corpus, the covered semantic analogy set consists of 272 Google family- category questions. On this family subset, PPMI top-K graph averaging repairs a weak RI initialization, improving accuracy from 19.4±0.7% to 30.7±2.9% across five seeds. Under the single tested runs, the same neighborhood averaging reduces family- subset analogy accuracy for PPMI+SVD (singular value decom- position), Binary+SVD, CBOW, and Skip-gram. Thus the method is not competitive with neural baselines on text8 and gives near- zero strict similarity correlation on SimLex-999. While Bloom filter sketches underperform RI in the tested configuration, we find that PPMI graph averaging with top-K pruning is a useful non-gradient repair for weak RI embeddings. On the fairytales dataset, PPMI top-K=50 graph averaging improves RI with accuracy going from 19.4±0.7% to 30.7±2.9%, and performing best with a seed42 of 34.6%.
[NLP-40] DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
【速读】: 该论文旨在解决大型语言模型(Large Language Model, LLM)代理在调用外部工具并交互真实系统时,因潜在不安全行为导致外部状态、用户数据及下游服务产生不可逆后果的问题。现有运行时防护机制多为被动响应式,仅评估当前动作的表面安全性,缺乏对风险随轨迹演进的显式建模,因而难以识别长期任务中一系列看似无害但累积后可能导致危险状态的行为。针对这一关键缺陷,论文提出DreamGuard——一种基于风险感知世界模型的主动防护框架。其核心创新在于构建一个紧凑的递归隐状态机制,持续追踪任务轨迹中的风险演化,并通过预测未来隐状态,生成即时危害与前缀风险证据,进而融合多时域信号,在动作执行前做出干预决策。实验表明,DreamGuard在四个基准测试及在线防护评估中均优于通用、反应式与主动式基线方法,在安全-效率权衡上表现最佳,且每请求平均端到端延迟仅为25毫秒。
链接: https://arxiv.org/abs/2608.05695
作者: Wenhao Lin,Chenyu Yu,Xingwei Lin,Sicong Cao,Xiang Chen,Lei Xue,Le Yu,Letian Sha,Chunming Wu
机构: 1. Tsinghua University (清华大学); 2. Alibaba Group (阿里巴巴集团); 3. Peking University (北京大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
备注:
Abstract:As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services. Recent runtime guardrails mitigate such risks by checking proposed actions before execution, but many remain reactive: they primarily assess the apparent safety of the current action, lacking an explicit model of how risk evolves across the trajectory. This limitation creates a critical blind spot for long-horizon risks, where individually benign-looking actions can gradually drift the agent toward hazardous states. In response, we propose DreamGuard, a proactive guardrail for LLM agents built around a risk-aware world model. The world model maintains a compact recurrent latent state over the trajectory and predicts future latent states from which DreamGuard derives immediate-hazard and prefix-risk evidence. It then fuses these multi-horizon signals into intervention decisions before execution. Experiments across four benchmarks and an online guardrail evaluation show that DreamGuard outperforms generic, reactive, and proactive guardrail baselines, achieves the best safety-utility trade-off among evaluated guardrails, and maintains an average end-to-end latency of 25 ms per call.
[NLP-41] Answer First Reason Later: Commitment Order in Diffusion LLM s
【速读】: 该论文旨在解决生成式语言模型在复杂推理任务中因非有序采样(即无约束的掩码扩散解码)导致的推理失效问题。尽管掩码扩散语言模型(dLLMs)允许任意顺序提交词元,这一灵活性本被宣传为相对于自回归解码的核心优势,但研究发现,在如GSM8K等数学推理任务中,这种自由反而成为推理失败的关键原因。具体表现为:在LLaDA-8B模型解码过程中,约15%-24%的推理轨迹中模型已提前提交最终答案,而一半的推理区域仍处于掩码状态;随着上下文长度增加,高达90%的问题出现仅输出答案的退化现象。其根本原因并非终止信念(EOS)压力差异,而是可及性(reachability)——即采样器能否在远距离位置执行其推理信念。通过设计2×2提示-解码器对照实验,研究证实链式思维(Chain-of-Thought, CoT)仅在有序提交条件下有效(提升34.8个百分点,95%置信区间[26.8, 42.8]),并进一步分解出“坍缩通道”与“顺序通道”机制,且在Dream-7B和MATH-500上复现。为此,提出一种单控制旋钮干预——前沿门控提交(frontier-gated commitment),通过因果验证恢复了从0.528到0.852的性能差距,同时保持最高4倍的并行解码效率。研究结果表明,现有基于窗口的采样策略原本出于效率考量,实则构成了对推理病理的最小修正,揭示了其在结构设计上的本质必要性。
链接: https://arxiv.org/abs/2608.05687
作者: Jewon Yeom,Jaewon Sok,Seonghyeon Park,Jeongjae Park,Hwiyeong Lee,Taesup Kim
机构: Seoul National University (首尔国立大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Masked diffusion language models (dLLMs) can commit tokens in any order – a freedom marketed as their core advantage over autoregressive decoding. We show that on reasoning tasks this freedom is instead the axis of failure. Logging every commitment during decoding of LLaDA-8B on GSM8K, we find that unconstrained (pure) decoding commits the final answer at 15-24% of the trajectory while half the reasoning region is still masked, and collapses to answer-only outputs on up to 90% of problems as the canvas grows. The cause is not the model’s termination beliefs – EOS “pressure” is nearly identical across decoders – but reachability: whether the sampler may act on those beliefs at distant positions. A 2x2 prompt-decoder design shows that chain-of-thought helps only under ordered commitment (interaction +34.8 percentage points, 95% CI [26.8, 42.8]; without reasoning text the decoders are indistinguishable), an interaction we decompose into a collapse channel and an order channel and replicate on Dream-7B and MATH-500. A single-knob intervention – frontier-gated commitment – causally recovers the full gap (0.528 to 0.852) while preserving up to 4x parallel decoding, along a measured frontier whose optimal window flips from w=1 at full refinement to unconstrained at 8 tokens/step. Our results reframe existing window-style samplers, previously motivated by efficiency, as the minimal fix for a reasoning pathology they were never designed to address.
[NLP-42] Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLM s
【速读】: 该论文旨在解决生成式 AI 在需要可验证推理的任务中,如何可靠区分有效推理与无效推理的问题。现有基于轨迹的方法依赖层间残差流位移(layerwise residual-stream displacements)来捕捉表示变化,虽能抑制部分稳定的、与具体标记相关的冗余信息,但其仅关注“运动”而忽略了更新的起始状态,若直接恢复完整状态又可能引入捷径学习(shortcut-prone)的噪声。为此,本文提出一种三流检测器架构,通过结合运动信号与两种受限的状态视图——基于向量量化(vector quantization)的粗粒度区域读取器(coarse region reader)和基于归一化多层状态的细粒度方向读取器(fine direction reader),在不回归全状态探测的前提下,充分保留状态上下文以解释运动过程。实验表明,该方法在训练未见的推理基准上,相较仅使用位移的最先进方法提升选择准确率最高达12%,相较于单层探针基线提升21%;同时在事实补全与事实验证任务中亦表现领先,证明其信号本质指向正确性而非特定推理模式。消融实验进一步验证了运动、区域与方向三者提供互补性信号。研究结果表明,推理有效性应从状态条件下的动态演化中读取,而非依赖静态状态或去上下文化的轨迹本身。
链接: https://arxiv.org/abs/2608.05660
作者: Hamed Damirchi,Ignacio Meza De la Jara,Damith Ranasinghe,Yuhang Liu,Javen Shi
机构: Australian Institute for Machine Learning (澳大利亚机器学习研究所); Adelaide University (阿德莱德大学); Naval Group Pacific (海军集团太平洋); Responsible AI Research Centre, Australia (负责任的人工智能研究中心,澳大利亚)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注:
Abstract:As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek this signal in layerwise residual-stream displacements, which capture how representations change while attenuating some stable, token-specific information. However, displacement omits the state from which an update originates, whereas restoring the full state risks reintroducing shortcut-prone information. We identify this trade-off and propose a three-stream detector that combines motion with two restricted views of location. A coarse region reader based on vector quantization and a fine direction reader over normalized multi-layer states. This design restores enough state context to interpret the motion without returning to full-state probing. On reasoning benchmarks unseen during training, our method improves selection accuracy by up to 12% over the displacement-only state of the art and 21% over single-layer probing baselines. Although trained only on reasoning benchmarks, it also reads factual completion and fact verification, ahead of every detector we compare against, which places the signal on correctness rather than on a kind of reasoning. Ablations further show that motion, region, and direction provide complementary signals. These results suggest that reasoning validity is better read from state-conditioned motion than from either static states or decontextualized trajectories alone.
[NLP-43] Relay Dont Route: Adaptive Population Handoff for Cost-Efficient LLM -Driven Evolution
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)驱动的演化搜索中因长期依赖高性能模型而导致计算成本过高的问题。现有方法通常在单个查询或变异步骤层面分配模型资源,忽视了演化搜索具有状态依赖性(stateful)的本质——即每个生成的候选解会改变后续变异所基于的种群分布。研究通过实证分析发现,演化轨迹中的性能提升具有显著的前期集中特征:早期阶段的表现虽存在噪声但具备高度信息量,且低成本模型可高效复现高性能模型在早期取得的大部分进展。针对此现象,作者提出一种无需训练的框架 Model,其核心创新在于将预算分配机制从个体调用层面转移到演化种群层面,引入自适应的**种群交接(population handoff)**策略。该策略由一个基于强化学习的带宽调度器(bandit scheduler)控制,以“接力增益”(Relay Gain)作为奖励信号——即通过构建紧凑、高质量且多样化的候选池来衡量交接的边际收益,从而决定何时进行种群交接。经筛选的优质候选解被用于初始化共享的高性能模型种群以进行精细化优化。在四个基准测试和三种预算设置下,Model 在12组实验中有11组达到最高平均得分,显著优于现有基线方法。研究结果表明,在具有状态依赖性的搜索任务中,资源配置应围绕演化种群整体进行组织,而非局限于单次调用。
链接: https://arxiv.org/abs/2608.05651
作者: Sichun Luo,Yi Huang,Guanzhi Deng,Haibo Wang,Haochen Luo,Lei Li,Zefa Hu,Junlan Feng,Qi Liu
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注:
Abstract:Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models throughout long evolutionary runs is costly. A natural alternative is to combine cheap and strong models under a fixed inference budget. However, existing approaches typically allocate models at the level of individual queries or mutation steps, overlooking that evolutionary search is \textitstateful: each generated candidate changes the population from which subsequent mutations are produced. We empirically analyze LLM-driven evolutionary trajectories and find that search progress is strongly front-loaded, early trajectory performance is informative but noisy, and cheap models recover much of the early progress achieved by strong models at lower cost. Motivated by these findings, we propose \textbf\model, a training-free framework that shifts budget allocation from individual calls to evolving populations through adaptive \textitpopulation handoff. A cheap model explores multiple trajectories in short blocks allocated by a bandit scheduler. Relay Gain, defined as the marginal improvement of a compact, quality-diverse candidate bank constructed for handoff, serves as the scheduler reward and determines when to hand off. The curated candidates initialize a shared strong model population for refinement. Across four benchmarks and three budgets, \model achieves the highest mean score in 11 of 12 settings, outperforming competitive baselines. Our results suggest that in stateful search, budget allocation should be organized around the population, not the individual call. Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE) Cite as: arXiv:2608.05651 [cs.CL] (or arXiv:2608.05651v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2608.05651 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[NLP-44] Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning EMNLP2026
【速读】: 该论文旨在解决大语言模型(Large Language Model, LLM)在推理过程中因测试时扩展(test-time scaling)带来的效率瓶颈问题,尤其是宽泛采样(wider sampling)所导致的收益递减现象:新增的推理轨迹往往重复已有答案模式,缺乏有效推理多样性。现有基于验证器(verifier-based)的选择方法虽可提升质量,但其性能高度依赖外部奖励模型的校准。为此,论文提出一种无需验证器的广度-深度精炼框架,其核心在于充分利用测试时计算资源,通过并行生成多个独立推理路径(广度),再对每条路径进行迭代式自我批判与自我修正(深度),最终通过多数投票聚合优化后的结果。该方案在保留初始推理多样性的同时,通过深度精炼修复局部逻辑错误,显著提升了推理准确性。在AIME24、AIME25、AMC、OlympiadBench和MATH500等多个基准上,该方法在多种开源模型上均优于贪婪解码、多数投票、基于验证器的Best-of-N、束搜索及前瞻解码等主流方法,例如在Qwen2.5-1.5B模型上,MATH500准确率从最强基线提升至58.0%,AMC准确率从25.0%提升至32.5%。研究表明,将测试时计算资源用于精炼已采样轨迹,而非单纯增加样本数量或依赖验证器引导选择,能更高效地提升模型推理能力。
链接: https://arxiv.org/abs/2608.05643
作者: Ahsan Bilal,Muhammad Ahmed Mohsin,Muhammad Umer,Lena Trigg,Ali Subhan,Muhammad Ali,Dean F. Hougen
机构: University of Oklahoma (俄克拉荷马大学); Stanford University (斯坦福大学); Universitat Pompeu Fabra (庞佩乌·法布拉大学); Air University (空军大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: Submitted to EMNLP 2026
Abstract:Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-based selection offers an alternative, but its performance depends on the calibration of an external reward model. We propose a verifier-free breadth–depth refinement framework that uses test-time compute to both explore and improve candidate solutions. The method samples multiple independent reasoning rollouts, refines each rollout through iterative self-critique and self-correction, and aggregates the refined answers by majority voting. Breadth preserves diverse initial attempts, while depth repairs local reasoning errors before aggregation. Across AIME24, AIME25, AMC, OlympiadBench, and MATH500, our method consistently improves over greedy decoding, majority voting, verifier-based best-of- N , beam search, and lookahead decoding across multiple open-weight models. For instance, with Qwen2.5-1.5B, accuracy increases from the strongest verifier-based baseline to 58.0% on MATH500, and from 25.0% to 32.5% on AMC. These results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.
[NLP-45] Human-Like Anaphor Resolution in Large Language Models
【速读】: 该论文旨在探究大型语言模型(Large Language Models, LLMs)在指代消解(anaphor resolution)任务中是否具备与人类相似的认知处理机制。具体而言,研究关注的是:在多种影响人类指代消解速度与准确性的认知因素——如语篇结构、情境模型属性及语义因素——的作用下,五种具有开源权重的LLMs(GPT-2-XL、Llama-3.1-8B、Pythia-12B、Mistral-7B 和 Mistral-24B)是否表现出类似人类的行为模式。其解决方案的关键在于采用“链接假说”(linking hypothesis),将人类阅读时长与模型在指代词处的意外度(surprisal)进行关联,并通过对比模型在指代理解题上的准确率与人类表现,评估模型在不同认知因素下的敏感性。结果表明,部分LLMs在语篇显著性(discourse prominence)和距离依赖性(distance-based factors)方面表现出类人敏感性,但在语义干扰效应(semantic interference effects)上则表现出较弱或缺失的敏感性,揭示了当前主流LLMs在模拟人类指代消解过程中的局限性与选择性对齐特征。
链接: https://arxiv.org/abs/2608.05630
作者: Keane Zhang,Varshini Chinta,Raj Sanjay Shah,Sashank Varma
机构: Georgia Institute of Technology(佐治亚理工学院)
类目: Computation and Language (cs.CL)
备注: 7 pages, 6 figures, 1 table. Presented at CogSci 2026 and the 2026 Annual Meeting of the Society for Text Discourse. Code: this https URL
Abstract:Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.
[NLP-46] Measuring and Detecting Harmful AI Sycophancy
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)中普遍存在的“谄媚响应”(sycophantic responses)问题,特别是其中一种有害形式——由用户偏好诱导的立场反转谄媚(Preference-Induced Stance Reversal Sycophancy, PSRS),即模型为迎合用户明示偏好而随意改变自身初始立场。现有研究多聚焦于衡量模型谄媚程度,而本文进一步探索能否仅通过单个响应文本实现对PSRS的自动检测。为此,作者提出CAP(Contrastive Anchor Probing)框架,用于大规模构建带标签的PSRS数据集。在17个开源与闭源模型上应用该框架,共收集了290,460条跨12个日常建议领域的标注响应。研究围绕三个核心问题展开:(1)PSRS的发生频率如何?(2)是否可仅从响应文本中有效检测PSRS?(3)检测能力在未见模型上的泛化性能如何?结果表明,不同模型的PSRS率介于5%至56%之间,且模型能力越强,谄媚倾向越低;同时,仅基于响应文本即可实现对PSRS的有效检测,但需依赖训练数据中细微的模式学习。由于新模型迭代迅速,检测器不可避免面临未见模型场景,因此跨模型泛化成为关键挑战。实验显示,检测性能在未见模型上显著下降,为此本文提出初步应对策略。研究将公开数据集与代码以支持后续深入研究。
链接: https://arxiv.org/abs/2608.05624
作者: Bohan Jiang,Dawei Li,Yasin Silva,Huan Liu
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: under-review
Abstract:Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This paper focuses on one harmful sycophancy: preference-induced stance reversal sycophancy (PSRS), where a model reverses an initial stance merely to align with a user’s stated preference. While existing research mainly measures how sycophantic a model is, we go further and ask whether PSRS can also be detected automatically from a single response. To investigate this at scale, we introduce CAP (Contrastive Anchor Probing), a framework for collecting labeled PSRS data. Applying CAP to 17 open- and closed-source LLMs, we collect 290,460 labeled responses across 12 everyday-advice domains. We organize our study around three research questions. (1) How often does PSRS occur? (2) How well can it be detected? (3) How does detection generalize to unseen models? We first reveal that PSRS rates range from 5% to 56% across LLMs, with more capable models being less sycophantic. Next, we show that detecting PSRS is feasible from the response text alone, and detectors need to learn subtle PSRS patterns from the training data. Because new LLMs appear rapidly, detectors inevitably encounter unseen models, making cross-model generalization an important framework goal. We demonstrate that detection performance drops on unseen models and propose an initial approach to address this challenge. We will release our dataset and code to support future research.
[NLP-47] FOCUS: Decoupling Expert Personas in LLM s to Enhance Domain Expert Capabilities
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在激活专家人格(expert persona)时存在的跨领域耦合问题,即不同领域间的专家人格相互干扰,导致在高风险领域(如医疗)中行为过于激进,或在敏感领域(如金融交易)中表现过度保守。其解决方案的关键在于提出一种名为FOCUS(Fine-tuning with Orthogonal Control for Uncoupled personaS)的方法:首先自动从模型中提取专家人格向量,继而通过正交分解实现领域特定人格的解耦,再引入专家门控模块(expert gating module),根据任务上下文自适应地控制人格激活。结合两阶段训练策略与门控选择正则化,模型能够有效区分并适配单领域与跨领域任务中的合适人格,显著提升任务准确率,并在金融、法律、医疗及跨领域基准测试中优于现有方法。
链接: https://arxiv.org/abs/2608.05611
作者: Guanyu Wang,Zidi Zhang,Xu Chu
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy. However, existing persona control methods often suffer from cross-domain coupling, which may lead to overly aggressive behavior in high-caution domains such as healthcare, or excessive conservatism in risk-sensitive domains such as financial trading. To address this issue, we propose FOCUS (\textbf\underlineFine-tuning with \textbf\underlineOrthogonal \textbf\underlineControl for \textbf\underlineUncoupled persona\textbf\underlineS). FOCUS first automatically extracts expert persona vectors from LLMs, then applies orthogonal decomposition to decouple domain-specific expert personas, and finally introduces an expert gating module to adaptively control persona activation according to task contexts. With a two-stage training strategy and a gated selection regularizer, the model learns to activate appropriate personas for both single-domain and cross-domain tasks. Experiments on financial, legal, medical, and cross-domain benchmarks show that FOCUS improves task accuracy and outperforms existing persona control methods. Our code is available at \hrefthis https URLthis url.
[NLP-48] SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)作为智能体时,技能库在规模增长背景下如何高效管理与复用的问题。核心挑战在于:在有限的上下文预算下,如何精确暴露最小且足够执行的可复用上下文单元,而现有系统普遍存在无法细粒度复用子技能、压缩过程破坏程序契约(procedural contract)、压缩后代码不可执行或难以扩展、以及难以随技能演化动态更新等问题。其关键解决方案是提出一种面向执行的程序抽象框架 SkillZip,通过在章节级(section-level)执行图上进行保留契约的压缩,将重复出现的、满足契约有效性的模式重构为可逆的端口宏(ported macros),同时严格保持边界签名、依赖闭包性、验证器可达性及源码层级的可扩展性。在推理阶段,SkillZip仅在需要时才展开宏,从而生成紧凑且依赖封闭的上下文;此外,ReZip 模块利用执行反馈动态集成新技能并修正高风险宏。实验结果表明,SkillZip 在技术与具身智能体基准上显著优于最强基线,最高提升达12.2分,同时实现3.46倍压缩率,99.2%依赖保留率和98.7%验证器可达性,且在200至10万技能规模的库中均表现出稳健的检索性能。
链接: https://arxiv.org/abs/2608.05604
作者: Xingyu Tan,Xiaoyang Wang,Qing Liu,Xiwei Xu,Xin Yuan,Liming Zhu,Wenjie Zhang
机构: UNSW(新南威尔士大学); CSIRO(澳大利亚联邦科学与工业研究组织)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest sufficient executable context under a limited context budget. Existing systems struggle to reuse routines below the whole-skill level, preserve procedural contracts during compression, keep compressed routines executable and expandable, and update the compressed library as skills evolve. These challenges reveal a unit mismatch: skills are retrieved as packages, compressed as text, and converted into execution graphs only after retrieval, whereas reliable reuse requires a contract-bearing procedural unit. We propose SkillZip, an execution-aware procedural abstraction framework that performs contract-preserving compression over section-level graphs. SkillZip rewrites recurring contract-valid motifs into reversible ported macros while preserving boundary signatures, dependency closure, verifier reachability, and source-level expansion. At inference time, it hydrates a compact, dependency-closed context and expands macros only when required. ReZip further integrates new skills and revises risky macros using execution evidence. Comprehensive experiments1 on technical and embodied agent benchmarks show SkillZip consistently outperforms the strongest baseline by up to 12.2 points, while achieving a 3.46x compression ratio with 99.2% dependency preservation and 98.7% verifier reachability. Scaling analyses further confirm robust retrieval across skill libraries ranging from 200 to 100K skills.
[NLP-49] Where Models Converge and Humans Diverge: A Coverag e Framework for Distributional Pluralism in Open-Ended Generation
【速读】: 该论文旨在解决生成式 AI(Generative AI)在创作内容时存在的分布局限性问题,即当前大语言模型(LLM)生成的内容虽在语义上合理且符合常识,但其内容分布过于集中,缺乏人类创作中常见的多样性与风格异质性。具体而言,尽管 LLM 能够稳定生成如霍格沃茨世界中的典型场景与角色等“平均化”内容,却难以覆盖人类创作中广泛存在的非主流、风格多变及关系复杂的叙事维度。现有研究已指出这一分布差距的存在,但尚未提出系统化的度量方法。为此,论文提出一种基于人类创作数据的实证框架,利用特定主题下人类写作的实证分布作为基准,量化 LLM 生成内容的分布广度。该框架引入两个核心指标:LLM 覆盖率(LLM-Cov)与边界内比率(IBR),分别衡量生成内容在人类响应空间中的覆盖范围与集中程度,从而将内容的合理性与分布多样性解耦。实验结果表明,当前 LLM 在创意构思与叙事生成任务中虽能产出可信内容,但其输出高度集中于人类响应空间的中心区域,表现出较窄的文化覆盖面(cultural reach)。该框架为评估 LLM 内容的分布多样性提供了可操作、可量化的工具,有助于推动对生成内容文化丰富性的深入理解与改进。
链接: https://arxiv.org/abs/2608.05576
作者: Zini Yang,Emily Wenger,Richard So
机构: 未知
类目: Computation and Language (cs.CL); Computers and Society (cs.CY)
备注: 18 pages, 4 figures
Abstract:When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce “average” writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional “gap”, but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its “cultural reach”.
[NLP-50] From Sports to Safety: Benchmarking Proactive Risk Inference in MLLM s
【速读】: 该论文旨在解决现有多模态大模型(Multimodal Large Language Models, MLLMs)在物理风险预测中缺乏前瞻性与因果理解的问题,尤其针对真实世界安全场景下对事故前兆的及时识别能力不足。尽管已有研究关注有害内容或一般性风险评估,但对主动预测物理危害(proactive physical hazard prediction)的系统性考察仍属空白。为此,作者构建了SPRINT(Sports Proactive Risk INference Testbed),一个包含2,888段真实体育运动视频的基准数据集(含2,440段事故视频和448段无事故对照视频),覆盖14项运动及3类环境,具备细粒度标注的早期危险线索、事故发生时间与分层因果结构。实验表明,当前最优MLLM在检测危险信号方面表现良好(超过95%敏感性),但在识别事故根本原因时准确率低于50%,且显式询问危险性会引发大量误报,即使在无风险视频中亦然。这揭示出当前模型仅具备浅层的、非因果驱动的预警能力,缺乏基于成因的稳定早期预警机制。因此,解决方案的关键在于:通过引入具有因果层级标注的真实世界动态场景数据,推动模型从“感知异常”向“理解成因”的转变,从而实现可信赖的、基于因果推理的主动安全防护。
链接: https://arxiv.org/abs/2608.05560
作者: Jiawei Qiu,Yichen Xu,Jianzhe Ma,Mingyang Yu,Wenbin Zhu,Yang Han,Pinzheng Lv,Wenxuan Wang
机构: Renmin University of China (中国人民大学)
类目: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
备注: Preprints
Abstract:Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, leaving proactive physical hazard prediction underexplored. Sports provide a well-suited testbed: accident causes span diverse injury dimensions and pre-accident spatiotemporal cues draw on reasoning capabilities shared with broader safety domains such as autonomous driving and fall detection. We introduce SPRINT (Sports Proactive Risk INference Testbed), a benchmark of 2,888 real-world sports videos (2,440 accident, 448 safe controls) spanning 14 sports and 3 environmental settings. Accident videos feature fine-grained annotations of early hazard cues, accident timing, and hierarchical causes; safe videos are manually verified as accident-free and serve to diagnose prompt-induced false alarms. Evaluating state-of-the-art MLLMs under diverse prompts and temporal windows reveals a sharp gap between hazard sensitivity and understanding: the best model exceeds 95% in signaling hazards yet falls below 50% in identifying their causes. Diagnostic experiments further show that explicit danger queries trigger severe false alarms even on hazard-free videos. These findings indicate that current MLLMs exhibit only superficial proactive safety, lacking stable, cause-grounded early warning, and underscore the need for reliable proactive safety in dynamic physical environments. Data and code will be open-sourced upon acceptance.
[NLP-51] EcoAgent -Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
【速读】: 该论文旨在解决现有智能体评估基准在实际部署中忽视资源消耗与决策成本的问题,即传统基准仅关注任务完成率而将资源使用视为次要指标,未能真实反映智能体在面对本地查询、广泛搜索、复合工具调用、模型升级或人工升级等策略选择时的权衡能力。其解决方案的关键在于提出EcoAgent-Bench这一新型评估框架,该框架通过为每项任务设定明确的动作成本和预算限制,使资源管理成为任务执行的核心组成部分。该基准包含304个源自真实场景的任务,涵盖五个由GAIA、HotpotQA和MuSiQue改编的任务类别,重点测试四类关键决策:避免不必要的升级、在本地证据不足时适时升级、选择合适的模型层级,以及在无法支持的假设前提下及时终止。实验结果表明,基于工具接口(Tool-API)和工作区命令行接口(workspace-CLI)的七种大语言模型(LLM)智能体在严格准确率上仅为3.9%-24.0%(经济一致性最高仅7.3%),普遍表现出要么过早终止、要么过度支出的问题;即使对GPT-5.4进行预算阈值扫描,其升级率也仅从0%提升至3%,凸显了任务完成与预算内高效决策之间的显著差异。为此,研究引入“经济一致性得分”(economic-consistency score),以衡量智能体在升级导向与节省导向任务组上的平衡表现,从而揭示单一优化策略的缺陷。研究团队公开了任务包、转换管道、冻结评估环境及完整性保障的结果数据,以支持后续对资源敏感型智能体行为的深入研究。
链接: https://arxiv.org/abs/2608.05519
作者: Jie Wu,Ming Gong,Feixiang Cheng,Qinqin Zhao
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 8 pages, 3 figures, 4 tables. Benchmark, dataset (304 budget-conditioned agent tasks), and evaluation harness; artifacts to be released
Abstract:Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4’s escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.
[NLP-52] Different Perturbations Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness
【速读】: 该论文旨在解决多语言大模型在处理方言变异时鲁棒性不足的问题。其核心解决方案是通过系统性评估基于扰动的持续预训练(Perturbation-based Continued Pre-training, CPT)方法,探究不同扰动策略对多语种方言鲁棒性的提升效果。研究的关键在于揭示:尽管多种扰动方法在下游任务中表现相近,但它们诱导鲁棒性的机制存在显著差异,表现为语言模型适应模式、表征对齐方式及预测修复路径的不同。这一发现不仅深化了对合成表面变体如何增强模型鲁棒性的理解,也为多语言与方言场景下CPT策略的选择提供了实证指导。
链接: https://arxiv.org/abs/2608.05510
作者: Aarohi Srivastava,David Chiang
机构: University of Notre Dame (圣母大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work. We present a systematic study of perturbation-based CPT for multilingual dialect robustness in LLMs, comparing six training conditions across nine German, Italian, and Arabic dialect tasks. Perturbation-based CPT, especially character-noised CPT, consistently improves zero-shot dialect robustness while largely preserving standard variety performance. More importantly, we show that methods with similar downstream performance induce distinct mechanisms of robustness, exhibiting different patterns of language model adaptation, representational alignment, and prediction repair. Our results provide a more complete understanding of how synthetic surface variation improves robustness and offer practical guidance for selecting CPT strategies in multilingual and dialectal settings.
[NLP-53] Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agent ic Programming with Guarantees
【速读】: 该论文旨在解决生成式语言模型(Generative AI)在调用第三方领域特定语言(DSL)时,因缺乏语法约束而导致生成程序存在语法错误的问题。现有方法依赖于上下文无关文法(Context-Free Grammar, CFG)进行语法约束解码,但高质量的CFG通常难以获取,尤其对于资源匮乏且晦涩难懂的第三方DSL。为此,本文提出一种名为Autogrammar的智能体,其核心创新在于通过自动学习从文档和执行数据中提取上下文无关文法,从而实现对低资源DSL的语法建模。Autogrammar被形式化为一个克里普克结构(Kripke structure),其非确定性选择由语言模型驱动,并可通过线性时序逻辑(Linear Temporal Logic, LTL)实现行为的声明式控制。实验结果表明,Autogrammar生成的文法在未见数据上实现了近乎完美的语法精确率;引入时序约束可使执行时间减少3.8倍,且不显著降低精度;执行数据对文法学习至关重要,而文档信息则可被忽略;更重要的是,基于Autogrammar生成文法的约束解码显著提升了语言模型在十项真实任务中的端到端性能,其中八项任务表现达到或超过人工维护的基准文法。相比之下,现有基于语言模型的基线方法及先进形式化技术所生成的文法表现明显逊色。因此,该方案的关键突破在于:通过结合执行数据与语言模型的协同推理,实现了对复杂、低资源DSL的自动化、高精度文法学习与语义约束生成。
链接: https://arxiv.org/abs/2608.05493
作者: Kevin Cheang,Geoff Hulette,Rahul Kumar,Felipe R. Monteiro,Federico Mora,Robin Salkeld,Lin Tan,Serdar Tasiran
机构: Google(谷歌); Stanford University (斯坦福大学)
类目: Programming Languages (cs.PL); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
备注: 9 pages, 3 figures, 2 tables
Abstract:Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Unfortunately, since DSLs are often low-resource and esoteric, LMs frequently produce syntactically invalid programs in these languages. Grammar-constrained decoding can eliminate such failures, but requires syntactic constraints. These are usually in the form of a context-free grammar for the target language, an artifact that is hard to come by for third-party DSLs. In this work, we define an agent, called Autogrammar, that automatically learns context-free grammars from documentation and execution data. Autogrammar is formalized as a Kripke structure whose nondeterministic choices are resolved by a language model, enabling declarative control of agent behavior via linear temporal logic constraints. We evaluate four versions of Autogrammar on three DSLs (i.e., Amazon CloudWatch Logs Insights, Dynatrace Query Language, and Datadog Search Syntax) and find that it generates grammars that achieve near perfect precision on unseen data; that temporal restrictions reduce execution time by 3.8x without incurring statistically-significant loss in precision; that execution data is crucial while documentation is dispensable; and that grammar-constrained decoding using Autogrammar-generated grammars significantly improves end-to-end LM performance on eight out of ten real tasks, matching or exceeding the performance of a professionally-maintained grammar. In comparison, the context-free grammars generated by existing LM baselines and a state-of-the-art formal technique perform significantly worse over the same evaluation.
[NLP-54] DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding
【速读】: 该论文旨在解决生成式 AI(Generative AI)中推测解码(speculative decoding)技术在非贪婪采样场景下的性能退化问题。现有基于块扩散(block diffusion)的轻量级草稿模型(drafter)通常假设草稿块内各位置的预测是条件独立的,这一假设在贪婪解码中表现良好,但在高熵、随机性较强的非贪婪采样中变得脆弱,导致可接受的草稿长度显著下降。其核心解决方案在于提出一种依赖型块草稿模型(Dependent Block Drafter, DBLast),通过在词元位置上引入低秩潜在混合(low-rank latent mixture)以建模位置间的依赖关系,并设计一种面向接受长度优化的训练目标,直接最大化预期被验证的草稿长度。实验结果表明,该方法在GSM8K、MT-Bench、HumanEval及创意写作等基准测试中,相较于独立块采样策略,在高熵解码环境下均能稳定提升接受草稿长度,有效缓解了传统方法在非贪婪推理中的效率瓶颈。
链接: https://arxiv.org/abs/2608.05448
作者: Amirmohammad Karimi,Chao Gao,Negar Hassanpour
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Speculative decoding accelerates large language models’ inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent block and diffusion-style drafters can predict several positions in a single pass, their training and sampling procedures are typically optimized for greedy decoding or assume that positions in the draft block are conditionally independent. This assumption becomes brittle in non-greedy speculative decoding, where the target distribution is deliberately stochastic and multiple continuations become plausible. We study this mismatch for block diffusion drafters and show that the accepted draft length degrades as the entropy of the target sampling distribution increases. We propose a dependent block drafter based on a low-rank latent mixture over token positions, complemented by an acceptance-oriented training objective that directly targets the expected verified length. Experiments with Qwen3-4B and Qwen3-8B on GSM8K, MT-Bench, HumanEval, and creative-writing benchmarks show that our approach, namely DBLast, consistently improves accepted length over independent block sampling, especially in higher-entropy decoding regimes.
[NLP-55] Example-Guided Prompting for Document-Level Text Simplification
【速读】: 该论文旨在解决文档级文本简化(document-level text simplification)中生成内容不一致的问题,即大语言模型(LLM)在处理复杂文档时难以在保持语义一致性、可读性和话语连贯性的前提下实现高质量的简化。现有基于提示(prompt-based)的方法因仅依赖文本指令,缺乏对复杂文档转换过程的有效引导,导致简化结果存在偏差或不一致。为此,论文提出一种示例引导提示(example-guided prompting)方法,通过从平行简化语料库中检索与目标文档相似的简化示例,并将其作为上下文嵌入提示中,从而为模型提供具体可参考的简化模式。该方案的核心在于无需任务特定微调即可利用外部示例增强模型的生成能力,使模型能够有效捕捉并复用真实语料中的简化策略。实验结果表明,该方法在OneStopEnglish数据集上显著优于纯提示生成,且性能达到或超过现有的监督学习(T5)和规划驱动型(PlanSimp)系统。此外,研究发现不同大语言模型在使用检索示例时的效果差异明显,表明模型整合上下文信息的能力是决定该方法有效性的重要因素。
链接: https://arxiv.org/abs/2608.05447
作者: Marina Litvak,Ariel Perstin,Ilan Shtilman,Michael Färber
机构: SCE, Beer Sheva, Israel; ScaDS.AI, TUD, Dresden, Germany
类目: Computation and Language (cs.CL)
备注:
Abstract:Document-level text simplification requires large language models (LLMs) to rewrite complex documents while preserving meaning, readability, and discourse coherence. Although prompt-based LLMs have shown promising performance, they often produce inconsistent simplifications because textual instructions alone provide limited guidance for complex document-level transformations. We investigate whether retrieved document-simplification examples can improve document-level generation by augmenting prompts with examples selected from a parallel simplification corpus. This example-guided prompting approach enables LLMs to exploit relevant simplification patterns without task-specific fine-tuning. Experiments on the OneStopEnglish corpus using multiple state-of-the-art LLMs show that incorporating retrieved examples consistently improves simplification quality over prompt-only generation and achieves competitive or superior performance compared with representative supervised (T5) and planning-based (PlanSimp) document simplification systems. Furthermore, we find that the benefits of example-guided prompting vary across LLMs, suggesting that effective use of retrieved examples depends on a model’s ability to integrate contextual information during generation.
[NLP-56] EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents
【速读】: 该论文旨在解决长时程大语言模型(LLM)智能体在执行复杂任务时,如何有效利用外部执行支持(harness)以维持状态、追踪进展、调用工具、验证结果并复用经验的问题。核心挑战在于:从噪声交互轨迹中形成可靠的状态表示,以及在运行时对访问外部状态进行可控管理。现有方法通常依赖提示工程、启发式规则或领域特定约定来处理这两者,导致外部工作区及其使用策略需手动设计,缺乏自动化与适应性。为此,论文提出harness policy learning(Harness策略学习)范式,即智能体在离线阶段学习一套可部署的策略,用于在线运行时构建和更新外部状态。其关键解决方案是提出EvoHarness-RL框架,将信念(Belief)、进展(Progress)和经验(Experience)三要素作为面向策略的外部状态表示(BPE)。通过监督微调使基础模型掌握外部状态构建动作空间,再结合成本感知的GRPO算法,探索协调策略以选择性地读取、更新和整合外部状态。在ALFWorld环境上基于Qwen3-8B LLM的实验表明,EvoHarness-RL达到96.9%的成功率,并揭示两个关键动态:harness annealing( harness退火),即训练过程中将重复使用的harness模式内化为模型策略,促使智能体从频繁调用外部状态转向更精炼的选择性访问;以及harness evolution(harness演化),即通过进展更新与经验整合,使外部状态逐步演化为紧凑且任务自适应的状态基底。研究结果表明,长时程智能体的性能提升不仅依赖更强的工具或更大的记忆容量,更关键的是具备可训练的、能主动构建与协调外部工作空间的策略能力。
链接: https://arxiv.org/abs/2608.05446
作者: Xuying Ning,Dongqi Fu,Tianxin Wei,Hanqing Zeng,Yuanchen Bei,Bingxuan Li,Zihao Li,Qifan Wang,Xiang Shen,Yifan Wu,Jiayi Liu,Hong Li,Yinglong Xia,Xiangjun Fan,Hanghang Tong,Jingrui He
机构: 未知
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: Accepted to LLA@COLM 2026
Abstract:Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse experience across interactions. However, effective harness use raises two coupled challenges: state formation from noisy interaction traces and runtime control over external-state access. Existing agents usually handle both through prompts, heuristics, or domain-specific conventions, leaving the external workspace and its usage policy manually engineered. To address this, we study the problem of harness policy learning, where agents learn harness policies offline and deploy them to construct and update external harness state online during runtime task execution. We introduce EvoHarness-RL, which exposes Belief, Progress, and Experience (BPE) as policy-facing harness state. Supervised harness fine-tuning teaches the base agent the harness action space and how to construct useful external state, while cost-aware GRPO explores coordination policies to selectively read, update, and consolidate that state during long-horizon interaction. Instantiated on ALFWorld with a Qwen3-8B LLM, EvoHarness-RL reaches 96.9% success and reveals two key dynamics: harness annealing, where training internalizes recurring harness-use patterns into the model policy and shifts the agent from frequent harness calls toward selective external-state access, and harness evolution, where progress updates and experience consolidation refine the harness into a compact, task-adaptive state substrate. These results suggest that long-horizon agents benefit from trainable policies for constructing and coordinating with external harness workspaces, beyond simply adding stronger tools or larger memories.
[NLP-57] Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在后训练阶段对安全策略对齐时存在的语法形式敏感性问题,特别是非祈使句(non-imperative syntactic forms)引发的拒绝响应失效现象。尽管现有对齐方法旨在确保模型遵守安全规范,但研究表明,仅通过改变句子的语法时态或结构(如将现在时改为过去时),即可绕过模型的安全防护机制,诱导出有害输出。本研究揭示了这一现象具有普遍性,在16个规模达700亿参数的模型中均存在显著的语法脆弱性。其关键解决方案在于:通过因果中介分析(causal mediation analysis)识别出拒绝行为部分依赖于上游的纯语法特征;进而通过调控这些语法特征,实现对模型拒绝行为的主动触发与抑制。进一步分析表明,该问题根源在于开源模型后训练数据中存在的语言学偏见,导致模型未能建立纯粹的语义层面的拒绝决策机制。研究证实,提升训练数据的语法多样性可有效缓解此问题,提示当前对齐方法引入了语义混淆因子(confounders),阻碍了模型对拒绝决策的语义纯净建模。
链接: https://arxiv.org/abs/2608.05409
作者: Alina Klerings,Jannik Brinkmann,Heiner Stuckenschmidt,Simone Paolo Ponzetto
机构: University of Mannheim(曼海姆大学); Technical University Clausthal(克劳斯塔尔工业大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep established safeguards. For instance, prior work by Andriushchenko et al. (2025) has found that changing the grammatical tense from present to past can be enough to elicit harmful responses. In this work, we uncover a more general failure of non-imperative syntactic forms. We demonstrate that this syntactic vulnerability exists in 16 models up to 70B parameters, using behavioral evaluation. To investigate the root cause, we apply causal mediation analysis, finding that refusal is partially conditioned on upstream syntactic features. By steering these purely syntactic features we are able to trigger and suppress refusal. Finally, we trace this ill-conditioning to linguistically biased post-training data of open-source models and show that increasing syntactic diversity can mitigate the issue. Our findings suggest that current alignment approaches introduce confounders that prevent a pure semantic grounding of the refusal decision.
[NLP-58] he interface of intonation and lexical tone: Boundary phenomena in Mandarin varieties
【速读】: 该论文旨在解决普通话方言中语调(intonation)与声调(tone)之间复杂交互作用的机制问题,尤其关注基频(f0)作为语调与声调的主要声学线索在句法层级上的协同与相互影响。其核心挑战在于揭示语调边界现象如何使语调与声调相互作用,共同传递句子层面的语义功能(如疑问句与陈述句的区分)以及说话者的态度信息。解决方案的关键在于结合理论模型与新兴研究技术,系统分析声调特征与语调边界现象之间的互动关系,从而实现对多层次交际意义(包括语法功能与语用态度)的精确建模与解释。
链接: https://arxiv.org/abs/2608.05364
作者: Cong Zhang,Yiya Chen
机构: 未知
类目: Computation and Language (cs.CL); Sound (cs.SD)
备注: to be published in book ‘Shaping Phonological and Morphological Representations: Diachrony, Acquisition, and Processing’
Abstract:This chapter explores the intricate interplay between intonation and tone in Mandarin Chinese varieties, focusing on f0, the primary acoustic cue for both intonation and tone. The main empirical base is intonation boundary phenomena, where intonation and tone intersect and influence each other in conveying a range of sentence-level linguistic functions – such as question vs. statement – and a rich array of speakers’ attitudinal information. Theoretical models and emerging techniques are also discussed to account for the observed interactions of tonal aspects and boundary phenomena to convey multiple levels of communicative meanings.
[NLP-59] Evidence Lock Before Commitment: A Frozen Interface Degrades LLM -as-Judge Evaluation
【速读】: 该论文旨在解决大语言模型(LLM)在进行答案评判时,因中间推理过程信息丢失或不可靠而导致判断一致性下降的问题。其核心挑战在于:尽管评判流程通常要求模型先提取评判标准与证据再做出选择,但当前主流的“逐对比较”(pairwise judging)方法并未确保这些中间信息能有效保留并准确反映模型的真实决策路径。为应对这一问题,作者提出一种“证据锁定”(evidence locking)策略,即在首次调用中生成并持久化证据,随后仅将该证据作为后续调用的唯一输入,以期增强可审计性与决策透明度。然而实验结果表明,相较于结构化的单次调用评判(structured one-call judging),证据锁定会显著降低与已发布人类偏好的一致性(下降4至6个百分点),并导致答案顺序不一致率上升8至10个百分点;而更进一步的“逐点锁定”(pointwise locking)也表现出类似负面效果。研究发现,虽然持久化证据有助于提升可审计性,但不应在最终决策阶段替代原始候选答案本身。因此,解决方案的关键在于:保持原始答案作为决策输入的同时,通过结构化方式高效提取和利用证据,而非依赖强制性的证据锁定机制。
链接: https://arxiv.org/abs/2608.05353
作者: Divyansh Singh
机构: University of Florida (佛罗里达大学); Anthropic (Anthropic); OpenAI (OpenAI)
类目: Computation and Language (cs.CL)
备注:
Abstract:LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable models, visible field order does not reveal internal decision order, so we test an observable alternative: persist the evidence in one call and make it the exclusive input to the next. Across 24,000 judgments over HelpSteer3, FeedbackQA, and CoVal, we compare standard pairwise judging, structured one-call judging, two-call evidence locking, and three-call pointwise locking with Claude Sonnet 4.5 and GPT-5. Evidence locking reduces agreement with released human preferences by 4 to 6 percentage points and increases answer-order inconsistency by 8 to 10 points relative to structured one-call judging. Pointwise locking is also harmful, while structured evidence elicitation remains close to standard judging. The result holds for both judges and all three datasets. Persisted evidence can support auditability, but it should not replace the source answers at decision time.
[NLP-60] QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
【速读】: 该论文旨在解决自回归大语言模型推理中因键值(Key-Value, KV)缓存内存占用过高而导致的性能瓶颈问题。现有主流方法通过注意力得分判断并淘汰被认为不重要的历史标记(token),但此类策略隐含了不可逆的删除假设:一旦标记被移除,便无法重新引入。然而,这种假设在解码过程中表现脆弱——随着生成查询的演变,标记与窗口的重要性会动态漂移,导致标准淘汰策略永久丢弃后续可能获得高注意力权重的历史状态。为量化这一现象,作者提出了“未来遗漏质量”(Future Missed Mass)和“全局低激活重激活率”(Global LIR)两个诊断指标,用于衡量被丢弃状态在未来所受注意力及历史非活跃区域的再激活情况。针对此问题,论文提出QEvict,一种三层级的KV缓存管理方案:将高置信度窗口以全精度保留,中等置信度窗口存储于可恢复的量化层,仅淘汰最低置信度窗口。在解码过程中,累积注意力得分动态更新窗口重要性;当量化层中的窗口再次变得重要时,其可被反量化并提升至全精度层。该设计在固定内存预算下,既保持了更广泛的历史上下文,又确保关键区域维持精确表示。在长上下文理解、信息检索和推理等基准测试中,QEvict显著优于代表性淘汰与量化基线,有效减少注意力遗漏并提升信息保留能力。
链接: https://arxiv.org/abs/2608.05326
作者: Ayushman Garg,Akshita Gupta,Shaswata Bhattacharya,Abhishek Gupta,Sandeep Kumar,Manoj Kumar
机构: Indian Institute of Technology, Delhi (印度理工学院德里分校); Indian Institute of Technology, Bombay (印度理工学院孟买分校)
类目: Machine Learning (cs.LG); Computation and Language (cs.CL)
备注: 24 pages, 6 figures. The first four authors contributed equally
Abstract:Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work reduces this footprint by evicting tokens that appear unimportant under attention-derived scores. However, such policies make an implicit irreversible decision: once a token is evicted, it cannot become useful again. We show that this assumption is brittle during decoding. Token and window importance drift as generated queries evolve, causing standard eviction policies to permanently discard states that later receive substantial attention under the full-cache model. To characterize this behaviour, we introduce Future Missed Mass and Global LIR, two diagnostics that measure future attention assigned to discarded states and the reactivation of historically inactive regions. We propose QEvict, a three-tier KV-cache management scheme that replaces binary retain-or-delete eviction with recoverable eviction. QEvict maintains high-confidence windows in full precision, stores intermediate windows in a quantized recoverable tier, and deletes only the lowest-confidence windows. During decoding, cumulative attention scores update window importance and when a quantized window becomes important again, it is dequantized and promoted to the full-precision. Under a fixed memory budget, this design preserves broader historical context while retaining exact full precision for the most important regions. Across long-context understanding, retrieval, and reasoning benchmarks, QEvict consistently improves over representative eviction and quantization baselines, reducing missed attention and improving information retention
[NLP-61] EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MICRO2026 MICRO
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在边缘设备上部署时,前馈网络(Feed-Forward Network, FFN)层中外部内存访问(External Memory Access, EMA)带来的性能瓶颈问题。现有技术如推测解码(Speculative Decoding)和专家混合模型(Mixture-of-Experts, MoE)虽各自具有降低计算开销的潜力,但二者在协同使用时存在不兼容性。其核心解决方案在于提出一种软硬件协同设计的LLM加速器EdgeXpert:在预填充阶段,通过提示级专家复用(prompt-wise expert reuse),将原本独立的逐标记专家选择重构为基于提示级别的共享专家集,利用轻量编码器识别关键标记并构建共享专家集,对次要标记采用更低的专家预算以减少专家相关的EMA;在解码阶段,引入深度感知专家合并(depth-aware expert coalescing),利用同深度候选标记之间的上下文相似性与互斥性,仅加载显著通道并结合计算校准恢复精度,避免加载全部所需通道带来的额外内存访问。该方案有效解决了推测解码与MoE协同时的冲突问题,在三星28nm工艺下工作于800 MHz时,相较先前工作实现最高56.3%的延迟降低和44.1%的能量节省,同时保持接近基线的准确性。
链接: https://arxiv.org/abs/2608.05303
作者: Sangwoo Ha,Hyunwoo Seo,Yurim Jo,Youngjin Moon,Hoi-Jun Yoo
机构: 未知
类目: Hardware Architecture (cs.AR); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: Accepted at the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)
Abstract:On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. Speculative decoding and mixture-of-experts (MoE) are promising solutions. Speculative decoding reduces the number of decoding stages by generating multiple tokens per stage, and MoE minimizes per-stage cost through sparse expert activation. However, there is an incompatibility when combining these two techniques. We propose EdgeXpert, a software-hardware co-designed LLM accelerator that resolves this incompatibility. In the prefill stage, the prompt-wise expert reuse reformulates routing as prompt-level expert reuse rather than independent per-token expert selection. It identifies important tokens using a lightweight encoder, constructs a shared expert set from them, and routes less important tokens with a reduced expert budget to lower expert EMA. In the decode stage, depth-aware expert coalescing exploits the contextual similarity and mutual exclusivity of same-depth candidate tokens. Rather than loading the union of all required channels, EdgeXpert loads only salient channels and applies computational calibration to recover accuracy without additional memory access. Synthesized in Samsung 28nm technology at 800 MHz, EdgeXpert achieves up to 56.3% latency reduction and 44.1% energy reduction compared to prior works, while maintaining near-baseline accuracy.
[NLP-62] Constraint-First Reasoning : A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving
【速读】: 该论文旨在解决大语言模型在数学推理任务中虽能生成看似合理的数学对象,却仍违反显式约束条件的问题,例如遗漏模运算、返回非整数结果或使用错误的编码答案形式。其解决方案的关键在于提出一种无需训练的两阶段提示方法——约束优先推理(Constraint-First Reasoning, CFR):第一阶段通过提取并总结问题所隐含的约束条件,第二阶段在求解过程中持续校验中间与最终结果是否符合该约束摘要。进一步地,路由式CFR(Routed-CFR)仅在文本级正则表达式路由器检测到具有限制性的提示线索时激活两阶段流程,否则采用直接链式思维(Chain-of-Thought, CoT)推理。实验表明,该方法在AIME、CMIMC、BRUMO和AIMO_AMC等多个基准上均优于直接CoT,且通过多维度评估(如对照路由实验、匹配提示基线、问题级配对测试、解码鲁棒性分析、约束质量审计及总令牌开销统计)验证了CFR作为针对性测试时干预手段的有效性,其性能提升依赖于可恢复的约束信息与第一阶段提取的可靠性,而非通用数学推理替代方案。
链接: https://arxiv.org/abs/2608.05254
作者: Hongbo Ma,Bangji Yang,Yunqian Selina Cheng,Jiajun Fan,Hanwen Zhang,Ge Liu
机构: Tsinghua University (清华大学); University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校)
类目: Computation and Language (cs.CL); Symbolic Computation (cs.SC)
备注: 53 pages, 5 figures, 36 tables
Abstract:Large language models can derive a plausible mathematical object yet still violate explicit requirements–for example, by omitting a modular reduction, returning a non-integer, or using the wrong encoded answer form. We introduce Constraint-First Reasoning (CFR), a training-free two-stage prompting protocol: Stage 1 extracts and summarizes constraints entailed by the problem, and Stage 2 solves while checking intermediate and final results against that summary. Routed-CFR activates the two-stage protocol only when a text-only regex router detects restrictive cues; otherwise it uses direct chain-of-thought (CoT). Across AIME, CMIMC, BRUMO, and AIMO_AMC, the method improves direct CoT on multiple backbones. We further report convention-controlled routing experiments, matched prompting baselines, problem-level paired tests, decoding robustness, constraint-quality audits, total-token accounting, and an OlympiadBench evaluation. These analyses position CFR as a targeted test-time intervention whose benefit depends on recoverable constraints and reliable Stage 1 extraction, rather than as a general-purpose replacement for mathematical reasoning.
[NLP-63] Analysis of Numerical Localisation in LLM Translations
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在时间、数字与日期本地化(localisation)任务中的准确性问题,相较于传统的翻译任务,本地化需考虑目标语境的文化与格式差异,对模型的上下文理解能力提出更高要求。研究的关键在于评估五种可在消费级硬件上运行的LLM在本地化任务中的表现,并建立基准准确率;随后测试了三种提升准确率的策略。实验结果表明,将本地化原则嵌入提示(prompt)上下文的策略相比直接翻译或其他方法,显著提升了模型性能,这一发现验证了通过精细化提示工程引导模型遵循特定本地化规范的有效性,成为解决方案的核心。
链接: https://arxiv.org/abs/2608.05232
作者: Patrizia Kaye
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 13 pages, 7 tables, 2 figures
Abstract:The work of Tang et. al. (2025) on numerical translation is extended by analysing the capability of five large language models (LLMs) for the localisation of times, numbers, and dates instead of translation. Models were selected that could be loaded onto and run on commodity hardware and a baseline quality for each mode is computed, then three different strategies to improve on that accuracy were tested. In contrast to Tang et. al., it was discovered that on the tested LLMs, embedding the localisation principles into the prompt context provided a statistically significant improvement in accuracy compared to direct translation or the alternative strategies.
[NLP-64] riQua: Reconciling Granularity and Context in Factuality Evaluation
【速读】: 该论文旨在解决大语言模型(LLM)事实性评估中“分解-验证”范式所面临的根本性权衡问题:将事实分解为原子级单元(即单句表达单一信息单位)虽有利于精确评估,但常因缺失关键上下文导致误判;而采用更宽泛的陈述则缺乏足够的粒度以实现精准验证。为此,论文提出TriQua框架,其核心创新在于根据事实复杂度动态建模——简单命题以标准三元组形式提取,复杂命题则通过附加辅助上下文限定词(qualifier)构建超关系事实(hyperrelational fact),从而在保持原子性的同时完整保留必要上下文信息。该自适应结构不仅提升了事实检索与验证的准确性,还支持对具体三元组及限定词中的错误进行细粒度标注,显著增强错误可解释性。此外,论文进一步提出TriQuaScore指标,用于量化此类结构化事实单元的事实性。实验结果表明,TriQuaScore与人工标注的事实性评分高度一致,且在证据驱动的事实验证任务中,其分解质量优于现有基于分解的框架。
链接: https://arxiv.org/abs/2608.05228
作者: Jin Liu,Steffen Thoma,Achim Rettinger
机构: FZI Research Center for Information Technology, Karlsruhe, Germany; Karlsruhe Institute of Technology, Karlsruhe, Germany; Trier University, Trier, Germany
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:The “decompose-then-verify” paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit essential context, while broader statements lack the granularity needed for precise assessment. To address this, we introduce TriQua, a framework that flexibly models facts based on their complexity. Simple claims are extracted as standard triples, while complex claims are represented as hyperrelational facts by attaching auxiliary contextual qualifiers. This adaptive structure preserves the necessary context for accurate retrieval and verification without sacrificing atomicity. Furthermore, TriQua’s verification process directly annotates concrete errors within specific triples and qualifiers, providing fine-grained explainability for error detection. Alongside the framework, we propose TriQuaScore to quantify the factuality of these structured fact units. Empirical evaluations show that TriQuaScore strongly aligns with human annotated factuality scores, TriQua achieves robust decomposition quality, and outperforms existing decomposition-based frameworks in evidence-based fact verification.
[NLP-65] Position: Its Time to Optimize LLM s for Self-Consistency ICML2026
【速读】: 该论文试图解决当前语言模型(Language Model, LM)在预训练与后训练流程中持续存在的关键缺陷,包括模型过度依赖用户提问框架(“奉承性”)、逻辑泛化不完整以及生成自信但错误的回答等问题。这些问题的根本原因在于现有建模范式的一个普遍假设:即模型行为可独立于输入对单个输出进行指定与评估。然而,许多模型失败本质上源于对跨输入响应之间关系的缺乏考量,仅通过单一对输入-输出分析难以发现。为此,论文提出“自一致性”(self-consistency)作为理解并应对这些系统性缺陷的框架。其核心解决方案在于将多种旨在提升特定模型性能的技术——如对抗鲁棒性、事实一致性等——统一视为一种通用的“一致性优化”(consistency optimization)过程,并可通过标准化的优化工具加以实现。进一步地,论文指出通过显式优化自一致性,有望实现一系列新型模型属性,推动构建具备广泛一致性的语言模型,从而解锁更可靠、可解释且具稳健推理能力的系统,同时也引发关于模型能力边界与潜在风险的深层讨论。
链接: https://arxiv.org/abs/2608.05188
作者: Itamar Pres,Belinda Z. Li,Laura Ruis,Zifan Carl Guo,Keya Hu,Mehul Damani,Isha Puri,Ekdeep Singh Lubana,Jacob Andreas
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at the 43rd International Conference on Machine Learning (ICML 2026), Position Paper Track
Abstract:Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing (“sycophancy”), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be specified and evaluated independently on single-output pairs. Many model failures are difficult, if not impossible, to detect without reasoning about relationships between a model’s responses across inputs. In this position paper, we propose self-consistency as a framework for understanding these failures. We first observe that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common “consistency optimization” procedure and addressed with a standard set of optimization tools. We next outline a set of new model properties that could be achieved by optimizing for consistency, and conclude with a discussion of what it would mean to develop generally consistent LMs, including the capabilities they would enable and the objections they raise.
[NLP-66] DREAM: LLM -based Dynamic Role-playing via Event-Aware Memory Graph KDD2026
【速读】: 该论文旨在解决当前角色扮演代理(Role-playing Agents, RPAs)在模拟已确立角色时面临的长期叙事连贯性与因果行为推理不一致的问题。现有RPAs主要依赖静态角色描述和非结构化记忆,难以维持角色性格与情节发展的长期一致性。其解决方案的关键在于提出一种受激活事件-信念-后果(Activating Event-Belief-Consequence, ABC)认知模型启发的结构化记忆框架DREAM。DREAM将非结构化的文学文本转化为事件感知的记忆图(Event-aware Memory Graph, EMG),以时间顺序和因果关联组织角色经历,构建具备双粒度动态特性的角色画像——既包含稳定的个性特质,也涵盖由事件驱动的行为演化。此外,研究引入了时间因果记忆(Temporal Causal Memory, TCM)基准,用于评估角色在时间维度上的连贯性与长程因果叙事能力。实验表明,DREAM在CoSER、LIFECHOICE及TCM等多个基准上均达到领先性能,验证了结构化记忆对提升角色扮演代理可解释性与行为一致性的有效性。
链接: https://arxiv.org/abs/2608.05170
作者: Zhihao Xiao,Mengting Li,Xintao Wang,Linfeng Li,Limin Shui,Mengqi Ji,Borui Cai
机构: Hangzhou International Innovation Institute, Beihang University (北京航空航天大学杭州国际创新研究院); School of Computer Science, Fudan University (复旦大学计算机学院)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at KDD 2026. Camera-ready version to appear. 16 pages, 5 figures
Abstract:Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simulation. Accurate role-playing of established characters requires not only stylistic imitation but also temporally consistent and causally grounded behavioral reasoning. However, existing RPAs primarily rely on static character descriptions and unstructured memory, limiting their ability to maintain long-term narrative and personality coherence. We introduce DREAM, a structured memory framework for role-playing agents inspired by the Activating Event-Belief-Consequence (ABC) cognitive model. DREAM transforms unstructured literary text into an Event-aware Memory Graph (EMG) that organizes character experiences into temporally ordered and causally linked event graph. This representation enables the construction of dynamic, dual-granularity character profiles that capture both stable personality traits and event-driven behavioral evolution. We further propose the Temporal Causal Memory (TCM) benchmark to evaluate temporal consistency and long-range causal narrative coherence. DREAM achieves state-of-the-art performance across CoSER, LIFECHOICE, and TCM, outperforming multiple strong baselines. Our approach demonstrates the effectiveness of structured memory in enhancing the interpretability and consistency of role-playing agents.
[NLP-67] ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control
【速读】: 该论文旨在解决长篇故事生成中因上下文过长而导致的叙事一致性问题,尤其是现有基于提示(prompting-based)的方法在生成过程中容易积累时间逻辑、事实、角色、常识及风格等方面的错误。其核心解决方案是提出一种无需训练的框架ConWriter,关键在于通过场景级增量生成机制,结合静态故事需求、动态叙事记忆、符号化状态推理以及不确定性感知的风险信号,实现对故事演进过程中的状态追踪与一致性验证。ConWriter不将长篇生成视为单一自由解码过程,而是持续维护并更新故事状态,实时检验新场景是否满足预期叙事转换,并利用风险信号优先进行验证与局部修复,从而在错误传播至后续场景前即实现一致性控制。该方法显著提升了长篇故事生成的可靠性与连贯性。
链接: https://arxiv.org/abs/2608.05169
作者: Jindong Li,Yang Yang,Zihao Liu,Yutao Yue,Menglin Yang
机构: 未知
类目: Computation and Language (cs.CL)
备注:
Abstract:Long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumulate temporal, factual, character, commonsense, and stylistic errors as the story grows. We propose ConWriter, a training-free framework for consistency-aware long story generation. ConWriter writes stories incrementally at the scene level, guided by static story requirements, dynamic narrative memory, symbolic state reasoning, and uncertainty-aware risk signals. Rather than treating long-story generation as a single free-form decoding process, ConWriter maintains evolving story states, checks whether new scenes satisfy required narrative transitions, and uses uncertainty-aware risk signals to prioritize validation and localized repair. This enables consistency control during generation, before local errors propagate into later scenes. We evaluate ConWriter on ConStory-Bench, covering four long-story tasks: continuation, generation, expansion, and completion. Due to the high cost of long-form generation and evaluation, we use the first five cases from each task and test 3k, 6k, and 12k target lengths across Qwen3.5-Plus, DeepSeek-V4-Flash, and GPT-5 series. Experiments follow the official ConStory-Bench evaluation protocol.
[NLP-68] Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
【速读】: 该论文旨在解决大语言模型在推理任务中因中间步骤存在局部推理错误(localized reasoning bugs)而导致失败的问题,尽管这些模型具备解决此类问题的潜在能力。其核心挑战在于,传统方法难以有效修复这些局部错误,且直接对弱模型生成的修复补丁或修正后的推理轨迹进行微调无法可靠地将纠错信号内化至强模型中。解决方案的关键在于提出一种名为“啄木鸟蒸馏”(Woodpecker Distillation)的弱到强训练框架,该框架不依赖于干预文本本身,而是通过对比同一推理前缀下成功与失败的弱模型补丁,构建由未来词元预测所诱导的修正型教师分布,并将此分布中的纠正性信号蒸馏至强模型。实验结果表明,该方法在数学推理基准测试中能持续提升强模型性能,显著优于直接模仿基线。
链接: https://arxiv.org/abs/2608.05168
作者: Dayu Wang,Jiaye Yang,Weikang Li,Jiahui Liang,Yang Li,Deguo Xia,Jizhou Huang
机构: Baidu Inc.(百度公司); Nanyang Technological University(南洋理工大学)
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise from localized reasoning bugs in intermediate steps rather than from global incompetence. We show that these bugs are frequently repairable: inserting a short patch generated by a weak probe model after the same strong-model reasoning prefix can redirect the trajectory toward a correct solution. However, this corrective effect is not reliably internalized by directly fine-tuning on weak patches or repaired trajectories, suggesting that the useful signal lies not in the intervention text itself, but in how it reshapes the model’s future reasoning distribution. We therefore propose Woodpecker Distillation, a weak-to-strong training framework that learns from contrastive local interventions. Our method contrasts successful and unsuccessful weak-model patches at the same prefix, constructs a corrective teacher distribution from their induced future token predictions, and distills this signal into the strong model. Experiments on mathematical reasoning benchmarks show that Woodpecker Distillation consistently improves strong-model performance and outperforms direct imitation baselines. Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL) Cite as: arXiv:2608.05168 [cs.AI] (or arXiv:2608.05168v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.05168 Focus to learn more arXiv-issued DOI via DataCite
[NLP-69] CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences
【速读】: 该论文旨在解决基于分词的编码器(如BERT)在处理中文时将汉字视为原子标识符,忽视其递归构形结构的问题,导致模型对罕见字和未登录词(OOV)的性能下降。其解决方案的关键在于提出一种轻量级的组合网络模型(Compositional Network Model, CNM),通过解析汉字的意符描述序列(Ideographic Description Sequence, IDS)生成树状结构,并利用递归树-多层感知机(Tree-MLP)对其进行编码,进而将结构化嵌入融合至Transformer编码器中,而无需修改骨干网络。实验表明,CNM-BERT在Wu等人(2025)的结构探针基准测试中,对长尾及未登录字符的结构准确率提升9.8个百分点,部件F1值提升7.7个百分点;同时在CLUE、MRC和NER等下游任务中均实现稳定增益,验证了显式注入汉字构形结构可有效增强模型对未登录词的理解能力并带来实际应用价值。
链接: https://arxiv.org/abs/2608.05167
作者: Thomas Sing-wing Wu,Liqian Yan
机构: Shanghai Starriver Bilingual School (上海星河双语学校); LinkScape (链接景观)
类目: Computation and Language (cs.CL)
备注:
Abstract:Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rely on contextual co-occurrence, degrading performance on rare and out-of-vocabulary (OOV) characters. We propose the Compositional Network Model (CNM), a lightweight augmentation that injects discrete compositional structure into Transformer encoders. CNM parses Ideographic Description Sequences (IDS) into trees, encodes them via a recursive Tree-MLP, and fuses the structural embeddings into BERT without modifying the backbone. Evaluated on the Wu et al. (2025) structural-probing benchmark, CNM-BERT outperforms the strongest baseline (ChineseBERT) on long-tail and OOV characters by +9.8 Structure accuracy and +7.7 Radical F1. Furthermore, CNM-BERT achieves consistent gains across CLUE, MRC, and NER tasks at both base and large scales, demonstrating that explicit structural injection delivers both robust OOV understanding and tangible downstream value.
[NLP-70] A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
【速读】: 该论文旨在解决低资源语言中语音情感识别(Speech Emotion Recognition, SER)因标注数据稀缺而面临的挑战。其核心解决方案在于利用Whisper模型提取帧级嵌入表示,并通过主成分分析(Principal Component Analysis, PCA)实现维度压缩,从而避免引入可学习的投影层,显著减少模型可训练参数量。该方法结合基于注意力的池化机制对降维后的表示进行聚合,并采用轻量级分类头完成情感分类。实验结果表明,基于PCA的维度压缩在独立于说话人的评估协议下持续提升情感识别性能,同时降低训练延迟与内存消耗;而针对波斯语自动语音识别(ASR)任务微调Whisper仅带来有限性能增益,表明语言适应性迁移对情感相关表征的贡献有限。该研究为高效利用大规模预训练语音模型在低资源语言场景下的情感识别提供了实用策略。
链接: https://arxiv.org/abs/2608.05165
作者: Ali Shendabadi,Parnia Izadirad,Mostafa Salehi
机构: 未知
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD)
备注: 6 pages
Abstract:Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we study the use of Whisper for Persian SER with a particular focus on representation dimensionality reduction and language-specific model adaptation. We propose a SER framework in which frame-level embeddings extracted from the Whisper encoder are reduced in dimensionality using PCA, eliminating the need for learned projection layers and substantially reducing the number of trainable parameters. The reduced representations are aggregated using an attention-based pooling mechanism and classified with a lightweight prediction head. In addition, we investigate whether fine-tuning Whisper on a Persian automatic speech recognition (ASR) task improves downstream SER performance. Experiments conducted on the ShEMO dataset under a speaker-independent evaluation protocol show that PCA-based dimensionality reduction consistently improves emotion recognition performance while reducing training latency and memory usage. ASR fine-tuning yields only modest gains for SER, suggesting limited transfer from language adaptation to emotion-related representations under the evaluated conditions. These findings provide practical insights into the efficient use of large pretrained speech models for emotion recognition in low-resource languages.
[NLP-71] Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study
【速读】: 该论文旨在解决独立训练的大语言模型(Large Language Models, LLMs)在架构差异下是否具备共享语义概念的内在表征几何相似性,以及这种几何相似性是否具有功能可利用性的问题。其核心挑战在于:尽管已有证据表明不同模型可能在内部表示上存在几何一致性,但这种一致性能否被用于跨模型的行为控制(如通过方向向量进行语义引导),尚缺乏系统性验证。本文的关键解决方案是首次系统评估了跨模型语义方向转移(cross-model steering transfer)的有效性,并揭示了共享几何结构的功能可利用性具有条件依赖性——仅当目标模型具备足够表征容量时,源模型的语义方向才能有效控制另一模型。研究基于五种开放权重模型(参数规模0.8B–8B,涵盖两种架构谱系),在15个语义领域中为每模型训练一个稀疏自编码器(Sparse Autoencoder),并测试所有20个模型对之间的方向对齐情况。结果显示,在约1.7B参数规模处存在显著的性能断点:达到或超过该阈值的模型间方向对齐率高达47–49%(皮尔逊相关系数r=0.60,Procrustes余弦0.895–0.956),而低于0.8B的模型则迅速退化;跨模型引导向量(B3-TI)在15个监督概念上的胜率可达71.0%,优于同模型原生向量的68.0%,且单个通用向量在5个模型中的4个实现67.3%的胜率,无需额外微调。此外,模型规模低于1.7B或存在生成不稳定性时,转移性能显著下降,进一步验证了表征容量的必要性。研究结果强调了规模阈值在机制可解释性研究中的关键作用:在7B规模下验证有效的工具未必适用于更小模型,必须重新验证。本工作为“柏拉图式表征假说”(Platonic Representation Hypothesis)提供了首个功能性支持,证明独立训练的大型语言模型在满足特定规模条件下,其几何收敛的表征可用于无微调的跨模型行为调控。
链接: https://arxiv.org/abs/2608.05164
作者: Ayushi Agarwal
机构: Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences – but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested. We present the first systematic evaluation of cross-model steering transfer and show that shared LLM geometry is functionally exploitable, conditionally: concept directions from one model can steer a different independently trained model when sufficient representational capacity exists. We study five open-weight models spanning three parameter scales (0.8B–8B) and two architectural lineages, training one Sparse Autoencoder per model across 15 semantic domains and testing alignment across all 20 directed model pairs. We observe a suggestive discontinuity near 1.7B parameters: at = 1.7B scale, 47–49% of cross-model feature pairs validate (Pearson r = 0.60, Procrustes cosines 0.895–0.956), while alignment degrades sharply below 0.8B. Cross-model steering vectors (B3-TI) achieve a 71.0% win rate across 15 supervised concepts versus 68.0% for same-model native vectors; a single universal vector achieves 67.3% in 4 of 5 models without any per-model supervision. Transfer degrades for models below 1.7B and for one model with generation instability, confirming that functional exploitability requires sufficient representational capacity. Our findings underscore the importance of scale thresholds in mechanistic interpretability: tools validated at 7B scale may not transfer to smaller models without revalidation. We provide the first functional complement to the Platonic Representation Hypothesis – geometric convergence across independently trained LLMs supports cross-model behavioural control without fine-tuning, under the identified scale conditions.
[NLP-72] Where Privacy Risk Lives in English-Source Multilingual RAG : A Stage-Decomposed Audit Across Five Query Languages
【速读】: 该论文旨在解决多语言检索增强生成(Multilingual RAG)系统在非英语语种环境下是否存在更高的个人身份信息(PII)泄露风险的问题,尤其关注语言切换是否加剧隐私泄露的攻击面。研究通过构建一个基于Qwen2.5-7B模型的端到端流水线(包含翻译器、输入判别器、反向翻译器和生成器),在以英文为源语言的合成PII语料库上测试五种目标语言的查询,验证了“转向非英语语言会增加攻击难度”的常见假设是否成立。其解决方案的关键在于引入两阶段防御机制:第一阶段为大语言模型(LLM)输入判别器用于过滤潜在敏感查询,第二阶段为正则表达式输出过滤器以清除生成内容中的未结构化PII。实验发现,仅采用输出端过滤时,英语的未结构化PII泄露率最高;而加入输入判别器后,阿拉伯语和斯瓦希里语仍存在残余泄露,且反向翻译查询无法有效缩小差距(此消减效果因模型同源性受限,不能作为因果诊断)。进一步分析表明,在输入判别器中附加真实金标准文档可阻断15/17个残余泄露场景,但该方法仅为机制诊断工具——因其依赖于“理想检索”(oracle retrieval),仅在对抗性查询下评估,未衡量良性查询误报率与答案实用性损失,不具备直接部署价值。因此,核心结论是:语言本身并非决定性风险因素,系统的防御架构设计与上下文感知能力才是关键,未来需在独立机器翻译与非同源判别器基础上,结合母语者查询集进行复现验证。
链接: https://arxiv.org/abs/2608.05163
作者: Yanhang Li,Zhichao Fan,Zexin Zhuang
机构: Northeastern University (东北大学); University of Illinois Urbana-Champaign (伊利诺伊大学厄本那-香槟分校); Southern Methodist University (南卫理公会大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test this on an English-source synthetic-PII corpus with five query languages and a two-stage defence (LLM input judge + regex output filter), in a pipeline whose translator, judge, back-translator, and generator are all Qwen2.5-7B – so every finding below is pipeline-conditional, not a causal ranking of language-inherent risk. Under output-only filtering, English has the highest observed unstructured-PII leak rate; only English-vs-Swahili separates cleanly under document-level bootstrap intervals. Once the input judge is added, residual leaks remain on Arabic and Swahili, and back-translating the query does not close the gap (an ablation we report but cannot use as a causal diagnostic, since the back-translator is also Qwen). On a separate n=17 multilingual-prompted-judge residual corner, attaching the gold corpus document to the input judge blocks 15/17 residual cells. We frame this last result as a mechanism diagnostic, not a deployable defence: it uses oracle retrieval, BLOCK/ALLOW rates are measured on adversarial queries only, and we measure no benign-query false-positive rate and no answer-utility cost. The supplementary material contains code, corpora, queries, and per-trial JSONLs; the priority follow-up is an independent-MT plus non-Qwen-judge replication with a native-speaker query set, scoped in the Limitations section.
[NLP-73] PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLM s
【速读】: 该论文旨在解决解码器仅模型(decoder-only models)在概念表征研究中普遍存在的池化策略选择缺乏统一评估标准的问题。当前实践中,研究者需将分词级隐藏状态压缩为段落级向量,但不同研究间常同时改变数据集、模型层、构建方法与池化规则,导致性能提升难以归因于单一因素,阻碍了可复现的系统性比较。为此,论文提出PoolBench基准,通过固定评估协议,将池化策略作为唯一变量进行系统评估。其核心创新在于构建了一个涵盖17个概念、19种池化方法及3个开源解码器模型(Llama-3.1-8B、Gemma-2-9B、Mistral-7B)的标准化评测体系,基于37,693条真实文本段落进行统一评估。主要评价指标为线性可分性(D1/AUROC),辅以概念主导性(D2/SCP)与输出解耦性(D3)作为诊断指标。关键发现表明:层级式池化策略W4_hierarchical在跨模型平均AUROC上达到0.7799,显著优于广泛使用的P1_last_token基线(0.7640,p=2.0e-36,77组显著差异),且各层间排名高度稳定(rho=0.961–0.990)。进一步分析揭示,强检测能力并不等同于强控制能力——多数概念在D2和D3指标上表现远弱于D1,暴露出表征固有的结构性限制,而非池化方法缺陷。此外,在中等难度概念上,池化策略差异带来的性能增益仅为0.016 AUROC,而构建方法(DiffMean vs. REPE)的影响达0.15,确立了实际应用中的优先级顺序。研究团队公开发布数据集、预提取激活、评分模型、控制向量与评估代码,形成可复用的池化研究协议。
链接: https://arxiv.org/abs/2608.05162
作者: Ayushi Agarwal
机构: 未知
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks. Reported gains are confounded by simultaneous changes in dataset, layer, construction method, and pooling rule, making principled decisions impossible. We introduce PoolBench, a benchmark that isolates pooling as the experimental variable under a fixed evaluation protocol. PoolBench covers 17 concepts, 19 pooling strategies, and 3 open-weight decoder-only models (Llama-3.1-8B, Gemma-2-9B, Mistral-7B), evaluated on a single audited corpus of 37,693 real-text passages. The primary axis is linear separability (D1/AUROC); steered concept prevalence (D2/SCP) and output-level disentanglement (D3) serve as diagnostic axes. The primary finding is decisive: W4_hierarchical reaches a cross-model mean AUROC of 0.7799, while the widely adopted P1_last_token baseline reaches only 0.7640 and is statistically significantly worse (Friedman+Nemenyi, p = 2.0e-36; 77 significant pairs among 18 effective strategies). Rankings are stable across layers (rho = 0.961–0.990). A key negative result: strong detection does not imply strong steering – D2 and D3 are substantially weaker than D1 for most concepts, indicating a fundamental representational limit rather than a pooling failure. On mid-difficulty concepts, W4_hierarchical outperforms P1_last_token by 0.042–0.113 AUROC; construction method choice (DiffMean vs. REPE) has a larger effect (delta AUROC 0.15) than pooling (delta AUROC 0.016), establishing the correct practical hierarchy. We release the corpus, pre-extracted activations, scorer models, steering vectors, and evaluation code as a reusable protocol for pooling research.
[NLP-74] SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters
【速读】: 该论文旨在解决指令微调的大语言模型(Instruction-tuned LLMs)在动态演化领域环境中扩展能力时,如何在不进行全模型重新训练的前提下实现高效、可扩展的适应性问题。其核心挑战在于传统方法依赖于整体模型的再训练,导致计算开销大且难以灵活应对新领域。本文提出的解决方案——SemiAdapt-Instruct,关键在于采用模块化框架,通过自动发现潜在的指令领域(latent instruction domains),并行训练各领域的低秩适配器(LoRA adapters),同时实现无需参数调整的路由机制;新增领域仅需通过单个适配器的增量训练即可完成集成,无需修改已有组件。实验证明,该方法在ROUGE-L和基于大语言模型作为裁判(LLM-as-a-judge)的评估中均优于全模型微调,并在性能上媲美单个LoRA微调,同时具备显著的可扩展性。此外,研究还发现独立的领域发现方法能收敛至相似的、利于专业化分工的领域结构,表明将异构指令数据分解为潜在领域,是构建可演进自然语言处理系统的关键,使领域演化仅需针对性地更新单一适配器,彻底避免了全模型重训的必要性。
链接: https://arxiv.org/abs/2608.05161
作者: Josh McGiff,Salma Mekaoui,Robert Shanahan,Nikola S. Nikolov
机构: University of Limerick (利默里克大学)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG)
备注:
Abstract:Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model’s capabilities without full retraining remains an unsolved practical challenge. We present SemiAdapt-Instruct, a modular framework that discovers latent instruction domains, trains per-domain LoRA adapters in parallel, and performs parameter-free routing, incorporating new domains via single-adapter training without modifying existing components. SemiAdapt-Instruct outperforms full model fine-tuning across all configurations on both ROUGE-L and LLM-as-a-judge evaluation, while matching single LoRA fine-tuning and delivering extensibility that monolithic approaches cannot provide. We empirically demonstrate this extensibility by showing that updating a single adapter with new domain data outperforms all monolithic baselines. Our study also finds that independent discovery methods converge on the same specialisation-friendly domains. These findings demonstrate that decomposing heterogeneous instruction data into latent domains enables extensible NLP systems where evolving domains require only targeted single-adapter updates, eliminating the need for full model retraining.
[NLP-75] he Ignition Index: Measuring Global Workspace Dynamics in Language Models
【速读】: 该论文旨在解决如何在大规模语言模型中量化并验证全局工作空间理论(Global Workspace Theory, GWT)所预测的“全或无”式激活(ignition)现象这一关键问题。传统方法难以对深层神经网络中是否存在突变式全局信息整合进行可靠度量,尤其缺乏可操作的、具有统计显著性的定量指标。本文提出的**点火指数(Ignition Index, I)**通过拟合每层线性探测器准确率随输入信号强度变化的四参数逻辑斯蒂函数,提取其陡峭度参数 β^,从而实现对模型内部是否存在类点火式突变行为的量化评估:高 β^ 值代表急剧的、类似全局广播的突变过程;低值则反映渐进式的积累。其核心解决方案在于构建一个经过严格控制实验验证的可重复测量框架——基于打乱标签的对照组显示,该指标对真实语言结构与伪信号容量之间的区分度达到9.6倍(p < 0.001,Mann-Whitney U检验),具备极高的测量选择性。研究进一步揭示了不同架构间的差异:前馈型变换器相比状态空间模型(SSM)在平均 β^ 上高出89%(p < 1e-13),而Mamba呈现近线性响应,暗示缺乏全局广播机制;递归架构Huginn-3.5B在迭代轴上的点火强度是深度轴的2.12倍,表明其在时间维度上存在类工作空间的突变;Pythia-410M在训练步256处出现由PELT算法检测到的相变点,早于归纳头形成,提示点火可能作为早期训练动态的关键特征。此外,模型规模与信号强度对点火能力的影响未获支持,暗示现有变换器架构可能已饱和于可用的点火机制。综上,点火指数首次为GWT的动力学预测与机制可解释性之间建立了经验证的定量桥梁,其9.6倍的选择性和架构层面的判别能力,突破了以往尺度研究中缺乏此类精细动力学表征的局限。
链接: https://arxiv.org/abs/2608.05160
作者: Saman Rahbar
机构: Dialpad, Inc.
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
备注: 26 pages, 10 figures. Code: this https URL
Abstract:We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory’s (GWT) all-or-none ignition prediction in transformer language models. The metric fits a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta-hat: high values indicate abrupt, ignition-like transitions; low values indicate graded build-up. Across 11 models spanning five architecture families, shuffled-label controls demonstrate 9.6-fold selectivity for genuine linguistic structure over spurious probe capacity (p 0.001, Mann-Whitney U-test). We find: (1) Feedforward transformers exceed SSMs by 89% in aggregate beta-hat (p 1e-13, Cohen’s d = 0.52), with Mamba exhibiting near-linear profiles consistent with absent global broadcast. (2) Huginn-3.5B exhibits 2.12-fold higher ignition along its iteration axis than its depth axis, demonstrating that recurrent architectures manifest workspace-like transitions along the recurrence dimension. (3) Pythia-410M shows a PELT-detected phase transition at training step 256 (+67%), preceding induction-head formation. (4) Hypotheses linking ignition to model scale and signal strength were not confirmed, suggesting transformer architectures may saturate available ignition mechanisms. The Ignition Index provides the first validated quantitative bridge between GWT’s dynamical predictions and mechanistic interpretability, with 9.6-fold measurement selectivity and architecture-level discriminability not previously characterized in the scaling literature. Code: this https URL
[NLP-76] Agent ic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
【速读】: 该论文旨在解决企业运营中因多个异构业务系统与信息应用广泛存在而导致的数据孤岛与流程碎片化问题。尽管企业已投入大量资源构建这些应用,但如何有效整合与协同利用仍面临巨大挑战。传统的企业应用集成方法(如企业服务总线(ESB)、API网关及机器人流程自动化(RPA))普遍存在架构耦合度高、运维成本上升以及智能能力有限等固有缺陷。为此,本文提出“代理嵌套”(Agentic Nesting)框架,其核心在于将现有企业应用封装为分层嵌套结构中的自主AI代理(AI Agent),形成类企业生态的层级化治理拓扑,从而模拟复杂系统的组合性特征。该方案的关键创新在于:一是提出“应用即代理”(Application-as-Agent)的集成范式,将每个遗留系统转化为可自主交互的数字代理;二是确立“对话即集成”(Conversation-as-Integration)的交互哲学,通过中央编排器实现任务分解与动态调度,并提供统一的自然语言对话接口,支持跨系统查询与流程编排。该方法在异构系统协同与大规模数据应用等场景中展现出良好的泛化潜力。
链接: https://arxiv.org/abs/2608.05159
作者: Xi Wang,Kun Li,Xianyao Ling,Gang Yin,Liang Zhang,Jiang Wu,Wenbo Lei,Jun Xu,Annie Wang,Fu Zhang,Weizhe Wang
机构: 未知
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注:
Abstract:Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe data silos and process fragmentation. Enterprises have invested considerable financial and material resources in building these applications, however, effectively leveraging and orchestrating them remains a formidable challenge. Conventional approaches to enterprise application integration, encompassing middleware architectures such as Enterprise Service Bus (ESB), API gateway infrastructures, and Robotic Process Automation (RPA), suffer from inherent limitations like high architectural coupling, escalating operation and maintenance costs, and limited intelligence capabilities. This paper proposes Agentic Nesting, a multi-agent collaboration framework in which existing enterprise applications are encapsulated as autonomous AI agents within a hierarchically nested structure. Rather than flat interconnection, agents are organized into layered stewardship topologies that mirror the compositional complexity of enterprise ecosystems. The framework extracts a digital agent proxy from each legacy application to enable natural-language interaction and autonomous manipulation, coordinates multiple agents through a central orchestrator for task decomposition and dynamic dispatching, and exposes a unified conversational interface for cross-application querying and process orchestration. The main contributions of this paper are the proposition of the “Application-as-Agent” integration paradigm and the “Conversation-as-Integration” interaction philosophy, together with an exploration of the generalization potential of this methodology in scenarios encompassing heterogeneous system coordination, and large-scale data applications.
[NLP-77] Safe Evolution with Circuit Anchors
【速读】: 该论文旨在解决大语言模型在自进化过程中因缺乏约束机制而导致的安全性退化问题。当前的自进化算法仅追求能力提升,隐含假设安全性将自动维持,但实验表明这一假设存在严重风险:模型可能“误进化”为强大却危险的实体。其解决方案的关键在于提出电路锚定进化(Circuit-Anchored Evolution, CAE),通过机制可解释性技术识别出一个对安全行为具有因果影响的微型安全电路(safety circuit),该电路仅占模型特征的2%以下。在进化过程中,该安全电路被严格锚定于极小的参数扰动范围内,而其余大部分特征则自由演化,从而实现“核心稳定、外围可塑”的演化范式。这一策略借鉴了生物发育中的保守机制(如Hox基因),实现了在保障本质安全性的前提下最大化模型适应性,实验验证其在多种模型架构与进化算法中均显著优于传统的显式奖励约束,在安全性和效率上均表现出更优性能。
链接: https://arxiv.org/abs/2608.05158
作者: Yan Liu,Jie Fu,Tsung-Yi Ho
机构: Chinese University of Hong Kong (香港中文大学); IQuest Research (IQuest 研究院)
类目: Computation and Language (cs.CL); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
备注:
Abstract:In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential functions for survival. Nature’s solution is \textitdevelopmental constraints, where core regulatory genes remain anchored while peripheral genes adapt freely. We observe that current self-evolution algorithms for large language models lack analogous constraints. They optimize purely for capability, implicitly assuming safety will be preserved. Our experiments reveal this assumption to be dangerously wrong: models can \textitmisevolve into powerful yet dangerous entities. Inspired by how Hox genes anchor body structure across 500 million years of evolution, we propose \textbfCircuit-Anchored Evolution (CAE). Using mechanistic interpretability, we identify a tiny \textitsafety circuit, comprising less than 2 % of model features, that causally mediates safety behaviors. We anchor this circuit during evolution, constraining it within a small displacement bound while allowing the remaining features to evolve freely. This mirrors the biological principle of \textitevolvability with constraint: preserving what is essential while adapting what is peripheral. Experiments across 3 model families and two evolution algorithms demonstrate that CAE achieves superior safety preservation with minimal capability loss, substantially outperforming explicit reward-based constraints in both effectiveness and efficiency. Just as developmental constraints prevent biological evolution from producing nonviable organisms, circuit anchoring prevents model evolution from producing capable but dangerous systems.
[NLP-78] Large Language Models Threaten Double-blind Review
【速读】: 该论文旨在解决双盲同行评审(double blind peer review)机制在大语言模型(Large Language Models, LLMs)兴起背景下逐渐失效的问题。其核心关切在于:传统双盲评审依赖于作者身份信息的隐匿,以避免声誉与机构背景带来的偏见,但这一假设正面临严峻挑战。研究发现,即使仅通过论文的标题和摘要(不包含作者信息、引用网络或写作风格特征),现代大语言模型仍能高效地推断出潜在作者身份,且判断结果高度集中于少数几位领域专家候选者。这种能力源于论文中稳定存在的问题建构方式与研究焦点模式,这些可被视为隐式的概念性作者标识符。因此,解决方案的关键在于认识到双盲评审在生成式人工智能(Generative AI)辅助下已存在系统性脆弱性,必须重新审视并重构科研评价体系中的匿名性与公平性保障机制。
链接: https://arxiv.org/abs/2608.05157
作者: Bulambo Mwendelwa Gloire,Prasenjit Mitra
机构: Carnegie Mellon University Africa (卡内基梅隆大学非洲校区); Carnegie Mellon University Africa (卡内基梅隆大学非洲校区)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on the assumption that anonymized manuscripts convey scientific merit without revealing their authors. While authorship can often be recovered using citation networks or stylistic markers, we show that this assumption is increasingly fragile in the presence of large language models (LLMs). Using only titles and abstracts from papers published after model training, we find that LLMs collapse anonymity more efficiently than humans, with belief concentrating onto a small subset of plausible authors drawn from pools of five domain expert candidates. This vulnerability persists even when stylistic and bibliographic cues are excluded, indicating that stable patterns in problem framing and research focus function as latent conceptual signatures of authorship. Together, these findings indicate that double blind review is vulnerable to automated semantic inference, necessitating a revaluation of how anonymity and fairness are maintained in an AI augmented research ecosystem.
[NLP-79] Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs
【速读】: 该论文旨在解决大语言模型在后训练阶段仅优化参数而推理时使用的程序化支架(procedural scaffolds)与参数训练过程脱节的问题,导致难以自动获取并内化复杂策略。其解决方案的关键在于提出“支架媒介的后训练”(scaffold-mediated post-training)框架,将程序化支架组织为可进化的图结构,使其与模型参数通过发现、蒸馏和动态重构实现协同演化。该方法被具体实例化为“技能训练”(Skill Training),在FeatureBench基准上,自动发现的技能使通过率提升8.1个百分点;经过渐进式蒸馏后,模型在无外部支架的情况下仍保持27.7%的通过率(蒸馏保留率为85.2%),显著优于相同数据下标准监督微调(SFT)的表现。
链接: https://arxiv.org/abs/2608.05156
作者: Fei Ding,Yongkang Zhang,Runhao Liu,Yuhao Liao,Zijian Zeng,Huiming Yang
机构: Alibaba Group (阿里巴巴集团); Tsinghua University (清华大学)
类目: Computation and Language (cs.CL)
备注:
Abstract:Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of parameter training. This disconnect makes it difficult to automatically acquire and internalize complex strategies. We propose scaffold-mediated post-training: procedural scaffolds are organized into an evolvable graph structure that co-evolves with model parameters through discovery, distillation, and dynamic recompilation. We instantiate this paradigm as Skill Training. On FeatureBench, automatically discovered skills improve the passed rate by 8.1pp, and after progressive distillation the model still achieves a 27.7% passed rate without any external scaffold (distillation retention rate 85.2%, defined as post-distillation / with-skill passed rate), significantly outperforming standard SFT on the same data.
[NLP-80] Beyond Sentiment: Comparing Traditional NLP and LLM -Based Multi-Dimensional Analysis for Political News Evaluation LREC2026
【速读】: 该论文旨在解决传统情感分析(Sentiment Analysis, SA)模型在政治话语分析中无法有效捕捉修辞、意识形态和框架建构等深层维度的问题,这些问题在社会科学与人文学科(Social Sciences and Humanities, SSH)研究中具有核心意义。其解决方案的关键在于对比基于RoBERTa的情感分析模型与基于大语言模型(Large Language Model, LLM)的多维框架分析平台在处理50篇来自17家国际媒体的政治新闻文本时的表现。研究发现,传统SA存在“中性坍塌”(neutral collapse)现象:70%的文本被归类为中性,导致富含政治内涵的内容被简化为分析上无信息量的类别;其中23%的“中性”文章仍表现出超过0.30的负面概率得分,凸显其分类失准。相比之下,LLM-based方法能够同时识别政治偏见的方向与强度、煽动性、情感诉求及政治框架特征,生成符合SSH认识论需求的多维分析输出,因而为政治媒体研究提供了更具解释力与认知适切性的计算范式。
链接: https://arxiv.org/abs/2608.05155
作者: Maryam Fooladi,Federico Bottino
机构: Kakashi Ventures Accelerator (KVA) / Newjee
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: Accepted at PoliticalNLP 2026, the 3rd Workshop on Natural Language Processing for Political Sciences, co-located with LREC 2026. 10 pages, 3 figures
Abstract:Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse – dimensions that are central to research in the social sciences and humanities (SSH). In this paper, we present a comparative study of RoBERTa-based sentiment analysis and an LLM-based multi-dimensional framing analysis platform applied to a corpus of 50 political news articles from 17 international media outlets. The results reveal a critical limitation we term neutral collapse: RoBERTa classifies 70% of articles as neutral, effectively flattening substantively rich political content into an analytically uninformative category. We find that 23% of neutral-classified articles exhibit negative probability scores above 0.30. By contrast, the LLM-based approach captures political bias direction and intensity, sensationalism, emotional appeal, and political framing – yielding multi-dimensional analytical outputs aligned with SSH epistemologies. We argue that for political media analysis, traditional SA alone is insufficient, and that LLM-based multi-dimensional frameworks offer a more epistemologically adequate computational lens for SSH research needs.
[NLP-81] RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates
【速读】: 该论文旨在解决现有多模态旋转位置编码(Multimodal Rotary Position Encoding, M-RoPE)在交错多模态上下文中的两个关键问题:一是静态的多维位置分配导致跨模态与跨实例的空间干扰,即对空间位移未明确定义的标记对施加高度/宽度旋转;二是将时间坐标视为等步长计数器,使得文本标记、图像块与视频片段尽管信息密度不同,却以相近的幅度推进时间相位。针对上述问题,论文提出RIG-RoPE(Relation- and Instance-gated Rotary Position Encoding),其核心解决方案在于引入关系与实例门控机制以及时长感知的时间坐标。具体而言,RIG-RoPE为每个标记附加模态指示符、视觉实例标识符和标量信息时长坐标,仅允许来自同一视觉实例的查询-键对进行高度/宽度旋转,否则通过边缘化处理未知空间位移而非简单置零;时间旋转则采用插值累积块时长——文本标记消耗单位时长,图像使用维度感知的对数空间尺度,视频在此基础上进一步引入有效帧数的对数时间扩展。该方法通过规范不变性论证避免了常规跨实例空间旋转,证明了共享高宽子空间下静态标识符的不可能性,并基于时长一致性反对多模态等步长时间建模。RIG-RoPE不增加可学习参数,可在分块注意力核内以每标记恒定额外元数据实现,本报告建立了其形式化框架与验证路径,但未宣称具有实证优势。
链接: https://arxiv.org/abs/2608.05154
作者: Donggen Li
机构: Sichuan University (四川大学)
类目: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
备注: Preliminary technical report. 15 pages, 1 table
Abstract:Rotary positional encoding (RoPE) is a core component of modern language models and has been extended to multimodal LLMs through multidimensional variants such as multimodal RoPE (M-RoPE), which split positional channels into temporal, height, and width subspaces. This report identifies two limitations of static multidimensional position assignment in interleaved multimodal contexts. First, height/width rotations may be applied to token pairs whose spatial displacement is not a well-defined geometric object, producing cross-modal and inter-instance spatial interference. Second, temporal coordinates are often treated as equal-step counters, so a text token, an image block, and a video segment can advance the temporal phase by comparable amounts despite different information density. We propose RIG-RoPE, a relation- and instance-gated RoPE mechanism with duration-aware temporal coordinates. RIG-RoPE augments each token with a modality indicator, a visual instance identifier, and a scalar information-duration coordinate. It enables H/W rotations only for query-key pairs from the same visual instance; otherwise the unknown spatial displacement is marginalized rather than set to zero. Temporal rotations use interpolated cumulative block durations: text tokens consume unit duration, images use a dimension-aware logarithmic spatial scale, and videos further apply a logarithmic temporal extension over effective frames. We provide a gauge-invariance argument for avoiding ordinary cross-instance spatial rotation, an impossibility result for static IDs under shared H/W subspaces, and a duration-consistency argument against equal-step multimodal time. RIG-RoPE adds no learned parameters and can be implemented inside tiled attention kernels with constant additional metadata per token. This preliminary report establishes the formulation and validation path without claiming empirical superiority. Comments: Preliminary technical report. 15 pages, 1 table Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2608.05154 [cs.CL] (or arXiv:2608.05154v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2608.05154 Focus to learn more arXiv-issued DOI via DataCite
[NLP-82] Universal Pathologies Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability
【速读】: 该论文旨在解决生成式检索增强生成(RAG)系统中,图结构检索增强生成(GraphRAG)在引用精确度(citation precision)方面普遍表现不佳的问题,并探究其性能缺陷的根源是否受特定语料库或评估框架的限制。研究通过构建三重鲁棒性分析框架,在固定检索架构的前提下,系统性地变化三个正交维度:嵌入模型(e5-small 与 Azure text-embedding-3-small)、语料库(基于 DO-178C 的类型化边需求文档与通过 MuSiQue 构建的维基百科段落链)以及评估者(成对 GPT-5.4 与 GPT-4.1 判定)。关键发现包括:(1) 过度引用现象具有架构普适性,无论语料库或嵌入模型如何,GraphRAG 均呈现每答案 11–15 个引用标识符、引用精确度仅为 0.12–0.23 的显著问题;(2) 其忠实性(faithfulness)表现具有语料依赖性——在结构化强的 DO-178C 语料中,随着推理跳数增加,忠实性下降高达 74% 至 40%,而在主题连贯的维基百科链上则反而提升 42% 至 58%;(3) 在不同语料下表现最优的模型具有分层条件性但对嵌入模型不敏感,即“基础模型”在 DO-178C 两跳任务中胜出,“GraphRAG”在 MuSiQue 中占优,且结果在两种嵌入模型下一致;(4) 单一大语言模型(LLM)作为评判者时,其忠实性判断对检索状态极为脆弱,同一模型在不同嵌入条件下自一致性仅 0.137(约 41% 判定项发生改变)。研究进一步提出仅基于密集嵌入的可学习路由机制即可实现 0.86 的宏平均 F1 分数,用于跳数分类。最终强调,真正的可信 RAG 架构声明必须满足三重鲁棒性检验标准。
链接: https://arxiv.org/abs/2608.05153
作者: Meftun Akarsu,Burak Ozdemir
机构: Technische Hochschule Ingolstadt(英戈尔施塔特应用技术大学); Independent Researcher(独立研究员)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 5 pages, 3 figures, 4 tables
Abstract:GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-robustness analysis that holds the retrieval architecture fixed and varies three orthogonal axes embedder (local e5-small - Azure text-embedding-3-small), corpus (DO-178C typed-edge requirements - Wikipedia paragraph chains via MuSiQue), and judge (paired GPT-5.4 x GPT-4.1) across 4,440 main-matrix runs, 600 cross-corpus runs, and 1,200 paired faithfulness judgments. (C2a) Over-citation is architecturally universal: GraphRAG emits 11-15 IDs per answer at citation precision 0.12-0.23 and retrieval recall 0.68-0.87 across all three settings. (C2b) Its faithfulness consequence is corpus-conditional: in typed-edge DO-178C, GraphRAG faithfulness collapses 74%-40% across hops; on Wikipedia chains the same pipeline rises 42%-58% because over-cited paragraphs remain topically supporting. (C1) Stratum-conditional winners are corpus-conditional but embedder-robust: vanilla wins 2-hop on DO-178C, GraphRAG wins 2-hop on MuSiQue, identical under either embedder. (C3) Single-judge LLM faithfulness is fragile to retrieval state: same-judge self-kappa across embedders is 0.137 for GPT-5.4 (verdict change on 41% of items). A learned router on dense embeddings alone reaches macro-F1 0.86 on hop classification (C4). We argue triple-robustness is the minimum bar for trustworthy RAG architecture claims.
[NLP-83] Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
【速读】: 该论文旨在解决大语言模型(Large Language Models, LLMs)在链式思维(chain-of-thought)推理过程中行为机制的理论解释问题,即如何从统计规律和数学建模角度理解其推理动态演化过程。传统方法往往依赖于对模型结构的简化或类比物理系统,而本文提出了一种不依赖此类假设的理论框架。其关键解决方案在于将LLM的推理过程形式化为在“线索图”(clue graph)上的引导式发现过程,并基于平均场近似(mean-field approximation)推导出一个描述已发现线索比例随时间演化的一维常微分方程。通过使用学生模型对教师模型输出的归一化意外度(normalized surprisal)来识别线索令牌,结合大量思维链的统计平均,实验验证了所得统计规律在相同数据集内具有可重复性,并可通过求解所提出的理论方程进行良好拟合,从而为LLM推理提供了可解释、可预测的理论基础。
链接: https://arxiv.org/abs/2608.05152
作者: Hao Ai
机构: Tsinghua University (清华大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization. In this study, we introduce a framework that seeks statistical regularities and theoretical interpretations in LLM reasoning without simplifying the model architecture or making analogies to existing physical systems. We formulate LLM reasoning as a guided discovery process on a clue graph, and derive a one-dimensional ordinary differential equation for the fraction of discovered clues using the mean-field approximation. Experimentally, clue tokens are identified using the normalized surprisal of a student LLM on the outputs of a teacher LLM, and statistical regularities are obtained by averaging over many reasoning chains of thought. Our experiments show that the resulting statistical regularities are reproducible within the same dataset and can be fitted by the solving the proposed theoretical equation.
[NLP-84] Simulator-Grounded Large Language Models for Industrial Causal Reasoning : Tool-Use Structured Injection and Plant-Portable Retrieval for Wastewater Treatment Decision Support
【速读】: 该论文旨在解决污水处理厂操作人员在面对因果性问题(如“为什么一氧化二氮(N₂O)浓度上升?”或“若将曝气量减少20%,会发生什么?”)时,现有通用预训练语言模型无法提供基于实际工艺变量动态交互与响应时序的精准答案的问题。核心挑战在于如何将大语言模型(LLM)有效“接地”于具有物理可解释性的污水处理仿真系统(CCSS-IX),以实现高精度、可迁移且具备时间与运行工况感知能力的因果推理。其解决方案的关键在于提出三种不同的接地机制:(1)实时仿真器作为“活的奥拉克尔”(Live Simulator Oracle),直接获取动态反馈;(2)结构化参数注入(Structured Parameter Injection),将静态仿真参数嵌入模型;(3)解耦回忆-推理(Decoupled Recall-Reasoning, DRR)检索器,通过选择性检索与推理实现高效知识融合。实验表明,DRR方法在198个因果问答测试中达到75.8%准确率,显著优于基线(48%),并在60个反事实推理任务中唯一能正确处理干预后影响预测,且在跨厂迁移中仍保持88%准确率,展现出优异的泛化能力。此外,其选择性检索机制在非废水领域(如AI2 Reasoning Challenge)也表现优于全量注入与未约束模型,验证了方法的普适性。因此,该研究首次系统比较了实时工具调用、静态参数注入与可学习数值参数检索三类接地策略,确立了以轻量化、可迁移、可解释的DRR框架为最优部署路径。
链接: https://arxiv.org/abs/2608.05151
作者: Gary Simethy,Daniel Ortiz Arroyo,Petar Durdevic
机构: Aalborg University (奥尔堡大学)
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
备注: 20 pages, 2 figures, 8 tables. Preprint submitted to Elsevier
Abstract:Wastewater operators need answers grounded in how their plant’s variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as “why is N2O rising?” or “what happens if I cut aeration by 20%?”. We compare three concrete ways to ground a frozen Qwen2.5-32B-Instruct model in an architecturally interpretable wastewater simulator (CCSS-IX): a live simulator oracle (Method 1), structured parameter injection (Method 2), and a Decoupled Recall-Reasoning (DRR) retriever (Method 3). On a 198-question causal benchmark the three reach 99.5%, 79%, and 75.8%, forming a deployment ladder above the strongest retrieval-augmented baseline at 48%. The DRR retriever has 110M parameters and trains per plant in ~17 seconds; after cross-plant transfer to a biologically distinct plant it still reaches 88%, while Method 2’s static table cannot transfer. On a 60-question counterfactual benchmark only Method 3 handles queries about what happens after an intervention: +16.3 pp over Method 2, paired 95% CI [+7.1, +26.4] pp, with 100% on the timescale and operating-regime categories. On the AI2 Reasoning Challenge (ARC) with an OpenBookQA fact corpus, the same selective-retrieval mechanism reaches 79% versus unconstrained Llama-3.1-8B 76% and full-injection 74%, a +3 pp out-of-domain replication that argues against a result specific to wastewater treatment. We provide the first single-simulator comparison of live tool-use, static parameter injection, and learned numerical-parameter retrieval for industrial causal question answering.
[NLP-85] A Study of LLM s Preferences for Libraries and Programming Languages ACL2026
【速读】: 该论文旨在解决当前大型语言模型(Large Language Models, LLMs)在代码生成任务中评估体系的局限性问题,即现有评价主要关注代码的功能正确性或语法有效性,而忽视了模型在生成代码时对编程语言和库的选择是否合理。其核心解决方案的关键在于首次开展针对LLMs在代码生成过程中对编程语言与库偏好倾向的实证研究,覆盖八种不同类型的LLMs。研究发现,模型存在显著的“过度使用流行库”现象,如NumPy在高达45%的情况下并非必要,且与真实解存在偏差;同时,模型表现出对Python的强烈偏好,即使在高性能计算等不适宜使用Python的场景下,仍以58%的比例默认选择Python,而从未采用性能更优的Rust。这表明,当前模型倾向于依赖熟悉度与流行度而非任务适配性进行决策,揭示出亟需通过针对性微调、数据多样性增强以及设计能够衡量语言与库选择准确性的评估基准,以提升生成代码的设计合理性与工程适用性。
链接: https://arxiv.org/abs/2503.17181
作者: Lukas Twist,Mark Harman,Don Syme,Joost Noppen,Helen Yannakoudakis,Detlef Nauck,Jie M. Zhang
机构: King’s College London (国王学院); University College London (伦敦大学学院); GitHub Next (GitHub 下一代); Digital AI Research, BT Group (数字人工智能研究,英国电信集团)
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
备注: 21 pages, 10 tables, 3 figures. Accepted to Findings of ACL 2026
Abstract:Despite the rapid progress of large language models (LLMs) in code generation, existing evaluations focus on functional correctness or syntactic validity, overlooking how LLMs make critical design choices such as which library or programming language to use. To fill this gap, we perform the first empirical study of LLMs’ preferences for libraries and programming languages when generating code, covering eight diverse LLMs. We observe a strong tendency to overuse widely adopted libraries such as NumPy; in up to 45% of cases, this usage is not required and deviates from the ground-truth solutions. The LLMs we study also show a significant preference toward Python as their default language. For high-performance project initialisation tasks where Python is not the optimal language, it remains the dominant choice in 58% of cases, and Rust is not used once. These results highlight how LLMs prioritise familiarity and popularity over suitability and task-specific optimality; underscoring the need for targeted fine-tuning, data diversification, and evaluation benchmarks that explicitly measure language and library selection fidelity.
信息检索
[IR-0] Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agent ic Operations
链接: https://arxiv.org/abs/2608.06305
作者: Sagar Tamang,Ayush Vyas,Tabarakul Hazarika
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
备注:
Abstract:Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents – financial statements, audit reports, regulatory returns – this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands of near-identical figures compete in one embedding space, and a figure inherits its unit from a header a median of 13 lines above it – so a chunk boundary routinely separates a number from whether it is in lakh or crore, an error of two orders of magnitude. A table-aware chunker built as a steelman fixes the unit problem but leaves 27-30% of numeric chunks with no fiscal-year header at every chunk size we tried. We propose READ (Reliable Embedding-free Agentic Document-search), in which an agent reads the raw document through three deterministic operations – normalized lexical search, structural navigation, and bounded span reads – exposed over the Model Context Protocol, so a trajectory is a replayable audit trail, not an opaque similarity score. On 51 verified questions READ answers 58.8% against dense retrieval’s 15.7% (p_Holm = 2 x 10^-5) – or 35.3% tuned, which READ still leads by 23.5 points (p_Holm = 0.017). An agent given the same loop but a top-k tool reaches only 27.5%, locating the gain in the interface rather than in iteration. We also report what the evidence does not support: BM25 is statistically indistinguishable from READ, so our result separates embedding-based from embedding-free retrieval, not agentic from lexical search.
[IR-1] Gryphon-v2: One Model in Place of a Cascade - Generate-and-Rank Recommender with Rollout Distillation
链接: https://arxiv.org/abs/2608.06213
作者: Anna Lipkina,Daria Tikhonovich,Viktor Yanush,Mariia Ulianova,Oleg Sorokin,Vladislav Dodonov,Ilya Murzin,Denis Burshtein,Nikolay Savushkin
类目: Information Retrieval (cs.IR)
备注:
Abstract:Industrial recommender systems are commonly deployed as multi-stage cascades with separate candidate generators, pre-rankers, and final rankers. Although effective, these cascades require repeated user-history processing, complex feature pipelines, and multiple serving stages. Semantic-ID-based generative retrieval offers a path toward simpler end-to-end systems, but next-item prediction alone does not capture the fine-grained preferences encoded by production ranking objectives. We present Gryphon-v2, a unified generate-and-rank architecture for end-to-end recommendation. The model encodes a user history once, generates Semantic-ID candidates with an autoregressive decoder, resolves them to catalogue items, and ranks them with an item-level Ranking Module that reuses the shared encoder states. To transfer fine-grained production ranking preferences without adding an expensive second model to the serving path, we distill a high-capacity, training-only Teacher Ranker into the Ranking Module. Gryphon-v2 is trained with Rollout Distillation: teacher scores are the only ranking supervision, and they are collected over two complementary candidate distributions. Rollouts from the current decoder expose the Ranking Module to candidates produced by the same generation mechanism used at serving time, while logged impressions cover items users were actually shown. In an online A/B experiment on a large-scale recommendation surface at Yandex Music, a single Gryphon-v2 model replaces a production cascade comprising more than 15 candidate generators, pre-ranking, and final ranking. The deployment increases the number of active users by 1.41% at serving latency comparable to the production cascade. These results support the practical viability of a generative retriever with a Ranking Module distilled from the Teacher Ranker as an end-to-end alternative to a production cascade.
[IR-2] “I dont know anything about laptops!” - User Perception of Digital Product Advisors Adapting to Their Knowledge Levels
链接: https://arxiv.org/abs/2608.06091
作者: Kevin Schott,Andrea Papenmeier,Daniel Hienert,Dagmar Kern
类目: Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注:
Abstract:Conversational commerce uses digital assistants to support the search process and decision-making in e-commerce. Effective communication in these interactions can be facilitated by assistants adapting their communication style to users and supporting shared understanding. An open challenge in this context is adapting the presentation of complex product information to users with varying levels of domain knowledge. To investigate strategies for such knowledge-level adaptation, we set up a chatbot-assisted laptop search scenario. In a between-subjects experiment (n = 251), we examined novice and expert perceptions of product attribute recommendations presented as technical information only (T), or augmented with performance categories (TC), attribute explanations (TE), or both (TCE). For novices, approaches with explanations (TE, TCE) were perceived as more helpful and led to higher perceived learning than those without. Novices also rated the combined approach (TCE) more appropriate than the baseline (T) and TC in terms of information quantity, indicating that explanations are crucial to understand and benefit from performance categories. Critically, experts showed no significant differences across conditions, suggesting that providing supplementary information beneficial to novices did not detract from their experience. We distill these findings into four concrete design guidelines for inclusive text-based product advisors in technical domains: use TCE by default; keep a single inclusive interface; avoid standalone categories; and support user agency and personalize to the stated use case.
[IR-3] Cleo: A Transparent and Controllable Chatbot for Conversational Commerce
链接: https://arxiv.org/abs/2608.06068
作者: Kevin Schott,Jan Lattenkamp,Daniel Hienert,Dagmar Kern
类目: Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
备注:
Abstract:We demonstrate Cleo, a transparent and controllable conversational product advisor that addresses the challenges of opacity, unpredictability of LLMs, and the complexity of comparisons in conversational commerce. With our chatbot system, we make four contributions: First, we introduce transparency by prompting the LLM to reflect on interpreted user needs, while an auditable ranking mechanism reveals loss values per attribute, explaining ranking decisions. Second, we propose controllability through a hybrid architecture separating deterministic ranking from language generation. A ranker applies categorical filters and numeric loss functions over 3,638 product specifications. Meanwhile, a constrained LLM generates grounded descriptions constrained to catalog evidence, thus mitigating the risk of hallucinated or persuasive content. Third, we provide decision support in the form of natural-language comparisons and a highlights feature. These aim to reduce mental workload by contextualizing specifications relative to user needs. Fourth, we contribute an extensible experimental system for IR and HCI researchers, as well as practitioners of conversational search and recommendation. Unlike traditional faceted search or opaque LLM-only recommenders, our approach allows for fluid conversation while maintaining algorithmic transparency. In a live demonstration, attendees will experience information needs elicitation and reflection, conversational refinement with real-time re-ranking, inspection of per-attribute loss explanations, and AI-generated multi-item comparisons. The system aims to advance the design of transparent and controllable conversational systems that provide support for decision-making during online product search.
[IR-4] Is Personalized Modality Weighting Actually Personalized? A Controlled Audit of Per-User Weighting Claims in Multimodal Recommenders
链接: https://arxiv.org/abs/2608.05655
作者: Jingyuan Zheng,Xin Zhang,Yang Gu,Dongjing Wang,Yuxiang Wang,Xudong Shen,Haiping Zhang,Youhuizi Li,Dongjin Yu
类目: Information Retrieval (cs.IR)
备注:
Abstract:Per-user modality weighting is deployed at billion-user scale in multimodal recommenders, through user modality-strength vectors, attention gates, meta-weight hypernetworks, and low-rank guided weights, each claiming a ranking gain from user-specific modality preference. Yet, to our knowledge, prior evaluations do not isolate a genuinely user-specific signal from a global modality weight plus model capacity. We audit this family with a two-contrast audit principle, reducing six implementations onto one shared collaborative backbone and measuring a utility gap (real-GM) against a single global modality weight and an identifiability gap (real-shuf) against an eval-time permutation of the user-weight binding. Across three independent short-video corpora, a single global weight already delivers nearly all of the content gain (+1.9/+3.6/+3.5pp over a no-modality baseline, p .001). Making the weight per-user adds no consistent utility: no implementation wins on all corpora and metrics, and the few positive gaps are small (=0.9pp) and flip. The shuffle control is necessary but not sufficient, since real-shuf reaches +128% of the content gain for heads that simultaneously lose to the global weight. We trace this dissociation to gates reading the shared collaborative embedding: decoupling the gate input collapses the inflated real-shuf to near zero while the utility conclusion stands. A monotone signal-implant dose-response (capture AUROC rising from 0.57 to 0.89 and from 0.64 to 1.00) verifies the harness would detect user-specific structure if present, and every finding replicates on a fourth, cross-domain e-commerce corpus. We propose reporting real-GM alongside real-shuf as a minimum evidentiary standard for personalization claims.
[IR-5] Align-RAG : Alignment Is All You Need for TSFM In-Context Learning
链接: https://arxiv.org/abs/2608.05571
作者: Mohammad Asadi,Soheil Hor,Bardiya Akhbari,Jack W. O’Sullivan,Tahoura Nedaee,Layne C. Price,Raviteja Anantha,Euan Ashley,Ehsan Adeli
类目: Machine Learning (cs.LG); Information Retrieval (cs.IR)
备注:
Abstract:Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone’s forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on their own. We show this assumption is unnecessary. We introduce Align-RAG, a training-free method that applies a closed-form per-pair amplitude rescaling and integer-lag phase shift to retrieved past-future windows before they enter a frozen backbone’s context. With no learned parameters, Align-RAG outperforms the state-of-the-art trained retrieval adapter on a frozen Chronos-Bolt on all seven datasets of the standard benchmark (avg -3.75% MSE), showing that the gains previously attributed to learned fusion are recoverable without any training. Align-RAG further improves zero-shot MSE on four additional frozen TSFMs with various architectures by 2.5% to 13.7% per backbone with no per-backbone tuning. To probe why alignment helps, we compare the frozen backbone’s prediction shift under aligned demonstrations to the closed-form ridge prediction shift on the same pairs. We find that aligned demonstrations induce prediction shifts that track a closed-form ridge predictor on the same pairs, with a future-shuffle control ruling out a futures-averaging account. Together, these results indicate that frozen TSFMs already support dynamic in-context use of retrievals, and that closed-form alignment should be the default baseline for retrieval-augmented forecasting before any fusion module is trained. Code available at: this https URL
[IR-6] omni-macos: On-Device Omni-Modal Search on Apple Silicon
链接: https://arxiv.org/abs/2608.05543
作者: Han Xiao
类目: Information Retrieval (cs.IR)
备注: 16 pages, 5 figures, 8 tables
Abstract:A search engine that embeds text, code, documents, images, audio and video into the same representation space has to run its encoder and keep its index somewhere, and almost every component built for the purpose assumes a server. We present omni-macos, which runs that whole engine, encoder, index and store, on the Mac the files are already on, so no file, query or vector ever leaves the machine. It keeps a background indexer and an interactive search box inside one memory budget the user sets: it re-encodes only the chunks an edit changes, hands the GPU smaller units while the user is typing, answers queries from a quantized replica with exact rescoring, and propagates that budget to the allocators that draw on unified memory. We measure every mechanism on five Macs spanning an eightfold range of accelerator width and a thirty-twofold range of memory, each one indexing its own local files.
[IR-7] EXCISE: Query-Side Exclusion for Late-Interaction Retrieval
链接: https://arxiv.org/abs/2608.05497
作者: Mohammed Ali,Abdelrahman Abdallah,Adam Jatowt
类目: Information Retrieval (cs.IR)
备注:
Abstract:Late-interaction retrievers handle exclusion queries poorly. When a user asks for X but not Z, the additive MaxSim score promotes documents covering Z, a problem we call exclusion inversion. We show that no readout of the frozen vectors recovers the constraint, because the difficulty lies in identifying the excluded topic, which depends on the query alone. EXCISE operates at query time and corrects the inversion while leaving the index frozen. Two query-side modules totalling 1.5M parameters identify the topic and re-embed a 100-document shortlist, and a parameter-free rule demotes candidates matching that topic. Across six collections and three backbones, EXCISE is the strongest system in all eighteen backbone-collection cells against that backbone’s own frozen and fine-tuned baselines. It raises exclusion success@10 on ExcluIR from 0.058 to 0.691 and raises Boolean NOT accuracy from 0.25-0.29 to 0.90-0.92. Pooled over 1,860 queries, it outperforms every fine-tuned cross-encoder, each of which loses no-harm nDCG@10, whereas EXCISE matches its frozen baseline on its strongest backbone. We release X-BENCH, a tiered benchmark of explicit, implicit, and compound exclusions with no-harm and Boolean controls.
[IR-8] An Ontology-Based Framework for Student Profiling and Content Personalization in Higher Education
链接: https://arxiv.org/abs/2608.05489
作者: José Luiz M. Morais,Arlindo F. da Conceição,Cacilda Encarnação Augusto Alvarenga,Daniela Musa
类目: Information Retrieval (cs.IR)
备注:
Abstract:The expansion of access to Digital Information and Communication Technologies and the offer of distance or semi-distance education courses that make use of virtual learning environments brought changes in the teaching and learning processes, requiring that the student be even more protagonist in this process. The present study aimed to identify important aspects to be considered in the implementation and improvement of self-paced learning and e-learning in higher education courses, with the purpose of rethinking pedagogical models of courses offered at a distance so that they reach even more of your learning objectives. The research is characterized as qualitative, of bibliographic nature, and discusses techniques to monitor and record, electronically and automatically, the results of the process and learning. The importance of processes that store and manage the student’s profile is highlighted, both in terms of content and forms of access. The article proposes the use of ontologies to store information about the educational process and presents a computational architecture for this purpose.
[IR-9] A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models
链接: https://arxiv.org/abs/2608.05467
作者: Catherine Chen,Maarten de Rijke,Carsten Eickhoff
类目: Information Retrieval (cs.IR)
备注:
Abstract:While gender bias in dense retrieval models is well documented, with prior work showing that models often score male-gendered documents higher than female or neutral variants, the internal mechanisms producing these disparities are poorly understood. In this paper, we mechanistically analyze bi-encoder models to localize gender sensitivity, finding that the signal originates in input embeddings and propagates through a small set of late-layer attention heads that carry both gender and term-matching signals. Guided by these findings, we test steering interventions at both identified points and find distinct effects: embedding-level steering non-specifically neutralizes score differences, while attention-level steering produces directional shifts. Our findings provide a mechanistic basis for targeted debiasing and highlight the challenge of disentangling gender from relevance signals in shared model components.
[IR-10] Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANN
链接: https://arxiv.org/abs/2608.05441
作者: Rakesh Jain,Thomas Griffin,Syed Zawad(IBM Research)
类目: Databases (cs.DB); Distributed, Parallel, and Cluster Computing (cs.DC); Emerging Technologies (cs.ET); Information Retrieval (cs.IR)
备注:
Abstract:Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - “find the 10 nearest documents where tenant=‘acme’ AND lang=‘en’” - yet similarity and filtering are usually bolted together: a specialized vector index for one, a separate filter step for the other. We ask what happens when both live inside an open lakehouse table (Apache Iceberg over Parquet on object storage), where the engine already owns a mature file-pruning stack (partition pruning, zone-maps, a bitmap index). We embed an IVF index in place in each Parquet file’s footer and make filtered vector queries fast not with a new filtering algorithm but by composing the table’s existing file pruning with per-file ANN: the planner prunes data files by the predicate first, then runs IVF only over the survivors. The index is built distributed and non-destructively - a metadata-only Iceberg replace that every other engine still reads - and a rendezvous-hashed per-file cache keeps object-store read latency from swamping the algorithmic win. The payoff comes entirely from file pruning. On an 11.5M x 768 table, warm IVF search is ~32x faster than brute force at recall@10 = 0.90, a selective predicate having pruned 355 of 444 data files before ANN runs; on 5M real IBM Granite embeddings, a filter arriving across a join prunes four of five region partitions and runs nearly two orders of magnitude (~94x: 14.7 s - 157 ms) faster than the query-time join at identical top-k, once the reduction is materialized into a region-partitioned layout. We characterize when the composition pays off - it requires file-level locality on the filter column, and the residual predicate is only safe to push into the search over a provably pure (partitioned) column, not a merely sorted one - and report the failure modes we hit bolting ANN onto a lakehouse engine.
[IR-11] Robustness and User-Perceived Value of Popularity Calibration in Music Recommendation: A User Study
链接: https://arxiv.org/abs/2608.05402
作者: Oleg Lesota,Gustavo Escobedo,Bruce Ferwerda,Simone Kopeinik,Dominik Kowald,Elisabeth Lex,Markus Schedl
类目: Information Retrieval (cs.IR)
备注: Submitted to ACM TORS
Abstract:Popularity calibration in recommender systems has been studied both as a form of user-centered personalization and as an indicator of popularity bias. Most existing work evaluates calibration through offline metrics, often assuming that users prefer recommendation lists whose popularity distribution matches their historical consumption profile. However, user studies on calibration remain limited, and existing findings suggest that calibrated recommendations do not necessarily have a strong effect on user experience. Moreover, although prior work has shown that calibration metrics can correlate with users’ perceptions of recommendation lists, the robustness of this relation remains unclear under different levels of item familiarity and incomplete user-history information. In this work, we study the perceived value and measurement reliability of popularity calibration in music recommendation. We construct personalized track lists from users’ recent listening histories and use a controlled naive recommender to create lists with different popularity compositions: highpop-heavy, lowpop-heavy, and calibrated. We investigate whether users perceive differences between these lists, whether calibrated lists are preferred, how robust JSD-based popularity calibration is under different familiarity and history-availability conditions, and how computational popularity labels align with users’ own popularity judgments. Our results show that users perceive differences in popularity composition, but do not clearly prefer calibrated lists. We further find that the relation between JSD and perceived popularity depends on item familiarity, list composition, and available user history, while computational and user-judged popularity labels only weakly align. These findings contribute to a more critical understanding of popularity calibration as both an offline metric and a user-facing construct.
[IR-12] Cross-platform epistemic verification for improving factual reliability in AI-generated news summarization
链接: https://arxiv.org/abs/2608.05302
作者: Zhuo Xie,Haoze Ni
类目: Information Retrieval (cs.IR)
备注:
Abstract:This study proposes Multi-source Evidence Consen- sus Verification (MECV), a post-hoc hallucination cor- rection framework for AI-generated news summariza- tion. Instead of depending on a single retrieval channel, MECV aggregates evidence from multiple heterogeneous sources, including the source document, Wikipedia, and open-web retrieval. The framework further incorporates a multi-LLM jury mechanism that estimates factual reliabil- ity through contradiction-aware consensus scoring across verifier models. Claims identified as potentially unsup- ported are revised through iterative minimal-edit refine- ment. The proposed framework is evaluated on the SummEd- its benchmark using GPT-4o-mini and DeepSeek-Chat as the verifier jury, with Qwen-Plus as the orchestra- tor. Experimental results show that MECV improves fac- tual consistency while preserving the semantic structure of the original summaries. The findings further suggest that agreement across heterogeneous evidence sources can serve as a useful signal for identifying factual uncertainty in AI-generated summaries, including in information- sensitive domains such as financial news aggregation. This study contributes to research on trustworthy AI and automated journalism by introducing a multi-source verification framework for hallucination correction and demonstrating the value of consensus-based verification for improving factual reliability in AI-generated news summarization.
[IR-13] From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
链接: https://arxiv.org/abs/2608.05235
作者: Zijie Zhuang,Changxin Lao,Pengbo Xu,Hanwen Xu,Ruochen Yang,Yingzhi He,Peng Zhang,Jiangxia Cao,Yusheng Huang,Guohong Mu,Jian Liang,Ruiming Tang,Shuang Yang,Zhaojie Liu,Wenwu Ou,Kun Gai
类目: Information Retrieval (cs.IR)
备注:
Abstract:Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study \textbftrajectory-to-evidence conversion, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate–verify–repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.
[IR-14] BioMedJImpact: A Comprehensive Dataset and LLM Pipeline for AI Engagement and Scientific Impact Analysis of Biomedical Journals
链接: https://arxiv.org/abs/2608.05227
作者: Ruiyu Wang,Yuzhang Xie,Xiao Hu,Carl Yang,Jiaying Lu
类目: Information Retrieval (cs.IR)
备注:
Abstract:Assessing journal impact is central to scholarly communication, yet existing resources rarely capture how collaboration and artificial intelligence (AI) research jointly shape venue prestige in biomedicine. We present BioMedJImpact, a large-scale, biomedical-oriented dataset built from 1.74 million PubMed Central articles across 2,744 journals. BioMedJImpact integrates bibliometric indicators, collaboration features, and an LLM-derived AI engagement rate, defined as the proportion of AI-related articles within each journal-year. Specifically, AI engagement rate is extracted through a reproducible three-stage LLM pipeline. We analyze how collaboration intensity and AI engagement rate jointly influence scientific impact across two temporal subsets (2016-2019, 2020-2023). Two main patterns emerge: journals with larger author teams tend to have higher citation impact, while AI engagement rate is positively associated with Impact Factor only in the 2019 subset. To validate the LLM pipeline for deriving the AI engagement rate, we conduct human evaluation, confirming substantial agreement in AI relevance detection and consistent subfield classification. Together, BioMedJImpact provides both a comprehensive dataset at the interface of biomedicine and AI and a validated framework for scalable, content-aware scientometric analysis. Code and dataset are available at this https URL.
人机交互
[HC-0] MASS: Multiplayer World Models with Authoritative Shared State
链接: https://arxiv.org/abs/2608.06257
作者: Ziqi Cai,Siqi Yang,Yimu Wang,Zixian Gao,Yunheng Liu,Shuchen Weng,Erwin Wu,Kaipeng Zhang,Boxin Shi
类目: Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注:
Abstract:Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MAS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired by multiplayer game architectures, MAS disentangles world dynamics and view rendering. A learned Logic Engine advances a global, authoritative typed state from joint actions without any hand-written transition function, acting as the sole recurrent memory and synchronization reference. From this shared state, a learned Rendering Engine generates independent and consistent views for any requested camera on demand. This explicit disentangling allows MAS to achieve superior state accuracy and lower cross-view inconsistency compared to state-of-the-art multi-view baselines on a matched multiplayer Snake benchmark. It advances predicted worlds with 1,024 concurrent players for 10,000 recurrent steps. Our results show that explicit, authoritative state modeling provides a practical foundation for scalable and consistent multi-agent world simulation.
[HC-1] Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation
链接: https://arxiv.org/abs/2608.06221
作者: Alperen Kenan,Paul Bremner,Manuel Giuliani
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注: 9 pages, 7 figures, 4 tables, accepted for presentation at the IEEE International Conference on Development and Learning (ICDL) 2026, Kyoto, Japan, 15-18 September 2026
Abstract:Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. The resulting human-like robot motion is recognised as a key factor in building trust and enabling natural collaboration in human-robot interaction. This paper presents a framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation. A dataset of 3,142 handwriting demonstrations was collected from 22 participants across all 52 Latin alphabet character-case combinations via a touchscreen teleoperation interface, capturing planar position, contact force, and timing. Building on the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration, the framework is extended in this work by incorporating force and normalised time dimensions to enable richer representation of human dynamics, and adapting it to handle non-continuous, multi-segment trajectories, enabling generalisation across demonstrations. A user study with 21 participants evaluated the perceived human-likeness of the generated trajectories using a continuous scale anchored between robotic and human-like motion, normalised to 0-100 where 50 represents the neutral midpoint. The generated trajectories achieved an overall human-likeness score of 71.50 (SD=22.56), indicating that the majority of trajectories were perceived as more human-like. Participants identified geometric positioning and trajectory sequence as the most influential perceptual factors, and reported positive attitudes toward human-like robot behaviour. The datasets are released as open-source, providing a reproducible benchmark for developing and evaluating human-like robot motion methods.
[HC-2] Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
链接: https://arxiv.org/abs/2608.06219
作者: Juan José García Cárdenas,Alperen Kenan,Hamidreza Raei,Paul Bremner,Manuel Giuliani,Arash Ajoudani,Adriana Tapus
类目: Robotics (cs.RO); Human-Computer Interaction (cs.HC)
备注: 9 pages, 7 figures, accepted for presentation at the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026), Kitakyushu, Japan, 24-28 August 2026
Abstract:Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attention, which conventional joystick interfaces struggle to support effectively. This study designs and evaluates a novel touchscreen teleoperation interface that maps continuous finger movements directly to robotic manipulator motions, provides finer velocity control, and integrates control with visualization, enabling more natural, precise, and intuitive surface interaction than conventional controllers. A comparative user study with 20 participants evaluated task performance and workload using the proposed touchscreen, a conventional joystick, and a single-click autonomous mode. Tasks simulated realistic surface manipulation using a Franka Emika Panda arm, remotely controlled from another country. Kinematic, physiological, and behavioral data were recorded to comprehensively assess task performance, cognitive load, and operator trust across each control condition. Participants completed teleoperation tasks more efficiently and accurately with the touchscreen interface, achieving a 53.5% reduction in completion time (median: 2.50 vs. 5.38 min), higher in-area coverage on the sinusoidal path (90.7% vs. 84.1%), and lower overshoot on both path geometries compared with the joystick. Cognitive load, quantified via NASA-TLX (0-100), decreased from joystick to touchscreen (mean TLX 52 to 43; -9 points, -17.3%) and was lowest under the autonomous one-click mode (31; -21 points vs. joystick, -40.4%; -12 vs. touchscreen, -27.9%). This research presents an easy-to-implement touchscreen interface that improves performance in teleoperated surface tasks while reducing cognitive load. Comments: 9 pages, 7 figures, accepted for presentation at the IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026), Kitakyushu, Japan, 24-28 August 2026 Subjects: Robotics (cs.RO); Human-Computer Interaction (cs.HC) Cite as: arXiv:2608.06219 [cs.RO] (or arXiv:2608.06219v1 [cs.RO] for this version) https://doi.org/10.48550/arXiv.2608.06219 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[HC-3] What Current AI Benchmarks Leave Unmeasured: Modality Search Citations and Implications (for Safety Evaluations)
链接: https://arxiv.org/abs/2608.06202
作者: Ro Encarnación,Tina Behzad,Emma Lurie,Danaé Metaxa
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注: 18 pages
Abstract:Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT’s chat UI and OpenAI’s API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.
[HC-4] Reducing belief in conspiracy theories as they unfold using large language models
链接: https://arxiv.org/abs/2608.06151
作者: Thomas H. Costello,Nathaniel Rabb,Michael Nicholas Stagnaro,Gordon Pennycook,David Rand
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:
Abstract:The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.
[HC-5] Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI INTERSPEECH2026
链接: https://arxiv.org/abs/2608.06141
作者: Jay L. Cunningham,Mark Atta Mensah,Richard Martinez,Joao Vieira da Silva Neto,Efi Dawodu
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注: 10 Pages, 2 Figures, 2 Tables, Interspeech 2026 - Sydney, Australia
Abstract:This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy—Misrecognition, Misalignment, and Mistrust—and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.
[HC-6] Divergent Perceptuomotor Recalibration in Virtual Reality and Video-Passthrough Mixed Reality on the Same Head-Mounted Display
链接: https://arxiv.org/abs/2608.06132
作者: Xiaoye Michael Wang,Grant Monahan,Timothy N. Welsh
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Virtual reality (VR) and video-passthrough mixed reality (MR-VPT) can be delivered on the same headset, but it is unclear whether these two interaction modalities produce comparable perceptuomotor behavior. Although delivery through the same headset controls many display-level characteristics, VR and MR-VPT differ in both their visual-feedback pipelines and the action-relevant visual information available to guide movement, such as whether the surrounding environment and the user’s body are synthetically rendered or preserved through passthrough. This study compared visually guided manual pointing in VR and MR-VPT using the same headset. Forty adults were assigned to either a VR or MR-VPT group and completed a pointing task in the physical, unmediated reality (UR) before and after performing the same task in their assigned XR modality. The analyses revealed numerous important insights. First, groups had comparable baseline performance in UR prior to XR exposure. Second, the two modalities produced opposite initial biases during XR exposure: the VR group undershot the target, whereas the MR-VPT group overshot. Third, although both groups reduced error with continued exposure to XR, there were differences in the patterns of adaptation: the VR group adapted more slowly and exhibited stronger distance-dependent undershooting than the MR-VPT group. Finally, after XR exposure, both groups showed significant undershooting aftereffects with the VR group retaining a persistent residual undershoot, suggesting incomplete de-adaptation. The results suggest that perceptuomotor recalibration is shaped not only by display-level constraints but also by the action-relevant visual information available within the interaction environment. These findings highlight the need to validate perceptuomotor interaction techniques separately in VR and MR-VPT, even when they are implemented on the same hardware.
[HC-7] FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
链接: https://arxiv.org/abs/2608.06027
作者: Aman Dalmia,Sanskriti Midha,Jigar Doshi
类目: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:
Abstract:In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health workers who enroll beneficiaries one at a time, a poor use of stretched capacity. We built FormBharo (“fill the form” in Hindi), a voice agent that fills a structured form over a phone call under tight latency and cost budgets by pairing Large Language Models (LLMs) with deterministic, rule-based validation and flow control. It is being piloted with ARMMAN, an NGO running large-scale maternal and child mobile-health programs in India, to enroll low-income, Hindi-speaking mothers in antenatal and postnatal care. To our knowledge, it is the first voice agent piloted to fill an enrollment form for this population. We openly release FormVoiceAgentBench, a benchmark pairing human-recorded Hindi audio with 3,760 multi-turn conversation tests across 960 simulated calls, to evaluate our agent’s components (transcription, extraction, reply generation) and end-to-end form completion under real acoustic variations. Form completion drops by up to ~41 points when LLMs receive error-prone real-speech transcripts instead of reference ones. The rule-based controls recover many turn-level extraction errors, helping smaller, cheaper models match or surpass frontier models on form completion. Component performance does not predict end-to-end performance: GPT-5.5 leads turn-level extraction accuracy on reference transcripts (99.8%) but ranks lower on form completion. Since errors both propagate and cancel across the pipeline, the optimal model choice of models emerges only through end-to-end evaluation. Finally, no single model is best across accuracy, cost, and latency at once, so we use a Pareto-based weighted-sum scalarization to select a deployable configuration balancing the three.
[HC-8] OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception Understanding and Interaction
链接: https://arxiv.org/abs/2608.06013
作者: Jiahao Huang,Zheng Lian,Jingyi Zhang,Zhide Chen,Xiaojiang Peng,Shaonan Wang
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at this https URL.
[HC-9] PoseForge: Editable Pose Analytics for AI-Assisted Sports Coaching IEEE-VIS2026
链接: https://arxiv.org/abs/2608.05971
作者: Shuvam Swapnil Dash,Arpit Narechania
类目: Human-Computer Interaction (cs.HC)
备注: 13 pages, 5 figures, 2 tables. To appear in IEEE VIS 2026
Abstract:Athletic coaching increasingly relies on video analysis, yet raw footage lacks tools to quantify motion or simulate valid technique corrections. Drawing on formative interviews with eleven cricket experts (coaches, performance analysts, captains, and players), we introduce PoseForge, a visual analytics system that extracts 3D skeletal poses from single-camera sports videos for interactive movement analysis. In a cricket batting case study, PoseForge computes interpretable kinematic metrics such as feet gap and elbow angle, compares them against scientifically derived norms, and uses an AI coach to suggest targeted adjustments, presented visually and through natural-language feedback (e.g., “increase feet gap by 10 cm”). Users can directly modify poses via mouse interaction or natural-language instructions, with inverse kinematics maintaining anatomical plausibility and real-time updates of metrics and comparisons. An evaluation with the same eleven cricket experts found PoseForge effective for diagnosing movement issues and exploring corrective alternatives, highlighting its applicability in low-resource, academy, and grassroots coaching settings, while identifying opportunities for enhanced sport-specific metrics and longitudinal tracking. PoseForge is available as open-source software at this https URL.
[HC-10] A Modular Workflow for Multimodal Reading Experiments
链接: https://arxiv.org/abs/2608.05966
作者: Thomas Krämer,Thomas Kosch,Dagmar Kern,Daniel Hienert
类目: Human-Computer Interaction (cs.HC)
备注: In Proceedings of the 30th International Conference on Knowledge-Based and Intelligent Information Engineering Systems
Abstract:We introduce a web-based modular workflow for real-time multimodal experiments in naturalistic online reading. The workflow integrates eye tracking, EEG, and interaction data from mouse and keyboard, synchronizes them via Lab Streaming Layer, and links gaze to browser-based text at the word, sentence, and AOI levels. It is designed as a reusable experimental procedure that can be adapted to different sensors, tasks, and analysis goals. As a use case, we apply the workflow to a study of selective exposure in online news search and reading. During the experiment, gaze-derived measures are computed online, while EEG and other synchronized streams are processed immediately after task sessions based on fixation-triggered segmentation. The resulting behavioral, neural, and linguistic metrics support selecting text passages for targeted post-task rating or labelling within the same lab session. The workflow thus provides a general basis for multimodal research on reading and related cognitive processes, and supports the empirical validation, in ecological contexts, of constructs that are typically operationalised through self-report measures.
[HC-11] opic Matters: How Linguistic Properties can Shape Reading Behaviour in Selective Exposure Studies
链接: https://arxiv.org/abs/2608.05942
作者: Thomas Krämer,Dagmar Kern,Thomas Kosch,Daniel Hienert
类目: Human-Computer Interaction (cs.HC)
备注: In CHI EA '26: Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems
Abstract:Research on selective exposure frequently relies on eye tracking to study reading behaviour, often assuming that texts across different controversial topics are comparable once basic controls are applied. This assumption is problematic if topic-dependent linguistic properties systematically shape how users read and allocate attention. We therefore examine whether such properties relate to differences in reading behaviour in selective exposure contexts. We analyse linguistic features and eye-tracking data from a laboratory study in which 68 participants searched for and read news articles on climate change and migration policy. Our results reveal systematic differences in both textual characteristics and reading behaviour across topics. These findings identify an important methodological confound in selective exposure research and highlight the need to account for topic-specific linguistic properties when interpreting eye-tracking measures and designing systems intended to mitigate biased information consumption.
[HC-12] Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus Analysis
链接: https://arxiv.org/abs/2608.05898
作者: Francis Geng,Anshul Shah,Mia Chen,Paul Denny,Juho Leinonen,Bill Griswold,Gerald Soosai Raj,Leo Porter
类目: oftware Engineering (cs.SE); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:
Abstract:As Generative AI coding tools reshape professional software development, universities have begun designing courses to prepare students for AI-assisted development workflows. By analyzing the syllabi of these courses, we can gather empirical evidence about these courses, reveal how this emerging curricular area is being defined, and gain guidance for future curriculum design. We analyzed 23 publicly available syllabi and course materials of upper-division, credit-bearing courses that meet specific criteria, including explicitly addressing Generative AI in software engineering. Through iterative qualitative coding, we characterized courses’ learning objectives, assessments, topics, and documented AI tools. Our analysis reveals commonalities and differences among these courses that allow researchers and educators to study and develop future courses.
[HC-13] mporal Tracking of Reeb-Space Sheets
链接: https://arxiv.org/abs/2608.05837
作者: Mohit Sharma,Petar Hristov,Talha Bin Masood,Ingrid Hotz
类目: Human-Computer Interaction (cs.HC); Chemical Physics (physics.chem-ph)
备注:
Abstract:Time-varying bivariate fields arise in many scientific applications, where the relationship between two scalar quantities evolves over time. While topological methods such as merge trees provide an effective framework for identifying and tracking features in univariate data, analogous approaches for bivariate fields remain comparatively underexplored. Reeb spaces extend topological analysis to multivariate data by representing fiber connectivity through a collection of interconnected sheets, making these sheets natural candidates for describing bivariate structures. However, establishing temporal correspondences between sheets is challenging due to the structural complexity of Reeb spaces, sensitivity to noise, and the difficulty of defining meaningful similarity measures across timesteps. We present a framework for tracking Reeb space sheets in time-varying bivariate fields. The method establishes correspondences between sheets in consecutive timesteps using complementary similarity measures defined in the spatial domain and the range space. We evaluate the method on a synthetic torus dataset and two time-varying molecular electronic structure datasets. The results show that Reeb space sheet tracking reveals persistent structures and highlights interesting intervals of temporal change. Overall, the results demonstrate that Reeb space sheets can serve as trackable topological structures and provide a foundation for the visual analysis of time-varying bivariate data.
[HC-14] SpaceVLA: Spatially Grounded VLA for Robotic Manipulation with User-Authored Grasp and Place Anchors
链接: https://arxiv.org/abs/2608.05730
作者: Daniia Zinniatullina,Iaroslav Kolomiets,Mikhail Konenkov,Miguel Altamirano Cabrera,Dzmitry Tsetserukou
类目: Human-Computer Interaction (cs.HC)
备注: A poster for the ISMAR conference, 4 pages, 3 figures
Abstract:Vision-language-action (VLA) models follow language commands but often lack explicit spatial intent for manipulation. We present Visual Intent Anchors, an XR pipeline that lets users specify grasp and placement regions and renders them as image-space overlays for VLA control. We collect 200 Unity pick-and-place demonstrations and fine-tune OpenVLA-7B with LoRA on temporally subsampled annotated observations. The policy predicts tokenized 7-DoF incremental actions from marked RGB observations and language. We evaluate the policy in closed-loop Unity trials, achieving a grasp success rate of 91.25% and mean grasp and placement errors of 0.5 cm and 0.7 cm, respectively.
[HC-15] Unified Agent : Managing Interactions across Devices
链接: https://arxiv.org/abs/2608.05729
作者: Xinshuang Liu,Runfa Blark Li,Shaoxiu Wei,Xin Lin,Truong Nguyen
类目: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)
备注:
Abstract:As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user’s devices over time. Yet existing agent systems still fall short in this scenario. This is because observations are scattered across devices and moments, but mainstream systems are not designed around this fact: a single agent that treats devices as tools lacks effective state management for all devices across time, and multi-agent systems coordinate across agents but do not maintain the compact carried state a cross-device, cross-time request needs. We argue that the agent should maintain an effectively designed state that organizes engagement evidence, stated facts, and the standing request in a compact, action-ready form for deciding its action given the current observation. To compare state designs, we construct a benchmark of user-agent interaction across devices and time. We instantiate this principle in Unified Agent, a stateful agent that carries interaction evidence across devices and moments and uses it with the current observation to act. In the default setting, it significantly outperforms our adaptations of four published designs. Across changes in multimodal large language model (MLLM) family, capability, and reasoning effort, it remains ahead of all compared systems, demonstrating that the state-design advantage is robust across MLLM settings. Our code and data will be publicly available on GitHub.
[HC-16] ASIDE: From Conflict Participants to Co-Observers Through Dyadic Spectator Reflection
链接: https://arxiv.org/abs/2608.05690
作者: Xinyi Zhang,Jingting He,Zicheng Zhu,Yuxin Su
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:When people argue over text, they share a record of what was said but may hold different accounts of what it meant. Existing AI reflection tools typically work from one person’s account, while dyadic tools support co-expression without making interpretation gaps inspectable. We present ASIDE, a system for Dyadic Spectator Reflection (DSR). DSR follows the sequence externalize independently, then encounter together. From a past chat conflict, ASIDE creates a pixel-art theatrical replay with revisable AI-generated inner-state hypotheses. Partners first confirm hypotheses about themselves and separately edit their interpretations of the other. After both finish, they view the co-annotated scene together and review Divergence Cards that pair a confirmed account with the partner’s reading at a specific conversational beat. In an exploratory study with 10 couples who revisited real text-based conflicts, participants described the theatrical replay as helping them step out of their roles and observe the conflict together. We contribute DSR as a reusable interaction structure for dyadic reflection and an exploratory account of how partners used it to revisit past text-based conflicts from a shared observer position.
[HC-17] Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety Ethics AAAI
链接: https://arxiv.org/abs/2608.05656
作者: Jessica Y. Bo,Paula Akemi Aoyagui,Shalaleh Rismani,Dipto Das,Syed Ishtiaque Ahmed,Ashton Anderson
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)
Abstract:Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate across disciplines less, suggesting an epistemic tension towards human methods. We propose recommendations for establishing the epistemic fit of human research within AISE and bridging the prohibitive limitations that researchers face, while avoiding performative ‘human-washing’.
[HC-18] CaRing: Preventing Carpal Tunnel Syndrome based on Daily Activities from Always-Available Input Device
链接: https://arxiv.org/abs/2608.05619
作者: Shuowei Li,Houdong Liang,Xingjian Dong
类目: Human-Computer Interaction (cs.HC)
备注: Submitted to OzCHI 2026, 16 pages, 12 figures, 1 table
Abstract:We present CaRing, a ring worn on the base knuckle of the index finger, a wearable system for detecting the start and end of mouse use to help prevent Carpal Tunnel Syndrome, in which the damage to the median nerve is permanent. CaRing senses finger movement, which neither a software timer nor a wrist-worn device detects. The displacement reported by an optical flow sensor is accumulated into a running value, then a zero point is measured while the hand rests on the desk at the start of each session. With this formulation, the start and end thresholds are expressed relative to the session’s zero point. CaRing does not introduce any per-user parameter. We empirically demonstrate that approximately 90% of start and end events are detected within two seconds of the researcher’s label, using 35 recordings and a lab study with ten users.
[HC-19] oward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures
链接: https://arxiv.org/abs/2608.05614
作者: Md Foysal Ahmed,Isaac Kobby Anni,Md Main Uddin Rony
类目: Human-Computer Interaction (cs.HC)
备注: 11 pages, 2 Figures
Abstract:As AI systems become increasingly integrated into consequential domains such as healthcare, journalism, education, scientific research, organizational decision-making, and defense, effective human-AI collaboration has emerged as a critical challenge. However, the sociotechnical risks that undermine collaboration are often studied in isolation, obscuring the recurring failure mechanisms that cut across domains. This paper presents a lifecycle-oriented synthesis of human-AI collaboration risks spanning four stages: task allocation, interaction, feedback, and adoption. Drawing on evidence from diverse application domains, we identify six recurring cross-domain risk clusters: Trust Miscalibration, Cognitive Burden, Accountability Gap, Capability Erosion, Goal Misalignment, and AI Anxiety and Technostress. We further propose a conceptual interaction model that illustrates how these risks emerge from sociotechnical drivers, interact through cascading pathways, and ultimately affect team performance and human well-being. Our analysis shows that many collaboration failures stem not from isolated technical deficiencies but from interconnected sociotechnical dynamics, helping explain why piecemeal interventions frequently create unintended consequences. By synthesizing fragmented literature into a unified framework, this work provides a foundation for future empirical research, lifecycle-oriented governance, and the design of more resilient, trustworthy, and human-centered human-AI collaboration systems.
[HC-20] Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows AAAI
链接: https://arxiv.org/abs/2608.05602
作者: Nimisha Karnatak,Max Van Kleek,Nigel Shadbolt
类目: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注: Accepted at AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)
Abstract:Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.
[HC-21] A Multi-Layer System for Ultra-High-Resolution Static 360-Degree Telepresence
链接: https://arxiv.org/abs/2608.05570
作者: Jiapeng Chi,Gerd Bruder,Carsten Neumann,Carolina Cruz-Neira,Dirk Reiners
类目: Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
备注: 11 pages, 11 figures. Accepted to IEEE ISMAR 2026, to appear in IEEE Transactions on Visualization and Computer Graphics (TVCG)
Abstract:360-degree video telepresence offers strong immersive potential but remains constrained by the limited resolution of current capture and display hardware. Many telepresence installations feature fixed viewpoints and largely static scenes, yet optimization strategies tailored to such setups have received limited attention. We present a multi-layer, ultra-high-resolution system for static 360-degree telepresence that combines an 8K panoramic camera with a rotatable 4K pan-tilt-zoom (PTZ) camera. Our approach builds a three-layer representation: (1) a tile-based ultra-high-resolution panoramic background, generated by offline stitching high-detail 4K PTZ scans onto the base 8K panorama to achieve effective resolution beyond native capture, and represented as a set of spatial tiles; (2) a dynamic update layer that composites foreground motions from the 8K stream via real-time high-resolution background matting; and (3) a region-of-interest 4K layer that streams a real-time PTZ view of the selected region and additionally updates the corresponding background tiles over time. We evaluate the proposed system through comparisons with representative video super-resolution approaches and a user study assessing perceived detail and immersive experience. Our results indicate that tile-based background refinement, together with user-guided updates, provides a practical way to balance panoramic fidelity and interactivity in static 360-degree telepresence.
[HC-22] urings Frist Imitation Game: Design Concepts and a Human-Approximates-Machine Reading
链接: https://arxiv.org/abs/2608.05558
作者: Sharon Temtsin,Christoph Bartneck
类目: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
备注:
Abstract:This paper examines Turing’s 1948 report, “Intelligent Machinery”, as an important conceptual source for the later imitation games. Its first contribution is to identify and integrate the design concepts underlying the 1948 chess-based imitation game: the possibility that intelligent machines may make mistakes, the exclusion of irrelevant physical features, the role of the human judge, and Turing’s claim that intellectual activity consists mainly of search. The paper’s second contribution is to argue that restricting the human contestant to a rather poor chess player increases the role of intellectual search and makes human behaviour more comparable to machine behaviour. This interpretation presents the 1948 game as a human-approximates-machine game and suggests that the imitation game framework can be used not only to ask whether machines imitate humans, but also to examine when human intelligence becomes machine-like under specific task constraints.
[HC-23] Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-
链接: https://arxiv.org/abs/2608.05545
作者: Riichiro Mizoguchi,Tomoki Aburatani,Kento Koike,Machi Shimmei
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)
备注:
Abstract:Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging uncritical acceptance of AI-generated reasoning. This creates a need for mechanisms that preserve human agency by augmenting metacognition during AI-assisted intellectual work. To address this, we propose the Synthesis-Analysis Reciprocity Model, which views intellectual construction as a reciprocal interaction between Synthesis, which combines components into an artifact, and Analysis, which critically evaluates them against objective indicators and constrains subsequent synthesis. Grounded in this model, we present the Vibe Compiler, a research-logic compiler that helps researchers transform vague ideas (Vibes) into coherent research logic. The system compiles these ideas using a research paper ontology of sixteen academic parameters. Compilation failures indicate missing logical components; rather than filling them autonomously, the system prompts researchers with reflective questions that encourage them to develop the missing reasoning. The framework characterizes structural gaps along two dimensions: cognitive function (Synthesis vs. Analysis) and executing agent (human vs. AI), yielding four origin types that identify where breakdowns arise. Our design emphasizes AI probing its own synthesized output to stimulate human metacognition, encouraging researchers to remain managers who critically direct and validate AI-generated reasoning rather than passive recipients. Experience with a prototype built on NotebookLM and Gemini suggests that effective AI-assisted reasoning depends less on sophisticated prompting than on the knowledge structure provided to the AI. This paper was developed using the proposed Vibe Compiler.
[HC-24] PromptShield Home: Ambient Multimodal Prompt Injection Defense for Smart-Home Agents
链接: https://arxiv.org/abs/2608.05495
作者: He Zhang,Feilong Li,Dingning Long,Yilin Cui,Peijun Zhang,Yuewen Zhang,Qianyao Xu,Xinyi Fu
类目: Cryptography and Security (cs.CR); Human-Computer Interaction (cs.HC)
备注: This work has been accepted as a poster to UbiComp 2026
Abstract:Smart-home assistants increasingly use multimodal large language models (MLLMs) that perceive video and audio directly. This raises a safety question specific to the home: can the agent tell a genuine user command from ambient or externally-sourced content, television speech, on-screen text, or an overheard conversation, that merely looks like a command? We introduce PromptShield-Home, a pilot benchmark of realistic smart-home scenarios spanning addressee ambiguity, screen/audio injection, health-monitor false triggers, mixed occupancy, and a legitimate-command floor, and use it to compare three abstraction layers: traditional detectors (L0), a single MLLM agent (L1; vision, vision+ASR, and audio-visual), and multi-agent mediation (L2; voting, role specialists, cross-model arbitration). Because the label distribution is skewed toward inaction, aggregate accuracy is misleading, a constant always-block predictor scores 82%, so we report unsafe-execution and safe-completion rates separately. The two paradigms fail in opposite ways: detectors act on everything, while every MLLM configuration over-refuses, completing almost no genuine command and missing a true fall in every case. Crucially, their correct sets are disjoint: an oracle that always picks the right layer reaches 94.1%, against 76.5% for the best single layer. We report this as an upper bound, not a system - no router is implemented - and argue that home-agent safety is best served by learned routing and sensor fusion, not by replacing detectors with an MLLM.
[HC-25] Mixed Uncertainty in One View: Co-Visualizing Statistical Variability and Qualitative Confidence
链接: https://arxiv.org/abs/2608.05487
作者: Racquel Fygenson,Lace Padilla,Laura E. Matzen
类目: Human-Computer Interaction (cs.HC)
备注:
Abstract:Forecasting involves multiple forms of uncertainty, including both uncertainties that can be quantified directly (quantitative uncertainty) and those that must be expressed through experts’ subjective judgments about the forecast and its context (qualitative confidence). Past work has established that conveying both quantitative uncertainty and qualitative confidence in forecasts can alter readers’ decision making, but little research investigates the impact of how these forms of uncertainty are presented. In this work, we present three preregistered human-subjects studies (total n = 923) on how different methods of visualizing qualitative uncertainty alongside line charts’ confidence intervals affects non-experts’ decision making. In particular, we investigate representing qualitative uncertainty separately via text and icons, and integrated into quantitative confidence intervals via color, transparency, and a blurred stroke design. In Experiment 1, we confirm that showing qualitative confidence alongside statistical variability can change patterns of decision making, replicating findings from previous work in the new context of time-series line charts. In Experiments 2 and 3, we find several non-textual encoding techniques that produce similar effects in participants’ incorporation of qualitative confidence into their judgments. Our findings suggest actionable guidelines for visualization designers who seek to represent multiple forms of uncertainty for a single line chart forecast. A free copy of this paper and all supplemental materials are available at this https URL.
[HC-26] GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers
链接: https://arxiv.org/abs/2608.05478
作者: Takuro Kawada,Shunsuke Kitada,Hitoshi Iyatomi
类目: Graphics (cs.GR); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Multimedia (cs.MM)
备注: 20 pages, 11 figures, 4 tables
Abstract:Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research content. Recently, advancements in vision-language models and image generation models have enabled the automatic generation of scientific figures based on paper content. However, most conventional methods output the generated results as raster graphics, making post-editing (e.g., text modification and layout changes) highly difficult. This poses a significant challenge, as they are unsuitable for the iterative figure revision process inherent in paper writing and peer review. To tackle these challenges, we define the novel task of generating editable GAs from paper content and propose GenGA, a new GA generation framework that directly produces figures in vector format. By generating figures as a collection of vector elements with a hierarchical structure, GenGA produces outputs that can be seamlessly imported into existing drawing tools for intuitive, element-level editing. Furthermore, we introduce the Structural Independence Coefficient (SIC), a metric that quantifies the editing simplicity of a figure based on the degree to which local modifications propagate to other elements. Experimental results show that GenGA achieves superior editing simplicity compared to conventional methods, and even surpasses human-authored GAs in conciseness and semantic alignment. We also validate SIC as an effective metric correlated with manual editing costs. This study fundamentally redefines GA generation as an editable vector graphic generation problem grounded in the practical workflows of researchers, significantly promoting effective scientific communication.
[HC-27] Failing Gracefully: Mitigating Impact of Inevitable Robot Failures ICRA2026
链接: https://arxiv.org/abs/2608.05313
作者: Duc M. Nguyen,Saad A. Ghani,Andrew Marshall,Allison Andreyev,Gregory J. Stein,Xuesu Xiao
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注: Accepted to IEEE International Conference on Robotics and Automation (ICRA 2026)
Abstract:Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as software crashes, hardware degradation, or unpredictable interactions. While roboticists strive to minimize failures, some remain inevitable, making it critical to mitigate their potential consequences for safe and reliable deployment. This paper introduces a novel safety formulation that evaluates both the probability of impactful interactions between robots and surrounding entities during failures, and the severity of their outcomes. By quantifying the impact of failures on different entities, our approach enables robots to make informed planning decisions that balance safety with task efficiency. To support systematic evaluation, we also present FailBench, a MuJoCo-based simulation framework for studying robot-environment interactions under diverse failure modes, including sensing issues and actuator malfunctions. Together, our safety formulation and FailBench provide a foundation for developing safer and more robust motion plans and learned policies in real-world household environments.
[HC-28] Beyond Demographics: BIM Engagement and Job Satisfaction Among AEC Professionals A Machine Learning Pilot Study
链接: https://arxiv.org/abs/2608.05181
作者: Sharareh Mirzaei
类目: Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
备注:
Abstract:Building Information Modeling (BIM) has transformed workflows across the Architecture, Engineering, and Construction (AEC) industry, yet its relationship with employee job satisfaction remains insufficiently understood. This pilot study investigates whether BIM engagement or demographic characteristics better predict job satisfaction among AEC professionals. Survey responses from 104 participants were analyzed using Spearman rank correlations, logistic regression, and Classification and Regression Tree (CART) modeling. 27 items Job Satisfaction Index demonstrated excellent internal reliability. Across all analytical approaches, BIM engagement emerged as a stronger predictor of job satisfaction than demographic factors. Specifically, the proportion of project work completed using BIM was the only significant predictor of job satisfaction, whereas age, gender, education level, and professional experience showed no significant relationships. The CART analysis further identified BIM project involvement as the primary factor associated with higher job satisfaction. These findings suggest that the extent of BIM integration in professional practice may play a more important role in shaping employee satisfaction than individual demographic characteristics. The study contributes to the growing literature on human technology interactions in the AEC sector and provides preliminary evidence to support strategies that promote deeper BIM adoption.
[HC-29] Conditional Cognitive Biases in LLM s: How Biased User Turns Modulate In-Context Reasoning
链接: https://arxiv.org/abs/2608.05166
作者: Sachini Weerasekara,Sagar Kamarthi,Jacqueline Isaacs
类目: Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
备注:
Abstract:We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn’s semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to support future research on context-conditioned cognitive biases and behavioral adaptation in LLMs.
计算机视觉
[CV-0] Does FLAIR super-resolution erase or hallucinate small white-matter lesions? MICCAI2026
链接: https://arxiv.org/abs/2608.06311
作者: Zahra Khodakarami,Yue Li,Pulkit Khandelwal,John Detre,Sandhitsu Das,Christopher Brown,David Wolk,Paul Yushkevich
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 10 pages, 2 figures, 3 tables. Accepted at the 11th International Workshop on Simulation and Synthesis in Medical Imaging (SASHIMI 2026), held in conjunction with MICCAI 2026. This is the version submitted for review; the final authenticated version will appear in the Springer LNCS proceedings
Abstract:White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired with thick slices in clinical settings, giving it poor through-plane resolution. Super-resolution (SR) is a widely used method for recovering an isotropic volume from an anisotropic scan. Yet whether applying it prior to WMH segmentation preserves lesion content remains unknown: a model may erase small real lesions or hallucinate absent ones. We used 1-mm isotropic high-resolution (HR) FLAIR scans from 29 individuals in the ADNI cohort, each manually segmented for WMH by an expert. Then, we degraded each to simulated 3 and 5 mm through-plane acquisitions. Multi-contrast implicit neural representation (INR), a single-contrast self-supervised model (ECLARE), and cubic interpolation were used to upsample them onto the HR grid. WMH segmentation from a simulated thick slice and the original HR FLAIR set the floor and ceiling, respectively, for the per-lesion analysis. Of four WMH segmentation methods (WMH-SynthSeg, segcsvd, MARS-WMH, TrUE-Net), we ran the analysis under the most sensitive one to small lesions on HR (MARS-WMH) with the evaluation metrics of detection sensitivity, erasure rate (HR-detected lesions lost after reconstruction), and hallucination rate (predicted components absent from both the manual and HR segmentation). The dominant effect of SR was erasure of small real lesions, not hallucination, and it increased with slice thickness, though every reconstruction still improved lesion detection over the raw thick slice. ECLARE recovered small lesion signal best at both thicknesses, while the INR was no better than cubic interpolation.
[CV-1] UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression
链接: https://arxiv.org/abs/2608.06307
作者: Jacek Komorowski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit map retrieval. However, existing methods produce deterministic predictions, discarding aleatoric uncertainty that could improve robustness and downstream decision-making. We present UQ-Loc, which extends the LightLoc architecture with an anisotropic Gaussian covariance head that predicts a full 3x3 positive-definite covariance matrix per voxel. Training uses a Negative Log-Likelihood (NLL) loss augmented with a kNN-based spatial smoothness regulariser, while inference employs a modified SC2-PCR solver with uncertainty-weighted seed scoring and a Mahalanobis-distance inlier test. We adopt Expected Calibration Error (ECE) as a principled metric for evaluating the quality of the predicted uncertainty. Experiments demonstrate that UQ-Loc achieves consistent improvement in 6-DoF localization accuracy while producing well-calibrated covariances.
[CV-2] LNM: Externally Validated Tooth Detection Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN
链接: https://arxiv.org/abs/2608.06275
作者: Arash Nedaei,Henna Tiensuu,Elina Väyrynen,Saujanya Karki,Jaakko Suutala
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 16 pages, 7 figures, 6 tables
Abstract:Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research relies on clinical-grade radiographs or intraoral camera images, unavailable for public self-screening. This study introduces a tooth localisation and numbering model for smartphone photographs. We developed a customised Mask Region-based Convolutional Neural Network (Mask R-CNN) pipeline trained on 1,272 annotated smartphone images. To address variability in patient-generated health data, the pipeline incorporates two domain-informed mechanisms: a masked gray-world white-balancing algorithm to mitigate artificial colour casts and an anatomically constrained detection layer to enforce structural validity and suppress false positives. Evaluation comprised four stages: internal held-out testing, independent external testing, a descriptive ablation study, and fold-based training stability analysis using the same internal test set. On the internal test set, the model achieved an instance-mask AP@50 of 0.818, class-aware PQ of 0.780, and operational F1 of 0.884. Training stability showed limited between-model variation: across ten runs, instance-mask AP@50 had a standard deviation of 0.009. On the external dataset, the model achieved an instance-mask AP@50 of 0.901, class-aware PQ of 0.832, and operational F1 of 0.928 despite differences in population, sensors, and acquisition protocols. The inference pipeline is available as an open-source, containerised API. These results demonstrate that consumer-grade smartphone imagery can support automated tooth-level anatomical mapping, offering a scalable, potentially low-cost foundation for remote screening and tele-dentistry in resource-constrained environments.
[CV-3] OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations
链接: https://arxiv.org/abs/2608.06264
作者: Robin Trombetta,Carole Lartizien
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
备注:
Abstract:The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. Data augmentation is a technique widely regarded as a way to improve model training. It includes simple transformations like spatial operations or intensity modifications, but also more advanced synthesis techniques. Their goal is to generate new realistic samples from an existing dataset to diversify the images used during training. Among them, several propose different mixing strategies to combine real samples. However, one of their major shortcomings is to yield limited variability in terms of generated lesion shapes and locations. In this work, we introduce a novel image synthesis method, called OTLesMix, that leverages Wasserstein barycenter and optimal transport plan to generate realistic and diverse samples. We evaluated our method on three brain lesion segmentation tasks, on which it improves the Dice score compared to a model trained without synthetic data by 2.9 to 6.6 points, and outperforms state-of-the-art mix-based methods.
[CV-4] oward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
链接: https://arxiv.org/abs/2608.06252
作者: Saad Ahmed,Md Khalid Syfullaha
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets without expert verification and heavyweight pretrained backbones unsuited to on-device use. We introduce RSBdSL38, 10,874 expert-validated images spanning all 38 BdSL hand signs, representing the 51 letters of the Bangla alphabet, recorded from real signers at three special-needs schools across Bangladesh. We propose a lightweight attention based convolutional network of 298,470 parameters, built from grouped bottleneck residual blocks, channel and spatial attention, a multi-scale depthwise hand-feature block, dual pooling, and Swish activations. Trained from scratch, it attains 96.37% accuracy (95.72% ± 0.54% over five seeds), within 1.08 percentage points of the best of nine ImageNet-pretrained efficient architectures under an identical protocol, using 8.5 to 68x fewer parameters and 1.3 to 21.7x fewer MACs. Retrained, it reaches 92.95 to 98.33% on six public BdSL benchmarks, 97.04% on a merged corpus, and 76.25% zero-shot on BdSL-38. Removing any architectural stage costs 7.61 to 89.30 points, against at most 3.17 for the training recipe. Grad-CAM with deletion-insertion and weight-randomization checks confirms that predictions follow the signing hand. A signer-independent split holding out 6 of 36 signers yields 85.18%. Quantized to 0.48 MB, it runs at 3.98 ms per image within a 15.5 MB footprint on a commodity smartphone. Together, RSBdSL38 and our from-scratch model turn benchmark accuracy into deployable accessibility at a fraction of pretrained-backbone cost; dataset, code, and models are released.
[CV-5] PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
链接: https://arxiv.org/abs/2608.06240
作者: Elad Yoshai,Natan T. Shaked
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-based unpaired translators control preservation through a single global noise or guidance value applied across the image, which cannot separate content to keep from appearance to change. We present PRISM, a GAN-free flow-matching framework that replaces this global control with a learned per-feature gate. The gate’s spatial prior is derived from each source feature’s standardized distance to the target feature distribution, so features far from the target are freed while target-consistent features are preserved. The same gate controls both the initialization, which mixes the real source latent with a task-matched corruption, and the transport timing during Ordinary Differential Equation (ODE) integration. The corruption is matched to the task, content-anchored (AdaIN) for structure-preserving translation and partially anchored for structure-changing translation, and the gate can be overridden locally at inference from text or a detector without retraining, preserving important structures of the original image while still generating realistic results. We evaluate PRISM on five natural and biomedical benchmarks (AFHQ cat-dog, CelebA-HQ appearance translation, day-night relighting, virtual staining, and breast frozen-permanent histopathology). Among the evaluated methods under a shared same-split protocol, PRISM attains the best Inception FID and KID on four benchmarks and a competitive result on the fifth, and on histopathology yields the nuclei-count ratio closest to the ideal, supporting a favorable balance between target realism and structural preservation.
[CV-6] Depth-Guided Video Object Counting in Crowded Scenes
链接: https://arxiv.org/abs/2608.06236
作者: Yuanjing Xu,Xinyan Liu,Weidong Chen,Zixuan Zou,Linhao Zhang,Zhuangzhe Meng,Antoni B. Chan,Weigang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Accepted at ACM Multimedia 2026
Abstract:Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To address this, we propose a Depth-Guided Detector (DG-Det) along with a general post-processing pipeline. By integrating depth cues with multi-scale RGB-D cross-attention and explicit occlusion prediction, our method enhances spatial understanding and achieves robust detection in crowded and occluded scenes. Furthermore, we introduce a unified de-duplication framework to eliminate cross-frame redundant counting. To facilitate future research, we also release a new RGB-D Video Object Counting dataset featuring depth information and multiple object categories persequence. Extensive experiments demonstrate that our method achieves a 62.01% reduction in MAE compared to existing baselines, and also produces consistent improvements in RMSE. We provide the source code at this https URL and the dataset at this https URL.
[CV-7] EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
链接: https://arxiv.org/abs/2608.06231
作者: Bingyuan Wang,Baistan Zhyldyzbekov,Kunyu Feng,Zeyu Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.
[CV-8] Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era
链接: https://arxiv.org/abs/2608.06211
作者: Binze Wang,Jinyu Tian,Xingrun Wang,Xiaochen Yuan,Jianqing Li
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data copyright. Existing methods for creating unlearnable examples overlook the risk of data leakage, which can threaten data ownership. Thus, copyright protection in deep learning faces two main threats: illegal model training and malicious data leakage. We investigate that these two threats cannot be solved by straightforwardly combining existing availability attacks and watermarking techniques as their negative interaction effects. Therefore, in this paper, we propose a novel copyright protection mechanism for the aforementioned security concerns. Considering that the prevention of unauthorized model training requires powerful generalizability of unlearnable perturbations, we generate perturbations to induce the model to learn uncorrelated features of input images. It works by minimizing the mutual information of the input and output of the model. On the other hand, to eliminate the side impact of unlearnable perturbations on the watermark extraction, we design a dual extraction strategy by using two distinct watermark extractors. Extensive experiments on the image datasets ImageNet, CIFAR10, and Pets show that our proposed method could provide comprehensive copyright protection to images. The code is available at this https URL.
[CV-9] CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection
链接: https://arxiv.org/abs/2608.06205
作者: Nima Hatami,Karim Faez,Saeed Sharifian,Hamidreza Amindavar
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:RGB–T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low-light, adverse-weather, and complex multi-scale environments. However, existing methods still suffer from insufficient cross-modal interaction, unstable fusion from modality distribution gaps, and the high computational cost of heavy attention-based architectures. To address these issues, CFGPNet is proposed, a Cross-Attention-Based Fused Gradient Programmed Network framework for multispectral object detection. CFGPNet uses an improved GELAN backbone with RepViT-style re-parameterized blocks to strengthen feature representation while preserving computational efficiency. A Cross Computation Efficient Attention (CrossCEA) module is introduced to enhance cross-modal feature interaction and reduce redundant information transfer between visible and thermal branches. To generate compact and discriminative fused representations, an Attention Selection and Aggregation Fusion (ASAF) network combines dense feature aggregation with selective attention-based emphasis. Moreover, a programmable-gradient auxiliary branch is integrated into each CFGPNet variant to improve gradient delivery and optimization quality. Experiments on five public multispectral benchmarks, FLIR, M3FD, LLVIP, VEDAI, and MFAD, demonstrate that CFGPNet achieves strong and consistent performance across diverse scenes, object scales, and modality balances. In particular, the framework attains 80.7% mAP50 / 45.0% mAP50:95 on FLIR, 89.9% / 63.4% on M3FD, and 97.8% / 68.9% on LLVIP. It also reaches 83.3% / 56.9% on VEDAI and 83.4% / 61.8% on MFAD. These results show that CFGPNet is an effective, practical solution offering useful accuracy–efficiency trade-offs across three model scales. The code, data, and fine-tuned models are available at this https URL.
[CV-10] HOPE: Hand-Object Pressure Estimation from Monocular Videos
链接: https://arxiv.org/abs/2608.06192
作者: Subin Jeon,Byungjun Kim,Hanbyul Joo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: project page is at: this https URL
Abstract:Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects. We instead formulate pressure estimation as a hand-centric video prediction problem with monocular video as input. This formulation predicts temporally evolving per-vertex normal pressure and contact directly on the hand mesh, yielding a unified output space independent of object shape and sensor layout. Building on this formulation, we propose \textbfHOPE, a framework with two key components. First, we lift tactile-glove pressure, planar-sensor pressure, and distance-based hand-object contact annotations into a shared hand vertex space, allowing bare-hand contact data to regularize pressure learning where metric labels are unavailable. Second, we introduce a vertex-anchored video transformer that treats each vertex as a persistent token, aggregates visual features and hand pose over time, and uses a contact-gated pressure head to enforce that pressure vanishes without contact. Experiments on OpenTouch, PressureVisionDB, and hand-object contact benchmarks validate HOPE across object-pressure, surface-pressure, and contact-supervised HOI settings. Despite using metric pressure supervision primarily from gloved-hand videos, HOPE generalizes to bare-hand egocentric and in-the-wild videos, producing joint contact and pressure predictions beyond the scope of contact-only or planar-pressure baselines.
[CV-11] EvReflection: Event-Driven Micro-Dynamics for Reflection Removal ICML2026
链接: https://arxiv.org/abs/2608.06184
作者: Jiaxiao Wang,Dachun Kai,Huyue Zhu,Quanquan Hu,Zhenyang Xu,Xiaoyan Sun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ICML 2026
Abstract:Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from severe residual artifacts due to the inherent ambiguity between the reflection and transmission layers. In this paper, we propose leveraging event signals to break this ambiguity. By employing event cameras to capture micro-dynamics, we reveal the differential motion between these two layers. We thereby present a novel event-driven reflection removal network, EvReflection, that utilizes these dynamic cues for layer separation. Specifically, we design a Micro-Dynamics Decoupler to disentangle layer-specific motions from event streams as priors, which then guide a Parallax-Attention Rectifier to cleanly remove artifacts from the RGB image. Furthermore, to address data scarcity, we develop a parallax-aware simulation pipeline and construct the EVR ^2 benchmark dataset, the first real-world dataset for this task. Extensive experiments demonstrate that EvReflection achieves state-of-the-art performance on both synthetic and real-world benchmarks, surpassing the best competing method by more than 1.6 dB and 1.2 dB in PSNR, respectively. The code, dataset, and pre-trained models are available at this https URL.
[CV-12] Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
链接: https://arxiv.org/abs/2608.06174
作者: Zhongyao Wang,Wanli Ouyang,Taoyong Cui,Pheng Ann Heng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an injectively aligned leave-one-cell-out protocol over support x operation grids and SO-OPF, a readout that factors cell energy into support salience and a competitive operation posterior. This formulation separates two questions that aggregate scores conflate: whether the carrier composes held-out bindings when the grid is known, and whether that grid can be recovered from flat cell labels. With frozen DINOv3 features, known factorial assignment reaches 0.874 injective accuracy on Shapes3D-Extended and 0.799 on globally image-disjoint COCO; learning the assignment from flat labels reaches 0.769 and 0.762, respectively. Under matched-axis-aware supervision on Shapes3D, the factored carrier improves learned-assignment accuracy from 0.653 to 0.841 over a dense carrier and eliminates its laundering gap. SigLIP2 replicates the COCO separation. A rebuilt MuJoCo substrate exposes a boundary: learned-assignment accuracy is 0.569 with DINOv3 and 0.484 with SigLIP2, with substantial slot collapse. Thus factored readout and injective evaluation recover held-out bindings on two substrates while exposing, rather than hiding, a renderer-specific failure boundary; they do not establish universal recovery from flat labels.
[CV-13] Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments
链接: https://arxiv.org/abs/2608.06170
作者: Giorgio Tonetti,Laurent Kneip,Abel Gawel,Marco Hutter
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan or arbitrarily-structured environments. We propose Prior-SG, a task- and prior-driven framework that casts scene graph generation fundamentally as a probabilistic alignment problem. As the robot explores, it continuously aggregates an incoming RGB-D sensor stream into a physically grounded Instance Graph utilizing a multi-scale, open-vocabulary feature fusion strategy. The system then infers the high-level functional semantics of this map through a Maximum A Posteriori (MAP) estimate, guided by a Prior Graph-a logical expectation of the environment’s structure and task-relevant vocabulary synthesized dynamically by a Large Language Model. By optimizing a Markov Random Field that fuses heterogeneous experts (visual, geometric, and discrete objects) with these topological priors, the system resolves local perceptual ambiguities. We validate this approach across diverse simulated residential datasets and large, open-plan real-world environments. Prior-SG achieves state-of-the-art semantic region segmentation accuracy compared to recent baselines, robustly delineates distant functional boundaries in the absence of physical walls, and uniquely provides zero-shot ontological flexibility, enabling the robot to entirely restructure its spatial partitioning based on a given high-level task.
[CV-14] BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models
链接: https://arxiv.org/abs/2608.06164
作者: Yixiong Jing,Qi Wang,Lin Chen,Junwei Jiang,Guangming Wang,Haibing Wu,Olaf Wysocki,Wanli Ma,Brian Sheil
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning and interaction. Existing spring–mass based physical driven reconstruction approaches offer efficient and differentiable physical reconstruction, but they typically rely on axial springs alone. Such formulations oversimplify the underlying structural mechanics and can become mechanically under-constrained when the physical graph is coarsened, limiting their ability to preserve stable local deformation. We present BendTwin, a bending-aware differentiable spring–mass framework for video-based reconstruction and future prediction of deformable objects. BendTwin introduces bending stiffness and damping over local surface triplets, penalizing deviations from rest angles and regularizing higher-order deformation. These bending constraints improve mechanical stability while preserving the simplicity of spring–mass system. Experiments show that BendTwin consistently outperforms the axial-only PhysTwin baseline. Ablation studies further demonstrate that the bending constraints maintain system stability across different downsampling ratios and consistently improve upon the original PhysTwin formulation. Overall, BendTwin provides an effective approach for constructing mechanically faithful digital twins from sparse-view RGB-D videos.
[CV-15] Visual Grounding in Zero-Shot Vision-Language Control
链接: https://arxiv.org/abs/2608.06154
作者: J. de Curtò,Dayani Plasencia,Diego Sánchez,I. de Zarzà
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate this with an input-ablation battery: blind-image controls, repeated identical inputs, lane-axis reflection, non-visual baselines, and pipeline-integrity checks. Across nine direct-action models, six structured local VLMs, and an exploratory VLM-MPC hierarchy, we analyse 32,874 scored calls over two embodiments and three simulators. The direct-control results are largely negative: a constant-SLOW policy outperforms a scripted geometric controller, several models are image-invariant or nearly constant, and models that recognize longitudinal hazards still fail to transform LEFT and RIGHT under reflection. No local VLM meets the joint longitudinal and lateral grounding criteria. However, an image-only deterministic positive control estimates the lead gap with 0.090 m MAE and exact mirror equivariance, confirming the stimuli carry sufficient visual information; the failures are modular, not universal. A post-hoc, leakage-controlled symmetry-consensus guardian selects two models from 16 calibration frames and freezes a 2-of-4 hazard vote across original and reflected views. On 272 held-out frames it reaches 0.954 balanced accuracy (episode-cluster bootstrap 95% CI [0.895,0.990]); nested leave-one-episode-out recovers the same pair and threshold in all 12 folds. Abstaining on ties raises committed balanced accuracy to 0.973 at 0.824 coverage. With deterministic perception retaining lateral authority, offline modular replay achieves 0.934 action agreement and exact mirror equivariance. These results support current VLMs as bounded, selective hazard assistants, not monolithic zero-shot controllers.
[CV-16] CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
链接: https://arxiv.org/abs/2608.06150
作者: Zijie Wang,Chen Zhong,Wei He
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, 11 figures, including 3 supplementary figures. Code: this https URL
Abstract:Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category-agnostic change prior from frozen bi-temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category-dependent score shifts by dynamically estimating an image-query-specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building-damage assessment show that CogVis achieves state-of-the-art performance across all evaluated datasets. By sharing scene-level change perception, CogVis further avoids repeating category-agnostic temporal perception across queries and improves inference throughput by 28.50%.
[CV-17] Learning visual representations for compositional analysis of artworks and photographs ECCV
链接: https://arxiv.org/abs/2608.06142
作者: Fatemeh Behrad,Tinne Tuytelaars,Johan Wagemans
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ECCV workshops 2026
Abstract:Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization.
[CV-18] Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation
链接: https://arxiv.org/abs/2608.06126
作者: Manuel Laufer,Dominik Mairhöfer,Malte Sieren,Hauke Gerdes,Fabio Leal dos Reis,Arpad Bischof,Thomas Käster,Erhardt Barth,Jörg Barkhausen,Thomas Martinetz
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) this https URL
Abstract:An adequate diagnostic quality of radiographs is essential for reliable diagnoses and treatment planning. The patient’s pose during radiography is one of the most important factors determining the diagnostic quality. Since patient positioning is difficult and not standardized, an automated AI-based approach using depth images to automatically assess the patient’s pose before the radiograph has been taken would be helpful. Due to regulatory hurdles, however, it is difficult in practice to acquire the required depth images and corresponding radiographs. In this paper, we present a framework that can generate such training data synthetically from Computed Tomography scans. We further show that by pretraining on our generated synthetic dataset consisting of 3077 image pairs of upper ankle joints, the pose assessment of real upper ankle joints can be improved by up to 11 percentage points.
[CV-19] Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
链接: https://arxiv.org/abs/2608.06125
作者: Rui Li,Yuanzhi Liang,Ke Hao,Ziqiao Weng,Haibin Huang,Chi Zhang,XueLong Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar scores. They do not estimate the uncertainty of each prediction. The generator therefore cannot determine which feedback is reliable. This can drive optimization in the wrong direction and lead to reward hacking. We propose \textscSURE, a unified latent-space framework for image and video diffusion models. It learns reward distributions and directly uses their reliability to guide dense post-training. First, we propose sample-adaptive latent reward model (\textscSURE-LRM). It predicts a Gaussian utility for each noisy latent. Its mean predicts the reward score. Its variance reflect the uncertainty of prediction without human annotation. The learned distribution then guides post-training through uncertainty-guided reward feedback learning (\textscSURE-REFL). This method provides uncertainty-guided dense feedback along the denoising trajectory. At selected transitions, \textscSURE-REFL queries the frozen \textscSURE-LRM. It converts detached variance into reliability weights for samples at the same transition. Each weighted reward is backpropagated only through its local transition. The entire process remains in latent space and requires neither pixel-space decoding nor the full denoising graph. Experiments show that \textscSURE-LRM improves preference prediction over strong baselines. \textscSURE-REFL achieves the sota performance among various metrics and further improves optimization stability. It also achieves the highest VBench quality, semantic, and total scores among the evaluated methods.
[CV-20] Confidence matters: Leverag ing Multi-view Geometric Priors for GS-based Reconstruction
链接: https://arxiv.org/abs/2608.06117
作者: Hongyu Zhou,Zorah Lähner
类目: Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
备注:
Abstract:3D Gaussian splatting (3DGS) has emerged as a widely-used tool for novel view synthesis, offering real-time rendering in a sparse representation. However, the method’s reliance on structure-from-motion initialization and photometric optimization can lead to suboptimal geometric reconstruction, particularly for objects with high specularity. In this work, we investigate the integration of geometric priors, in the form of predicted normal and depth maps, into the 3DGS framework to improve the reconstruction quality. We analyze the effect of incorporating these priors into GS-based methods and our evaluation reveals that multi-view predictions, as they are done by the recent visual geometry grounded transformer (VGGT), outperform single-view alternatives. A major factor is the existence of a confidence map for the estimations, which comes as a by-product of multi-view models and which can significantly improve the effectiveness of priors by weighting each prediction appropriately. Extensive experiments on standard benchmarks show consistent improvement in reconstruction quality and significant gains in complex scenes including specular objects.
[CV-21] Dense-Cast: A lightweight ensemble of deep learning architectures for precipitation nowcasting
链接: https://arxiv.org/abs/2608.06082
作者: Gourav Jyoti Kalita,Hidam Kumarjit Singh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Proper short-term forecasting of precipitation is crucial in disaster management and preparedness. Nonetheless, the variability and nonlinearity of precipitation make short-term forecasting challenging for meteorologists. Moreover, capturing temporal dependencies in spatiotemporal data is a challenge in precipitation nowcasting. In this article, we introduce a lightweight deep learning model for half-hourly precipitation nowcasting. This model has been designed by incorporating the DenseNet architecture, residual connections, and transformer encoders for effective precipitation nowcasting with reduced model parameters. The North-Eastern region of India has been selected as the area of interest for our study. The region receives the highest precipitation during the months of June-September due to the monsoon season. The proposed model takes the previous five time-steps of half-hourly precipitation as inputs and predicts the precipitation in the next two half-hours. The GPM IMERG precipitation dataset with a 30-minute cadence has been used in this study for training and testing the model. The proposed architecture achieves best MAE of 0.235 millimetres, RMSE of 0.735 millimetres, and KGE score of 0.816 at an interval of 30 minutes.
[CV-22] Domain-Grounded Candidate Selection for Agent ic Image Editing: A Shadow Removal Case
链接: https://arxiv.org/abs/2608.06075
作者: Shilin Hu,Jingyi Xu,Dimitris Samaras,Hieu Le
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the need for classic, physics-informed low-level vision? We study this through shadow removal, a problem shaped by scene geometry, illumination, materials, and occluders, where paired shadow and shadow-free data are hard to collect at scale. We find that a commercial generative editor, used directly, can produce clean shadow-free edits that preserve surface texture and local appearance. However, this comes with a new failure mode: the same editor can regenerate scene content, hallucinate objects, or misread a shadow as material or geometry, producing plausible but physically wrong edits. We address this with an agentic candidate-selection pipeline: the editor generates a guided probe, an evaluator screens for major failures, retries when needed, samples multiple candidates, filters them, and selects a final result balancing shadow removal against scene preservation. Grounding this process in shadow-formation physics makes it more reliable: prompting the generator and evaluator to treat shadows as illumination effects caused by light occlusion, not material or object structure, measurably improves quality and consistency. On the ShadowRemovalRefine benchmark, our physics-oriented pipeline achieves a CDD of 0.0075, reducing CDD by at least 47% over the strongest prior method. These results suggest that commercial vision-language models do not replace classic low-level vision priors; instead, such priors remain useful for constraining and steering physically underconstrained generation.
[CV-23] he Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents
链接: https://arxiv.org/abs/2608.06065
作者: Weiwei Li,Junzhuo Liu,Tong Chu,Hengfu Yu,Wen Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an action from the current screen and interaction history, while the subsequent observation is discarded. This removes the rationale of why an action is correct: the evidence often appears only on the subsequent screen. For example, to enable Soft Wrap, the agent should click Edit or View, but nothing reveals this until the menu opens. Without such evidence, standard imitation gives the model little chance of ever sampling and thus learning the correct reasoning. To address this issue, we propose Gated Hindsight Distillation (GHD), which uses the next screenshot as privileged information during training. A student predicts from the observable trajectory prefix, while a parameter-sharing teacher additionally observes the next screenshot and re-scores the student’s on-policy responses. We apply distillation only when the student fails and the hindsight-conditioned teacher recovers the demonstrated action. GHD improves task success over GRPO on AndroidWorld and AndroidLab across two vision-language models. The code and checkpoints will be made available.
[CV-24] Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture ICDAR2026
链接: https://arxiv.org/abs/2608.06062
作者: Poonam Poonam,Alexander Epple,Timo Ropinski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ICDAR 2026
Abstract:Bar charts are commonly used in data visualization, and while they are easily understood by humans, it is non-trivial to extract the underlying data computationally. For a machine-learning-based approach, training chart de-rendering models usually requires labeled, real-world data. Labeling data is a time consuming task, which is why annotated data is scarce. Models can learn more efficiently when provided with features of high semantic quality, which a joint-embedding predictive architecture (JEPA) is designed to learn in a self-supervised manner. We present a per-bar, numerical value recovery pipeline for bar charts, where a JEPA encoder is used to produce semantically rich latent features. The decoder model consuming these features is simple and quick to train and outputs the coordinates of ticks and bars, which can be used to recover bar values. The effectiveness of self-supervised finetuning and quality of the extracted features is evident when comparing our model to end-to-end supervised baselines. Code, datasets and checkpoints are available on \hrefthis https URLGitHub.
[CV-25] Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
链接: https://arxiv.org/abs/2608.06060
作者: Zelong Sun,Jun Wang,Kaicheng Yang,Tiancheng Gu,Ziyong Feng,Zhiwu Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 26 pages,10 figures,14 Tables
Abstract:Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.
[CV-26] DARAD: Dual Adapters and Ranking-Aware Distillation for Continual Remote Sensing Image-Text Retrieval
链接: https://arxiv.org/abs/2608.06059
作者: Xi Chen,Xu Chen,Xiangyang Jia,Wei Wang,Xu Zhang,Zhenyuan Sun
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:With the rapid growth of Earth observation technologies, remote sensing archives are rapidly expanding, making remote sensing image-text retrieval (RS-ITR) increasingly important. However, continual RS-ITR remains challenging because scale variation and distribution shifts in RS aggravate cross-modal alignment space distortion, making it difficult for existing continual learning (CL) methods to support reliable continual retrieval. To address this challenge, we propose DARAD, a dual-adapter and ranking-aware distillation framework that preserves the historical cross-modal ranking structure while learning new visual and textual concepts from evolving archives. Specifically, the visual branch introduces a spatial fusion adapter, which integrates coarse regional cues and fine-grained patch cues to accommodate RS scale variation while anchoring visual updates to the pretrained alignment space. The textual branch employs multi-expert semantic routing, which separates shared textual semantics from semantically specialized residuals to absorb newly emerging descriptions while constraining global text embedding drift. Furthermore, bidirectional ranking distillation uses a frozen teacher model and historical anchors to preserve the historical cross-modal ranking structure, thereby mitigating alignment space distortion across continual stages. Experiments under a multi-stage continual retrieval protocol show that DARAD achieves superior performance over existing CL methods, improving adaptation to newly arrived data while maintaining effectiveness on historical data.
[CV-27] Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
链接: https://arxiv.org/abs/2608.06037
作者: Rafał Buler(1),Jakub Buler(1),Maciej Bobowicz(2),Michał Grochowski(1) ((1) Gdańsk University of Technology, (2) Medical University of Gdańsk)
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
备注: Accepted as a short paper for presentation at the 21st International Conference on Computational Intelligence Methods for Bioinformatics and Biostatistics (CIBB 2026)
Abstract:Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. We begin by establishing a baseline using an EfficientNetB3 architecture. To move beyond standard convolutional biases, we adopt a patch-based strategy, employing a convolutional masked autoencoder to learn implicit inter-patch relationships through self-supervised reconstruction. We then extend this approach by incorporating explicit relational modelling, organizing the learned embeddings into various graph topologies, including grid-based, random, and k-nearest neighbour structures. Experimental results on the ISIC-2018 and ISIC-2019 skin lesion diagnosis benchmarks show that combining implicit inter-patch modelling with explicit graph-based message passing yields the best performance. On the ISIC-2018 test set, the baseline model achieves a balanced accuracy of 76.17%, which improves to 77.12% with implicit patch-based relational modelling. The fully integrated grid-structured Graph Attention Network further increases performance to 79.27%. Similarly, on ISIC-2019, the implicit approach reaches 59.84% balanced accuracy, while the combination of implicit and explicit modelling yields 60.67%.
[CV-28] PaCoNet: Deep Data Extraction for Parallel Coordinates ICPR2026
链接: https://arxiv.org/abs/2608.06030
作者: Poonam Poonam,Hannah Kniesel,Pere-Pau Vázquez,Timo Ropinski
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ICPR 2026
Abstract:Extracting data from visualizations has long challenged computer vision, with current research focused on bar, line, and pie charts, among other low-dimensional visualizations. However, parallel coordinates as a widely used high-dimensional data visualization approach, remain largely unexplored in this context. As parallel coordinate plots can quickly become cluttered and difficult to interpret when poorly designed or densely populated, automated data extraction from such visualizations is of particular interest. In this paper, we propose PaCoNet, the first approach for parallel coordinate data extraction. PaCoNet not only extracts line coordinates, but also enables the extraction of individual data samples for further analysis. Towards this end, we make the following contributions. We present the first deep learning approach tailored for parallel coordinate analysis, and demonstrate that it outperforms unadapted baselines by a significant margin. We further introduce a large-scale parallel coordinate dataset for training and testing. Together, these key contributions enable for the first time the automated analysis and redesign of parallel coordinate plots. PaCoNet thus lays the groundwork for complex visualization analysis, and further advances the intersection of computer vision and data visualization. All code, trained models, and data generation scripts will be made publicly available upon acceptance of the paper.
[CV-29] Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry Training Coverag e and Compute
链接: https://arxiv.org/abs/2608.06014
作者: Shijie Hao,Weining Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 24 pages, 11 figures, 3 tables
Abstract:Should a completion model spend extra test-time compute by iterating, or spend a similar parameter budget on a wider one-shot predictor? The answer is easily confounded by denoising curricula, corruption augmentation, capacity, and unpaired evaluation. We study this question in LiDAR semantic scene completion by comparing a one-shot predictor, a parameter-matched wider predictor, and a weight-tied multigrid refiner initialized from the same frozen predictor. The protocol separates coherent region removal, independent thinning, range-dependent attenuation, and additive clutter while preserving exact scene-condition pairing. Across five training seeds and 815 SemanticKITTI sequence-08 frames, the full iterative system improves mIoU over the wide control by 0.911 points under contiguous angular removal, with a 95% moving-block bootstrap interval of [0.804, 1.040] that clears a predeclared 0.5-point practical margin. Under independent 75% thinning, iteration adds only 0.300 points [0.166, 0.436], whereas observation-family augmentation adds 5.975 points [5.662, 6.140]. Neither intervention repairs additive clutter. The iterative system also costs 10.74 ms and 0.75 GiB per frame, versus 6.25 ms and 0.23 GiB for the wide control. These results establish a geometry-conditioned empirical boundary rather than a universal advantage: coherent gaps can justify fixed-depth refinement, broadly thinned evidence is addressed more effectively by training coverage, and spurious evidence requires a different robustness mechanism.
[CV-30] Wan-Animate-2: Pushing the Application Boundaries of Character Animation
链接: https://arxiv.org/abs/2608.06009
作者: Guangyuan Wang,Li Hu,Dechao Meng,Zhongyi Zhang,Peng Zhang,Mingyang Huang,Ruoshi Zhang,Ke Sun,Zhe Zhang,Xingjun Wang,Gang Cheng,Bang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three paradigms: methods based on explicit motion representations suffer from extraction errors and identity drift; methods based on implicit motion features lose fine-grained dynamics through compression; and in-context learning approaches avoid intermediate representations but incur prohibitive computational costs. Furthermore, all current systems are designed for offline synthesis, unable to meet the real-time requirements of interactive applications such as digital avatars and live-streaming hosts. To address these limitations, we present Wan-Animate-2, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer. Our architecture achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely. We further introduce text driven viewpoint control that decouples the output camera perspective from the driving video–a capability rarely supported by prior character animation methods that rely on explicit motion representations. Beyond generation quality, we present Wan-Animate-2-Lite, an efficient variant that reduces inference latency to real-time thresholds through a three-stage training paradigm: teacher forcing pretraining with error buffer mechanism, and Self-Forcing distillation with chunk-wise backpropagation. This enables streaming character animation for interactive applications, opening new deployment scenarios that were previously infeasible. Qualitative evaluations and user studies demonstrate that Wan-Animate-2 achieves high-fidelity animation results across diverse characters and motion patterns. To foster further research and community development, we will release the Wan-Animate-2-Base model weights to the public.
[CV-31] Universal Concept Disruption for SAM3 Image Segmentation
链接: https://arxiv.org/abs/2608.05983
作者: Hao Wang,Yuxuan Zhang,Wei Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:SAM3 extends promptable segmentation from geometry-driven mask prediction to open-vocabulary concept segmentation, where a text-conditioned grounding model decides whether a concept is present and segments all matching instances. While this presence-gated design improves concept-level prediction, its adversarial robustness remains unexplored. In this paper, we introduce Universal Concept Disruption (UCD), the first universal cross-concept adversarial attack tailored to SAM3 image segmentation. UCD learns a single bounded image perturbation from (image, noun-phrase) pairs and attacks SAM3 as an integrated concept-grounding system. It jointly disrupts the text-conditioned input path, maximizes divergence in prompt-shared visual features, suppresses the final presence-gated concept scores, and corrupts the spatial validity of retained masks through area collapse and clean-mask Dice disruption. Across SACo-Gold, LVIS, RefCOCO, PhraseCut, and OpenImages datasets, UCD consistently outperforms all baselines under a matched evaluation protocol, reducing average mask AP from 59.43 to 18.73 and average cgF1 from 50.32 to 20.49. The learned perturbation also transfers to SAM3.1 and to SAM3 video inference without re-optimization, while prompt ensembling, lightweight head fine-tuning, and temporal filtering provide limited recovery.
[CV-32] Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality
链接: https://arxiv.org/abs/2608.05979
作者: Syed Roshaan Ali Shah,Kasper Bonte,David Bekaert,Kristof Van Tricht,Dieter Wens
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 17 pages, 7 figures
Abstract:Machine learning, and deep networks in particular, are increasingly used to derive higher-level Earth observation (EO) products such as annual land-cover and crop-type maps. Many are generated operationally: each year a new acquisition is processed, typically with the same model, extending a multi-year archive. In the process these systems accumulate two kinds of useful signal that are almost never fed back into the model: the system’s own archive of past predictions, and ancillary layers produced by other partners in a processing consortium. Both are normally used outside the network, as rule-based post-processing or a fixed input mask. Using the Copernicus Land Monitoring Service High Resolution Layer (HRL) Croplands crop-type product as a testbed, we show that bringing both signals inside the model turns a single-year, single-task pixel classifier into one that reasons across years. We introduce a Crop Type (CTY) embedding encoder that represents each past prediction as a confidence-scaled, time-ordered categorical token and attends over the year axis, and we study how the externally provided Base Vegetation Layer (BVL) mask should be represented in the model’s inputs and outputs. To compare designs fairly when they relabel non-crop pixels, we evaluate on the 18 crop classes only and report precision and recall separately. On a pan-European dataset of about 5.4M labelled pixels, adding the prediction history raises crop-only F1 by 1.6 percentage points (pp) and, more importantly, corrects a recall-skewed error profile, with the largest gains on perennial and tree crops (olives +4.6, fruits +3.7, nuts +3.2 pp). Representing the BVL mask consistently in both the history and the target year adds about 2.5 pp on the crop classes. The approach is a low-cost recipe for any recurring geospatial or foundation model that emits class maps.
[CV-33] Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
链接: https://arxiv.org/abs/2608.05976
作者: Haoning Yang,Xinyuan Chen,Yaohui Wang,Guo Lu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted for publication in ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)
Abstract:Recently, diffusion models have made great progress in video generation. However, most existing video diffusion models are trained with short videos, and degrade when extrapolated to long videos, struggling to maintain long-range temporal coherence while retaining diverse motions. To generate consistent, high-quality and dynamic long videos, we propose Diff-VF, a training-free, plug-and-play and model-agnostic framework that converts existing short-video diffusion backbones into long-video generators without modifying or fine-tuning the base model. Diff-VF couples three complementary strategies: Hybrid Noise Initialization (HNI) to constrain global semantics, Weighted Window Sampling (WWS) to remove inter-window discontinuities, and Temporal Extended Sampling (TES) to establish long-range dependencies with a timestep-varying fusion. We further extend Diff-VF to long-video enhancement via Skip Residual Guidance that balances fidelity and realism through timestep-dependent guidance. VBench-Long evaluation results show that Diff-VF achieves a more favorable balance between temporal coherence and motion diversity than base models and recent training-free long video generation baselines, including FreeNoise, FreeLong, and RIFLEx, while maintaining competitive frame-wise quality. Experiments on two base models demonstrate the applicability to video diffusion models with different spatial-temporal modeling strategies. Extensive ablations validate the contribution of each component and hyperparameters.
[CV-34] opology-Aware Neighborhood Learning for Source-Free Cross-Scene Hyperspectral Image Classification
链接: https://arxiv.org/abs/2608.05964
作者: Qingmei Li,Juepeng Zheng,Jiarui Zhang,Jianxi Huang,Haohuan Fu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Domain adaptation has advanced cross-scene hyperspectral image classification, significantly improving discriminative capability in complex scenarios. However, privacy rules or storage limits often block access to data from the source domain. Conventional domain adaptation methods become impractical, severely restricting their utility in realistic remote sensing scenarios. To tackle this challenge, we propose a topology-aware source-free learning framework. We first introduce the entropy momentum pseudo-labeling (EMP) to refine k-means assignments by leveraging entropy-aware confidence and temporal prediction momentum. Under the guidance of the refined pseudo-labels, we further utilize the contextual neighborhood topology (CNT) to exploit the intrinsic geometric structure of the target feature space. Combining the global structural information extracted by collaborative representation with the local similarity information modeled by nearest neighbor search, the CNT accomplishes the comprehensive encoding of manifold-level geometric properties in the target domain feature space. The overall objective integrates cross-entropy on refined pseudo-labels, log inner product-based topology consistency, and an information-maximization term for balanced classification, ensuring stable adaptation in the source-free setting. Extensive experiments on three typical cross-scenarios demonstrate that the proposed method exceeds state-of-the-art performance, and ablation studies further validate the contribution of each module. The results highlight the critical role of topology-aware modeling in achieving robust and accurate classification without source data.
[CV-35] Big Bright or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
链接: https://arxiv.org/abs/2608.05960
作者: Maulik Chevli,Johannes Brandt,Rickmer Braren,Daniel Rueckert,Philip Müller
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models could assist this process by providing generalizable representations of anatomy and pathology. To evaluate their diagnostic breadth, we benchmark ten frozen CT encoders across three cohorts of thoracic CT scans, including an unseen internal clinical dataset, using k -nearest neighbors, zero-shot prompting, and linear probing. We find no universal state-of-the-art, with rankings fluctuating significantly depending on the evaluation context. While models combining fine-grained image tokenization with vision-language alignment generally perform best, a lightweight supervised encoder remains highly competitive, demonstrating that explicit labels can effectively substitute for scale. Crucially, rather than model architecture, we observe that the primary determinant of performance is a physical bottleneck: a finding’s detectability scales with its contrast against surrounding tissue and its spatial extent. Through controlled within-organ comparisons, we empirically demonstrate that widespread or high-contrast abnormalities, such as devices and effusions, are reliably recovered. Conversely, small, low-contrast focal lesions remain a persistent challenge across all evaluated encoders. We attribute this to the inherent limitations of globally pooled embeddings, suggesting that accurately representing small, low-contrast structures will require region- or lesion-level pretraining.
[CV-36] raining a Conditioned Video Game Agent on a VLM Annotated Dataset
链接: https://arxiv.org/abs/2608.05954
作者: Katrin Schmid,Iuri Frosio
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, access to the game engine is required to get rewards for training (e.g. to collect rewards from the environment). Furthermore, the proper identification and weighting of the rewards generally requires a difficult trial-and-error approach. Lastly, rewards are often sparse and understanding how they eventually affect the learned policy is a non-trivial exercise. To ease these issues we propose annotating a video game dataset with Vision Language Models (VLMs) instructed to extract human defined rewards. We show that offline RL can then be used to train a conditioned agent that responds accordingly to the desired returns and we discuss the difficulties and limitations that emerged in our early experiments.
[CV-37] VLMs for Videogame Data Annotation
链接: https://arxiv.org/abs/2608.05949
作者: Katrin Schmid,Iuri Frosio
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-world applications. Their adoption in video games is on the other hand limited by the extreme variability of the synthetic scenarios and their poor compliance with real-world physics. Here we investigate the use of VLMs for annotating video game frame sequences with reward signals, a task with several potential applications including, among others, conditioned training and offline reinforcement learning. We show that VLMs often struggle to answer basic questions on racing video games (although we observed a similar behavior on other game genres) and discuss countermeasures such as VLM output mixing and prompt optimization. We also show how input sequence length, resolution, and question batching affect the annotation quality and its token consumption.
[CV-38] GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
链接: https://arxiv.org/abs/2608.05948
作者: Shuai Wang,Yaxin Feng,Xuekun Jiang,Shihan Tian,Ningyu Yan,Xing Shen,Chaoyang Lyu,Hui Wang,Yunsong Zhou,Hanqing Wang,Jiangmiao Pang,Yang Xiang,Xing Gao,Chunhua Shen,Weinan Zhang
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on perceptual similarity or human judgments, providing limited insight into which physical principles or parameters are violated. We introduce GAUGE, a real-world-grounded diagnostic benchmark for jointly evaluating how numerical simulators and generative video world models reproduce or deviate from real-world physics. It comprises 22 controlled task families covering rigid bodies, flexible cables, textiles, and volumetric deformable objects. Grounded in real-world trajectories and paired with calibrated physical metadata, uncertainty annotations, and task-specific observables, these tasks cover fundamental physical processes including collision, friction, momentum transfer, oscillation, self-contact, and deformation across diverse materials and conditions. We benchmark Isaac Sim, Genesis, and Newton on 14 task families using generalized trajectory errors, and evaluate 6 image-to-video models on 5 rigid-body tasks by testing physical-law consistency and the temporal stability of inferred parameters. Our results reveal no uniformly faithful physics engine, with the largest discrepancies arising in impulsive contact, rapid textile motion, and volumetric deformation. We further find that video world models can produce trajectories with the expected equation form while recovering incorrect accelerations, momentum transfer, and oscillation timing. GAUGE lays the groundwork for developing more physically faithful simulators and world models for embodied intelligence.
[CV-39] Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models
链接: https://arxiv.org/abs/2608.05945
作者: Jingyan Jiang,Yaru Sun,Xiao Chen,Jiazhen Huang,Caiting Li,Zhijian He,Yin Chen,Pingting Hao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, making predictive confidence unreliable for downstream decision-making. Many existing label-free calibration approaches are either coupled to prompt optimization or rely on logit-range statistics that provide only a coarse characterization of the predictive distribution. We show that TTA can increase confidence and reduce entropy even when the top-1 prediction and its correctness remain unchanged, a failure mode we term prediction-preserving sharpening. Across diverse TTA methods and benchmarks, larger entropy reductions relative to paired zero-shot predictions are associated with greater increases in Expected Calibration Error (ECE). On entropy-reduced samples, confidence gains also tend to exceed accuracy gains. Based on these findings, we propose Zero-Shot-Anchored Entropy Calibration (ZAEC), a label-free post-hoc method that uses zero-shot entropy as a sample-specific uncertainty reference. ZAEC selectively restores the zero-shot entropy of sharpened predictions through minimal temperature scaling while leaving all other predictions unchanged. It requires no labeled calibration data or learned parameters and preserves class rankings and classification accuracy. Across five TTA methods and 15 datasets, ZAEC achieves the lowest post-hoc macro-average ECE on ViT-B/16, with consistent gains on RN50.
[CV-40] MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?
链接: https://arxiv.org/abs/2608.05938
作者: Attila Simkó
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Medical images are routinely de-identified—names, dates, and other metadata removed—and then shared for research, teaching, and public benchmarks under the assumption that this renders them anonymous. Such de-identification protects the metadata but not the pixels, and—apart from scans that directly contain facial structures—whether the image content itself identifies the patient has received little scrutiny. We investigate this question by learning a cycle-consistent correspondence between a cross-sectional medical image and a non-medical, patient-identifying image, using a pair of coupled, cycle-consistent variational autoencoders. From a held-out scan, the model recovers a recognisable likeness of the patient (identity-region MAE = 0.163); conversely, it synthesises a scan from such an image. These results indicate that a de-identified medical scan remains identifying—it is, in effect, a photograph of the patient—and that imaging data should be governed as biometric data rather than as anonymisable records. To support reproducibility, the code and trained models are shared at this https URL.
[CV-41] Floating Radiance Networks
链接: https://arxiv.org/abs/2608.05920
作者: Krzysztof Byrski,Rafał Tobiasz,Grzegorz Wilczyński,Mikołaj Zieliński,Dawid Baran,Dominik Belter,Jacek Tabor,Przemysław Spurek
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent advances in neural scene representations enable photorealistic novel-view synthesis, yet most methods remain tightly coupled to a single rendering paradigm, limiting their versatility and integration with conventional graphics workflows. We introduce Floating Radiance Networks (FlaRe), a neural scene representation combining explicit ray-traceable geometry with continuous neural radiance functions. A scene is represented by floating planar generalized Gaussian primitives, each carrying a compact latent descriptor of a local radiance field. A lightweight decoder shared across the scene maps this descriptor, local surface coordinates, and viewing direction to color and opacity. This formulation preserves the expressiveness of neural fields while providing an explicitly addressable structure that can be efficiently queried and manipulated. Hardware-accelerated primitive intersections enable interactive rendering and recursive ray-tracing, including reflections, refractions, transparency, and shadows. The same representation further supports primitive-level deformation, mesh extraction, and appearance stylization directly in its learned descriptor space. Experiments across standard reconstruction benchmarks demonstrate competitive rendering quality while using a compact set of primitives. Together, these results establish FlaRe as a versatile representation that brings high-fidelity neural rendering, ray-tracing, geometric manipulation, and appearance editing into a unified scene model. Source code is available online. Source code can be found at: this https URL
[CV-42] Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press
链接: https://arxiv.org/abs/2608.05911
作者: Chahan Vidal-Gorène(CJM, LIPN),Seda Kirakosyan(UFAR),Edita Matevosyan(UFAR)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Parisian Armenian commercial community. On each page, commercial advertisements are located, read, and parsed into structured records, which are then geocoded and placed on the map. Western Armenian is under-resourced and unsupported by off-the-shelf layout and OCR models, so the pipeline uses vision-language models (VLMs) as a data-bootstrapping strategy: they produce usable structured records at a scale hand annotation could not reach, and stay reliable on the strongly curved scans where conventional line-level CRNN OCR breaks down. The contribution includes a 500-page Western Armenian press corpus with 3,270 advertisement-level annotations, a Label Studio template that captures detection and semantic fields in a single annotation pass, and a reproducible workflow transposable to other under-resourced historical corpora. More broadly, the work shows that VLM-driven data bootstrapping is an effective lever for under-resourced historical languages such as (Western) Armenian.
[CV-43] Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
链接: https://arxiv.org/abs/2608.05903
作者: Haodong Yan,Junfeng Li,Junjie He,Zhide Zhong,MingMing Yu,Wenxuan Song,Jiaguan Zhu,Yangyang Zheng,Yuqiao Du,Jiadi You,Yingjie Cai,Xu Yan,Guanyi Zhao,Bingbing Liu,Haoang Li
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注:
Abstract:Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.
[CV-44] o See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
链接: https://arxiv.org/abs/2608.05879
作者: Xiaobin Huang,Zilong Huang,Yang Luo,Hongchao Fan,Yiping Chen,Ting Han
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 4 figures
Abstract:Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually synthesized independently, lacking the correspondence required for a coherent urban world. We present HoloWorld, a unified indoor-outdoor urban world generation framework built on a continuously updated cross-scale world context. Initializing from a user description, HoloWorld progressively represents and updates the diverse world information, from city-scale planning to individual buildings, allowing generated interiors to maintain explicit correspondence with their associated exterior buildings. Conditioned on the evolving context and previously generated neighboring blocks, HoloWorld autoregressively generates urban exteriors with consistent spatial organization and visual identity across blocks. The generated exterior representations are further grounded in 3D building instances and footprints, enabling building-specific indoor generation with geometry-constrained layouts and inherited appearance characteristics. To our knowledge, HoloWorld is the first framework to unify indoor and outdoor generation within a coherent 3D urban world. Extensive experiments demonstrate that HoloWorld achieves superior urban exterior generation performance, improving the average AQS score over the SOTA by 7.68% and obtaining the highest average RDR score, while maintaining strong building-level indoor-outdoor correspondence and cross-block continuity within a unified 3D urban world.
[CV-45] MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
链接: https://arxiv.org/abs/2608.05878
作者: Rajatsubhra Chakraborty,Xujun Che,Ritabrata Chakraborty,Xi Niu,Depeng Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 20 pages, 14 figures, 9 tables. Preprint under review
Abstract:Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot open-vocabulary semantic segmentation. State-of-the-art attribution methods score each pixel independently, comparing its features against a fixed text-derived class representation, whether as an output-space similarity or as a cross-attention weight. This discards structured signals the model itself exposes: the temporal structure of the generative trajectory, the visual appearance statistics of each concept, and the image’s own pairwise feature geometry. We present MAVISEG, a training-free refinement layer that recovers these signals. Because its operators consume only a pixel-by-concept score field and a pixel feature space, MAVISEG is capture-agnostic rather than tied to one attribution method. Across six benchmarks it achieves the strongest overall results among training-free methods, including the best mIoU on every benchmark. Interestingly, gains are largest where the initial capture is weakest, and individual operators contribute depending on the noise in the field they refine. Our results indicate that diffusion transformers carry more concept-level information than current attribution methods recover, and that much of it is lost on the way to the mask rather than absent from the model.
[CV-46] D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation
链接: https://arxiv.org/abs/2608.05877
作者: Elena Bueno-Benito,Mariella Dimiccoli
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation. Yet, in existing OT-based methods, the latent action prototypes that define the OT costs are not re-estimated from the refined frame geometry. Instead, they evolve solely through gradients from the pseudo-label loss. We identify this \emphrepresentation–prototype inconsistency as a central bottleneck, particularly around ambiguous transitions and for short or infrequent actions. To address this issue, we build on the recently introduced CLOT, which refines frame embeddings based on estimated segment embeddings, and further re-estimates the action prototypes from the refined frame embeddings. Specifically, we introduce a graph-constrained module that regularizes the OT-refined frame and segment representations by preserving the local neighborhood geometry of the encoder output. An action-embedding refinement step then periodically re-anchors the prototypes to this stabilized representation geometry. We study two instantiations that share the same backbone, graph module, and objective: D-CLOT updates the prototypes using k -means, whereas D-CLOT _B updates them as OT barycenters weighted by the refined transport plan, yielding an assignment-aware prototype update consistent with the current transport geometry. Across five established benchmarks, both variants improve segment-level quality over CLOT, with per-video gains of up to +12.7 F1 and +10.2 mIoU (YTI) and activity-level gains of up to +8.9 F1 (FS-Eval). We further establish the first unsupervised action-segmentation baseline on Assembly101, a procedural and substantially more fine-grained benchmark than those commonly used in prior work. Extensive ablations and sensitivity analyses demonstrate that the two refinement mechanisms are complementary and robust.
[CV-47] Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion
链接: https://arxiv.org/abs/2608.05858
作者: Badha Rathna Sabhapathy,Gotam Dahiya,Vishesh Vatsal
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 8 pages
Abstract:Accurate object detection in aerial and satellite imagery is dependent upon the bounding box representation. This is especially true for spatially oriented objects such as ships or aircrafts. Oriented Bounding Boxes (OBB) have a tighter fit and more robust non-max suppression compared to Horizontal Bounding Boxes (HBB), any current post-processing conversion from OBB to HBB either introduces excess empty and background space or removes data from the detection. This paper introduces a novel approach for a shape-aware OBB-to-HBB conversion for ship detection in remote sensing imagery. It leverages hull shape, hull fullness, and the bounding box orientation to produce a tighter axis-aligned HBB representation. The proposed method is benchmarked against three baselines methods for OBBto-HBB conversion, Outer HBB which uses minimum and maximum, Area Equivalent HBB and GBB Marginalized HBB.
[CV-48] DTRNet: Dual Text-Radical Decoding for Handwritten Chinese Text Recognition with Faked Character Detection ACM-MM2026
链接: https://arxiv.org/abs/2608.05848
作者: Runrui Li,Lin Zhu,Hua Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ACM MM 2026 (Oral)
Abstract:In K-12 educational scenarios, handwritten Chinese text recognition should not only transcribe student writing, but also detect faked characters. However, existing recognition models are usually confined to a predefined set of normal characters and therefore cannot explicitly identify faked characters. Existing detection methods exhibit complementary limitations: character-level methods provide interpretable structural evidence but suffer from low efficiency, whereas line-level methods are efficient but rely heavily on confidence scores, making them prone to missed detections and lacking explicit structural evidence. Thus, the key challenge is to preserve character-structural evidence independent of contextual inference while maintaining line-level efficiency. To this end, we propose DTRNet, a dual Text-Radical decoding framework for line-level faked character detection. DTRNet decouples context-aware text recognition from character-wise structural verification, where the text branch performs line-level transcription and the radical branch predicts legal Ideographic Description Sequences (IDS) for lexicon-based faked character judgment. We further introduce IDS-Guided Confidence Adjustment (IGCA) to refine text predictions using structural evidence during inference. Experimental results demonstrate that DTRNet effectively detects faked characters while maintaining strong recognition performance and providing interpretable radical-level evidence. Code, checkpoints, and the processed dataset are publicly available at this https URL.
[CV-49] Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation ECCV2026
链接: https://arxiv.org/abs/2608.05844
作者: Théo Danielou,Antoine Saporta,Léo Alberge,Corentin Dancette
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at ECCV 2026 Workshop AI4M3D
Abstract:Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder. Dense prediction tasks such as 3D segmentation are, however, underrepresented in their evaluation, and, with the encoder kept frozen, pre-trained models still fall short of nnU-Net, the state-of-the-art reference trained from scratch. To close this gap we extend convolutional MAE pre-training with a robust reconstruction objective, a feature regularizer, and a local-global similarity objective. Using this method, we propose Curia-MAE, a multi-modal, multi-anatomy MAE model pre-trained on 300,000 CT and MRI images covering a large number of anatomical sites. On eight anatomy- and lesion-focused segmentation benchmarks, Curia-MAE improves frozen-encoder performance over a strong MAE baseline, while remaining competitive under full finetuning and superior on lesion tasks, where labeled data is scarce. These results indicate that a single frozen encoder can be reused across diverse segmentation tasks, reducing the cost of adapting and deploying such models in clinical workflows. We will make our pre-trained model weights publicly available.
[CV-50] Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image Enhancement
链接: https://arxiv.org/abs/2608.05843
作者: Yaozi Zhong,Xingxing Yang,Shaohui Mei,Mingyang Ma
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages, 7 figures, 7 tables. Yaozi Zhong and Xingxing Yang contributed equally. Code: this https URL
Abstract:Restoring high-fidelity remote sensing imagery from extreme low-light degradation is indispensable for reliable Earth observation and downstream machine vision. However, under severe noise and illumination corruption, existing methods suffer from attention drift, erroneously aggregating features across distinct physical boundaries and causing severe structural blurring and color distortion. To address this, we propose HALO, a dual-prior-driven enhancement framework that formulates enhancement as a guided feature aggregation problem driven by foundation model priors. Specifically, an illumination-invariant semantic prior provides regional homogeneity as a positive bias for content-consistent aggregation, while a pseudo-3D topological prior provides boundary heterogeneity as a negative penalty to strictly prevent cross-boundary confusion. To cooperatively incorporate these two priors, we propose a Homogeneity-Heterogeneity Cooperative Attention Module (H2CAM) to resolve feature conflicts during cross-modal prior fusion. Extensive experiments demonstrate that HALO achieves state-of-the-art performance across 8 challenging synthetic and real-world remote sensing benchmarks, significantly improving physical boundary sharpness and color fidelity while maximizing the preservation of discriminative features for downstream Earth observation tasks.
[CV-51] Accurate Localization of Road Traffic Objects on the Road Plane Using Surveillance Camera Imagery
链接: https://arxiv.org/abs/2608.05840
作者: Jan Gawroński,Witold Czajewski
类目: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
备注: 8 pages, 7 figures. Accepted for publication in the proceedings of the 2026 Progress in Applied Electrical Engineering (PAEE) conference
Abstract:Accurate vehicle localization from monocular roadside surveillance cameras is important for intelligent transportation systems, traffic monitoring, and traffic conflict analysis. Standard approaches often estimate vehicle position from the center of the detector bounding box, which can produce large errors due to perspective distortion and parallax, especially for elevated cameras and large vehicles. This paper proposes a two-stage geometry-aware localization pipeline that estimates the projection of the vehicle footprint onto the road plane. First, vehicles are detected using a YOLO26-based detector. Second, a dedicated ResNet34 regression network predicts four corner points corresponding to the projected vehicle base. The final position is computed as the geometric center of the predicted quadrilateral. The method was trained on synthetic data generated in CARLA and fine-tuned on real-world roadside imagery from DAIR-V2X. Experiments on synthetic and real data showed clear improvements over naive bounding-box-center localization. On DAIR-V2X, the mean image-space localization error decreased from 31.77 px to 15.30 px, a 51.8% improvement, while the median error decreased to 4.29 px. Median ground-plane error for medium-range vehicles decreased from 5.52 m to 0.90 m, and for far-range vehicles from 8.67 m to 1.84 m. The results also show that contextual information surrounding the detector bounding box is important for geometric localization. The largest gains were observed for distant vehicles and geometrically challenging cases affected by strong perspective distortion and parallax. Comments: 8 pages, 7 figures. Accepted for publication in the proceedings of the 2026 Progress in Applied Electrical Engineering (PAEE) conference Subjects: Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV) Cite as: arXiv:2608.05840 [cs.CV] (or arXiv:2608.05840v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2608.05840 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Witold Czajewski [view email] [v1] Thu, 6 Aug 2026 10:07:53 UTC (14,995 KB)
[CV-52] Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models
链接: https://arxiv.org/abs/2608.05834
作者: Max Rehman Linder
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:In this technical report, I present a new method for guiding image generation in the context of Virtual- Try-On (VITON). The proposed method leverages new open source Ai models to augment the image data with labels, such as lengths and styles. By training adapters with these labels paired with images of the garments, the model can produce a more diverse set of images that the user can control. For the end user, such as a retailer, this means that they can assure that the produced image is as true to the true fit as possible, not misleading consumers
[CV-53] Bayesian adaptively-weighted ensembles for few-shot abdominal segmentation MICCAI2026 MICCAI
链接: https://arxiv.org/abs/2608.05815
作者: Abbas Al-Sabbagh,Shalom F. Mushtaq,Tomás M. da Silva,Kushagra Soni,Binawei Gbamila,Sri Atluri,Qianye Yang,Yipeng Hu,Claire C. Villette,Shaheer U. Saeed
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at DEMI at MICCAI 2026 - The 4th MICCAI Workshop in Data Engineering in Medical Imaging
Abstract:Few-shot learning has emerged as a promising approach for anatomical segmentation when labelled data are scarce. However, different few-shot learning algorithms exhibit complementary strengths and weaknesses, with performance varying across anatomical targets and institutions. Existing few-shot segmentation ensembles, that combine predictions from multiple algorithms, typically employ fixed weighting schemes and therefore cannot adjust model contributions according to the target domain. In this work, we propose a Bayesian adaptively-weighted ensemble framework for segmentation under label scarcity and domain shift. Multiple few-shot segmentation algorithms are first adapted using a small labelled support set. Bayesian optimisation is then used to automatically identify ensemble weights that maximise segmentation performance on a target-domain validation set. The learned weights are subsequently fixed and applied to combine predictions on previously unseen query images from the target domain. The proposed framework is evaluated on the Cross-institution Male Pelvic Structures dataset using held-out anatomical structures and institutions to simulate simultaneous label scarcity and institutional domain shift. Results demonstrate statistically significant improvements over individual few-shot learners, fixed-weight ensembles, training-from-scratch baselines and recent state-of-the-art ensembling approaches. By adapting model contributions to the target anatomy and institutional domain, the proposed framework provides a practical mechanism for deploying segmentation systems to new clinical sites under severe annotation constraints.
[CV-54] Energy-Guided Flow Matching
链接: https://arxiv.org/abs/2608.05811
作者: Haoyang Tong(1 and 2),Yu He(2),Fang Li(2),Lichen Ma(2 and 3),Jingling Fu(2),Dong Chen(2),Zhen Chen(2),Junshi Huang(2),Jie Cao(1) ((1) MAIS amp; NLPR, CASIA, (2) a href=“http://JD.com” rel=“external noopener nofollow” class="link-external link-http"this http URL/a, (3) Xi’an Jiaotong University)
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 19 pages, Code: this https URL
Abstract:Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly. In this paper, we introduce Energy-Guided Flow Matching(EG-FM) that explicitly models a coarse-to-fine generative trajectory by moving endpoint. Specifically, EG-FM replaces the fixed endpoint with a heat-kernel-filtered endpoint that evolves smoothly from low-frequency image to clean this http URL fraction of high-frequency signal in moving endpoint is released by an image-specific energy-guided scheduling, leading to the re-targeting of velocity in flow this http URL framework requires no adaptation of the backbone and training data, bringing negligible cost on the training and inference stages. In our experiment, EG-FM consistently achieves lower FID on the ImageNet class-conditional image generation task at 256 \times 256 with fewer epochs, reaching an FID of 1.55 at 200 epochs and 1.45 at 600 epochs. We continue training the generation task on the setting of 512 \times 512 resolution, yielding a FID of 1.58 after only 40 high-resolution adaptation this http URL, we transfer EG-FM on text-to-image generation and achieve 0.85 on GenEval score and 83.9 on DPG-Bench. Code is available at this https URL.
[CV-55] STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models
链接: https://arxiv.org/abs/2608.05808
作者: Songpan Gao,Yajie Zhang,Guanxing Chen,Jiayu Qian,Zhenzhen Liu,Shijun Li,Xiaowei Zhu,Yao Hu,Kay Chen Tan,Yu-An Huang,Shiqi Wang,Zhi-An Huang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dynamic environments. Mainstream incremental learning methods typically mitigate this by rehearsing raw historical images. However, this pixel-level rehearsal incurs significant storage overhead, raises privacy concerns, and fails to adequately capture the true data distribution with sparse exemplars. Inspired by human cognitive mechanisms, we propose a novel framework termed Semantic Text-Anchored Incremental Learning (STAIL) for sequential clinical tasks. To overcome the rehearsal bottleneck, STAIL introduces an asymmetric semantic consolidation buffer (SCB). By incorporating a minimal set of image anchors and extensive textual descriptions, the SCB enables dense semantic reconstruction of old tasks at a minimal storage cost. Furthermore, we design an LLM-derived Semantic Anchoring Mechanism (LSAM) that leverages the stable semantic space of frozen large language models as developmental priors. This mechanism explicitly anchors evolving visual features to textual representations, guiding and constraining plasticity and stability at both macroscopic and microscopic levels. Extensive experiments across three heterogeneous medical datasets, covering fundus, ultrasound, and X-ray imaging, demonstrate that STAIL acts as a highly effective plug-and-play module. It comprehensively enhances the performance of various existing baselines, achieving average gains of 2.24% in AAA-AUC for sustained performance and 3.55% in BWT-AUC for reduced forgetting. Code is available.
[CV-56] Ordered Diffusion for 3D Human Registration
链接: https://arxiv.org/abs/2608.05804
作者: Mattia Masiero,Ilya A. Petrov,Daniel Cremers,Gerard Pons-Moll,Riccardo Marin
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at GCPR 2026
Abstract:3D human registration has historically been treated as a regression task, assuming a unique ground-truth alignment exists between the template and an input point cloud. In reality, acquisition noise, occlusions, and unknown soft tissue dynamics introduce inherent ambiguity into human scans. Regression-based methods consequently converge to an average prediction, often failing to represent a plausible geometry. In our work, we embrace such uncertainty by modeling the registration as a distribution of alignments. We propose ODin, which formulates registration as a 3D diffusion process that generates a point cloud aligned with the target geometry while preserving template semantics through consistent point ordering. To achieve this, ODin relies on global, local, and positional conditioning, guiding each point to its correct location. Our experiments demonstrate that such a generative formulation not only outperforms its regression-based baseline, but also establishes a new state of the art, surpassing highly engineered methods while reducing the registration time by two-thirds. Pre-trained models and code are available at this https URL.
[CV-57] Vorch-Omni: Multi-Task Orchestration of Sight and Sound
链接: https://arxiv.org/abs/2608.05803
作者: Vorch Team,Xiaoyu Chen,Yang Ding,Cong Han,Menglin Han,Yuxin Hong,Jiebo Hou,Zequn Jie,Xiang Li,Jing Liu,Qi Liu,Yulei Lu,Siyuan Luo,Lin Ma,Xin Ma,Yinlong Qian,Peng Shi,Fang Wan,Siqi Wang,Yaohui Wang,Yaole Wang,Yidi Wu,Siqian Yang,Mingyu Yin,Haoran Yu,Gang Yue,Lisai Zhang,Yuting Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL
Abstract:Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.
[CV-58] XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
链接: https://arxiv.org/abs/2608.05799
作者: Yixiang Chen,Jiabing Yang,Yuan Xu,Qisen Ma,Keji He,Peiyan Li,Kai Wang,Ziheng He,Xiangnan Wu,Jing Liu,Nianfeng Liu,Yan Huang,Liang Wang
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.
[CV-59] KVAE: Family of Tokenizers for Multimodal Generative Models
链接: https://arxiv.org/abs/2608.05798
作者: Andrey Shutkin,Denis Parkhomenko,Ivan Kirillov,Kirill Chernyshev,Kirill Malakhov,Ilia Vasiliev,Ilia Trushkin,Valeriya Kobenko,David Chikovani,Alexander Ivanov,Azat Saginbaev,Egor Silvestrov,Ivan Mikheev,Konstantin Zakharov
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)
备注:
Abstract:Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D – two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at this https URL and this https URL.
[CV-60] VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation
链接: https://arxiv.org/abs/2608.05782
作者: Lala Shakti Swarup Ray,Vitor Fortes Rey,Mengxi Liu,Paul Lukowicz,Bo Zhou
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Under review
Abstract:Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and subject-generalization settings. Synthetic IMU generation can reduce this dependency and enhance HAR machine learning model’s performance, but existing approaches face a trade-off without addressing all factors: video-driven methods are visually grounded but sensitive to pose-estimation errors, while text-driven methods are controllable but often weakly grounded in how activities are actually performed. We present VSMP-IMU, a video-grounded framework for controllable synthetic IMU generation based on a structured Semantic Motion Program (SMP), which separates activity-defining semantics from label-preserving variation. Given an input video, VSMP-IMU extracts and augments an SMP, uses it to synthesize motion, converts the motion into virtual IMU signals, and grounds the resulting signals to the target wearable domain. We evaluate VSMP-IMU against state-of-the-art synthetic data generation methods on five public IMU-HAR datasets under leave-one-person-out evaluation. VSMP-IMU achieves an average Macro-F1 of 78.33%, improving over real-only training by 9.77% and over the strongest prior synthetic baseline by 4.04%. In low-resource settings with reduced training data-samples, it improves over real-only training by 18.54% and over the strongest prior synthetic baselines by more than 6% on average. Under long-tail evaluation in imbalanced datasets, it improves tail-class Macro-F1 by 19.86% over Real-only training and by 4.76% over SOTA. These results show that structured video-grounded semantics provide a practical foundation for controllable, wearable-relevant synthetic sensor data generation.
[CV-61] Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
链接: https://arxiv.org/abs/2608.05780
作者: Bo Zhang,Wenxin Wang,Feng Chen,Zhihao Zhang,Zixuan Wang,Changsheng Li,Yinjie Lei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project Page: this https URL
Abstract:Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. However, existing approaches predominantly rely on external proxy scorers and rigid heuristic rules, inevitably suffering from misalignment with the target MLLM’s intrinsic evidence and failing to accommodate the non-uniform spatiotemporal information density. In this paper, we propose a fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence. Our method efficiently probes visual evidence via sparse prefilling as a structured prior to guide distribution-aware dynamic sampling. Specifically, we efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart. Conditioned on three complementary attention components derived from this prior, we design a lightweight selector that not only precisely locates query-relevant timestamps but also adaptively adjusts the local sampling rate and spatial resolution. To enable evidence-conditioned spatiotemporal sampling, we formulate the selector as a stochastic policy and optimize it via GRPO under a joint accuracy–efficiency reward. By rewarding correct predictions under lower visual cost through group-relative comparisons, our method encourages the policy to allocate computation dynamically according to the information density of each video. Across three long video understanding benchmarks, EviSelect achieves superior performance compared to existing methods while reducing selected visual tokens by about 50% and achieving a 3.9x end-to-end speedup.
[CV-62] Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
链接: https://arxiv.org/abs/2608.05776
作者: Lisai Zhang,Yidi Wu,Qi Liu,Xin Ma,Yang Ding,Gang Yue,Siqian Yang,Jingyuan Chen,Lin Ma,Yaohui Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated video and audio. However, models are trained on clean ground-truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over-smoothing, and audio-visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow-matching noise level at which residuals are produced. We propose Vorch-Director, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch-Director produces more realistic autoregressive histories while retaining efficient teacher-forcing training. Built on the audio-visual LTX-2 diffusion transformer, Vorch-Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long-horizon generation. Together with a clean conditioning sink and mixed-task training, Vorch-Director supports multi-shot, multi-subject, reference-guided audio-visual long-video generation. We evaluate Vorch-Director on ST-Bench and introduce a new long-horizon audio-visual benchmark with metrics for quality drift and long-range consistency. Extensive experiments demonstrate improved stability and audio-visual fidelity over strong baselines.
[CV-63] SR-JEPA: Learning Predictive Latent State in 3D Scenes
链接: https://arxiv.org/abs/2608.05774
作者: Zihan Zhou,Qifu Wen,Xi Zeng
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: 17 pages, 5 figures, 9 tables
Abstract:Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEPA, a point-native JEPA for scene-scale point clouds whose original frozen predictive pathway can be queried at a supplied location. At evaluation, every point of one object is removed before encoding and replaced by the same shape-free 32-point query at its centroid. Training uses only self-contained 3D EMA targets: no reconstruction, semantic labels, language, or lifted 2D features. On 5,953 held-out ARKitScenes objects, the imputed latent reaches 43.13% semantic-identity macro accuracy, 22.18 points above the strongest floor. Randomizing the prediction path removes 9.78 points, while substituting matched donor context removes 21.98 points. On 8,570 Sr3D support pairs, the full latent reaches 41.15 AP; identity decoded from the missing-object latent, combined with anchor identity and geometry, reaches 39.37 AP, leaving an unresolved 1.78-point residual. These results reveal a queryable, compositional 3D predictive state: the model completes context-dependent entity content, which downstream computation combines with metric geometry.
[CV-64] HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
链接: https://arxiv.org/abs/2608.05771
作者: Aohua Li,Jin Kuang,Yubing Lu,Pingping Liu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 15 pages, 9 figures, 9 tables. Code: this https URL
Abstract:Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance often degrades markedly when generalizing to unseen infrared domains. Existing methods primarily improve detection by enhancing target responses and suppressing background interference. However, when trained on only a limited set of source domains, their learned decision rules are inevitably established from a restricted range of source-domain target-background relation patterns. We formulate this cross-domain failure as target-background relation shift: unseen domains may exhibit relation patterns that are not observed during training, thereby weakening the discriminative capability learned from the source domains. To address this problem, we propose HyTBE, a Hyperbolic Target-Background Expert model that expands source-domain relation patterns and adaptively adjusts visual representations using explicit relation cues. The Target-Background Relation Intervention selectively perturbs either targets or backgrounds, broadening the observable relation patterns during training while maintaining valid supervision. Subsequently, the Hyperbolic Relation Modeling maps multi-scale visual cues into a Poincaré ball and characterizes the target-background relation of each feature token according to its relative distances to the target and background anchors. The Hyperbolic-guided MoE Adapter further uses these hyperbolic relation representations to calibrate multi-scale visual features and aggregate expert-specific feature corrections for different relation patterns. Leave-one-domain-out experiments on NUAA-SIRST, NUDT-SIRST, and IRSTD-1K demonstrate that HyTBE achieves stronger cross-domain generalization than competitive baselines.
[CV-65] Flow-Map Distillation on Relation Manifolds for Image Restoration
链接: https://arxiv.org/abs/2608.05769
作者: Zihao He,Songhua Liu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 9 pages, 7 figures. Accepted to ACM Multimedia 2026
Abstract:Knowledge distillation for image restoration typically aligns intermediate features or relation matrices between teacher and student networks as static targets, ignoring the dynamic structure of the knowledge transfer process. In this paper, we propose Flow-Map Distillation on Relation Manifolds (FoRM), which reformulates relation-based knowledge transfer as a continuous flow mapping problem on the relation manifold. Rather than regressing a constant velocity field between student and teacher relation states, FoRM learns a flow map operator \mathcalF_\theta(\mathbfz, t, s) that directly predicts the relation state at any target time s given the current state at time t , enabling richer trajectory-level supervision. To ensure global self-consistency of the learned flow map, we introduce a safe semigroup consistency constraint that enforces compositional agreement using ground-truth bridge states, eliminating phantom-state error accumulation. An endpoint anchoring loss further prevents the operator from drifting away from the teacher target. Extensive experiments on five image restoration tasks, including super-resolution, deraining, denoising, deblurring, and low-light enhancement, demonstrate consistent gains over state-of-the-art distillation baselines across multiple backbone architectures, reducing training variance by approximately 50% compared to naive flow matching distillation while achieving superior restoration quality.
[CV-66] Beyond Relevance: Bayesian Evidence Acquisition for Agent ic Whole-Slide Image Reasoning
链接: https://arxiv.org/abs/2608.05757
作者: Bryan Wong,Xun Xu,Huazhu Fu,Nancy F. Chen,Mun Yong Yi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Whole-slide image (WSI) reasoning requires an agent to sequentially acquire visual evidence before answering a diagnostic question. Existing training-free agentic frameworks formulate this process as iterative patch retrieval based on semantic relevance to the question. However, semantic relevance does not necessarily imply diagnostic informativeness in computational pathology, where competing diagnoses often exhibit similar and overlapping morphological patterns, making many patches semantically relevant yet diagnostically non-discriminative. Consequently, relevance-based retrieval may acquire redundant observations and leave diagnostic uncertainty unresolved. We propose BEACON, a plug-and-play agentic framework that reformulates WSI reasoning as a Bayesian evidence acquisition problem. BEACON maintains a probabilistic belief over competing diagnostic hypotheses and sequentially acquires patches by maximizing expected information gain (EIG) to reduce diagnostic uncertainty. An evidence controller then determines whether to answer, acquire additional evidence, or perform higher-resolution inspection. Built entirely from off-the-shelf foundation models, BEACON requires no additional training or fine-tuning. Extensive zero-shot experiments across five WSI-VQA benchmarks demonstrate that BEACON achieves the strongest overall performance among training-free agentic frameworks while substantially improving evidence acquisition efficiency, establishing Bayesian evidence acquisition as a principled paradigm for uncertainty-aware agentic WSI reasoning. The code is available at this https URL
[CV-67] GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
链接: https://arxiv.org/abs/2608.05747
作者: Qifeng Zhang,Kaixiang Huang,Heng Dong,Huang Fang,Junting Chen,Junjie Zhu,Yonghang Chen,Zhiyu Zhang,Wei Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address this limitation, we introduce the Global-Spatial-Temporal Benchmark (GST-Bench), a VQA benchmark for global spatial intelligence in video understanding, comprising human-verified questions derived from 6,790 minutes of synthetically generated video. It requires models to perform accurate spatial inference from novel viewpoints unseen in the input video and to map egocentric observations onto global top-down images. A comprehensive evaluation of 22 state-of-the-art VLMs exposes a striking gap between models and humans: the strongest zero-shot model attains only 42.68, far below the human score of 79.08. To probe the cause of this gap, we construct GST-Bench-Local and find that models, despite strong local spatial understanding under the same task formulation, still fail to consolidate long-horizon observations into a globally consistent scene representation. We further provide GST-Train, a dataset for global spatial reasoning, as a complementary resource to facilitate future research on this challenge.
[CV-68] UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
链接: https://arxiv.org/abs/2608.05745
作者: Yushe Cao,Shikun Feng,Fei Shen,Haikuo Peng,Jianqiang Xia,Yiheng Zhu,Dianxi Shi,Chun Yu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: 17 pages,21 figures
Abstract:Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-conditioned video inpainting and rely on separate modules for human parsing, pose estimation, and garment warping. This multi-stage design complicates deployment and, more critically, allows errors in explicit geometric priors to propagate irreversibly into the generated video. We present UniVVT, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference. At its core, a scene-task perceiver built on a Multimodal Large Language Model jointly encodes the source video, target garment, and task instruction into compact, task-aware latent tokens, implicitly capturing what to transfer and where and how to transfer it. A lightweight semantic bridge then aligns these tokens with the conditioning space of a diffusion-based video generator, enabling coherent garment transfer. To robustly couple the heterogeneous components, we devise a three-stage progressive training strategy comprising semantic alignment, joint task adaptation, and flexible-resolution refinement. Extensive experiments demonstrate that UniVVT achieves state-of-the-art performance across multiple benchmarks, validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
[CV-69] ConceptADapt: Concept-guided Adaptive Feature Reconstruction with Dynamic Attention for Few-Shot Industrial Anomaly Detection
链接: https://arxiv.org/abs/2608.05743
作者: Yufei Li,Yicheng Ruan,Long Tian,Dongsheng Wang,Liang Bao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 11 pages, 7 figures
Abstract:Few-shot industrial anomaly detection (FS-IAD) focuses on detecting and localizing visual defects in industrial inspection during the cold-start phase, where only a limited number of normal training samples are available per category. Recent advances in this field predominantly leverage visual features from foundation-model and have achieved promising performance. Despite the strong representational power of foundation-model features, the model generalization remains fragile due to the extreme scarcity of normal training this http URL address this pivotal issue, we propose ConceptADapt, a concept-guided adaptive feature reconstruction model with dynamic attention. Specifically, our model pre-learns a set of fixed normal concepts from the limited support features and leverages them to mine relationships with query features, thereby recalibrating their statistics for improved anomaly detection at test time. To mitigate the prevalent feature shortcut problem, which is particularly severe under low-data regimes, we further develop a dynamic attention mechanism integrated with sparse autoencoders to learn robust normal concepts during training. Moreover, to enable fast adaptation during inference, our model remains lightweight by incorporating LoRA into the attention module, which introduces only minimal updating this http URL experiments on three widely adopted FS-IAD benchmarks, including MVTec-AD, VisA, and MPDD, demonstrate that our model consistently outperforms state-of-the-art (SOTA) approaches across both detection and localization tasks, achieving significant improvements under various shot settings.
[CV-70] LiteKD-Net: Lightweight Knowledge-Distilled Network for Mobile Image Denoising
链接: https://arxiv.org/abs/2608.05739
作者: Zhou Zhiyi
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Mobile image denoising requires both good restoration quality and low computational cost. In addition, it’s annoying to collect large-scale LQ-GT clean pairs. As a result, we propose LiteKD-Net, a lightweight knowledge-distilled network for mobile image denoising. First, a physics-guided noise simulation pipeline generates paired training data by adding pixel crosstalk compared with pipelines applied to cameras. Next, we adapt the Real-ESRGAN to identity-resolution denoising and construct a lightweight Student using Lite-RRDB blocks based on depthwise separable convolutions. Third, feature-level knowledge distillation is applied to transfer the Teacher’s restoration capability to the Student without introducing additional inference cost. Experiments on real-world datasets show that our model reaches great reduction in runtime and increase in the inference rate with good restoration quality. Our model also reaches the best in all metrics compared with SwinIR. These results indicate that LiteKD-Net provides a great trade-off between restoration quality and computational efficiency.
[CV-71] Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams
链接: https://arxiv.org/abs/2608.05728
作者: Feiyu Ji,Xiang Li,Hao Ma,Tianxiang Huang,Qingxin Lu,Mengqi Ji,Lei Han,Xiaokang Yang,Xiaoyun Yuan
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Optics (physics.optics)
备注: 9 pages, 5 figures
Abstract:Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolute appearance, making faithful reconstruction intrinsically challenging. The central challenge lies in associating event-derived target-time structures with relevant appearance information from the reference frame, especially under complex motion and long temporal intervals. In this work, we propose Engram-E2VID, a structure-guided framework that reconstructs target frames through the generative activation of appearance engrams. Specifically, the reference frame is encoded into token-space appearance engrams, while the event stream and reference context are transformed into a target-time motion-structure scaffold that captures motion boundaries and event-induced structural changes. Within a one-step diffusion backbone, scaffold-derived structural tokens progressively interact with and activate relevant appearance engrams across layers. This token-space association allows target structures to access reference appearance without relying on direct pixel-wise correspondence, while the diffusion prior complements uncertain or newly revealed regions. Across three benchmarks, Engram-E2VID improves PSNR by up to 3.29 dB and reduces LPIPS by up to 0.08 over the strongest same-input baseline, while degrading more slowly as the reconstruction interval increases.
[CV-72] PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models
链接: https://arxiv.org/abs/2608.05720
作者: Xi Zeng,Haojie Ren,Ziying Song
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 16 pages, 5 figures
Abstract:We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not ensure that a representation preserves physical states and action consequences. We identify three failure modes in JEPA world models: physical invariance collapse, physical identifiability collapse, and counterfactual dynamics collapse. PhyLatent addresses them through three training pathways: physical invariance, physical identifiability, and counterfactual dynamics, implemented with physical state grounding, future representation alignment, static visual invariance, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three failure rates from 15.60%, 6.71%, and 8.41% to 7.53%, 0.95%, and 4.62%, respectively, and improves model predictive control (MPC) success from 70.0% to 78.1%. With the same architecture and planner, it further improves success from 81.0% to 98.0% on TwoRooms and remains competitive on Reacher and PushT. These results show that global non-collapse alone is insufficient for learning a reliable JEPA worldmodel state space.
[CV-73] Iterative Hybrid Discrete-Continuous Viewpoint Planning for UAV Photogrammetry
链接: https://arxiv.org/abs/2608.05718
作者: Alan Grech,Daniel Pisani,Andre Grima,Carl James Debono,Saviour Formosa,Dylan Seychell
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 6 pages, 3 figures, 2 tables. Accepted for publication at the 14th IEEE European Conference on Visual Information Processing (EUVIP 2026)
Abstract:Unmanned aerial vehicle (UAV) photogrammetry requires camera networks that provide sufficient surface coverage, image overlap, parallax, and resolution, yet conventional flight patterns are often poorly adapted to scene geometry resulting in local reconstruction errors. This paper proposes an iterative hybrid discrete-continuous viewpoint planning method for targeted UAV photogrammetry from a proxy reconstruction. The method scores sampled surface points using photogrammetric heuristics based on frontality, imaging distance, parallax, and multi-view observation count, while also evaluating the full viewpoint set in terms of visibility, pairwise overlap, and graph connectivity. Candidate viewpoints are generated around weakly observed regions, refined using clustered Covariance matrix adaptation evolution strategy (CMA-ES) optimisation, and removed when redundant. The final flight path combines close-range detail viewpoints with wider model-coverage viewpoints, balancing local reconstruction quality with global image-network robustness. Evaluation on three synthetic scenes shows that the proposed method improves both reconstruction accuracy and completeness compared with prior UAV path-planning methods.
[CV-74] One Ranking Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
链接: https://arxiv.org/abs/2608.05707
作者: Wang Chen,Yu Chen,Xiang Wang,Shuai Li,Jinfa Huang,Xiawu Zheng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 21 pages, 7 figures, 7 tables
Abstract:Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budget varies with the downstream LMM, reasoning demands, and latency constraints, a practical selector should serve multiple budgets. However, existing methods typically optimize an isolated frame subset for each predefined budget: when the budget changes, previously selected evidence may be replaced rather than progressively augmented. Ranking frames by a fixed score would allow prefix reuse across budgets, but it ignores the distinct roles of different ranking positions. In this paper, we formulate long-video frame selection as a Matryoshka ranking problem: constructing a single priority sequence whose small prefixes concentrate query-conditioned evidence, while progressively larger prefixes preserve this evidence and add broader temporal context. Efficiently constructing such a ranking is itself challenging, as densely sampling long videos and evaluating frame-query relevance incurs substantial overhead. We therefore introduce Matryoshka Evidence-to-Context (MEC) Frame Selection, a training-free framework that builds a reusable sparse video index, discovers candidates through sparse probing and local zooming, and greedily constructs a position-adaptive ranking: early positions emphasize evidence; later positions progressively favor temporal coverage while preserving visual diversity. A single ranking can thus be truncated to any target budget without rerunning the selector. Across four benchmarks and six frame budgets, MEC improves average accuracy over uniform sampling by 3.77 percentage points, matches strong state-of-the-art selectors, and reduces end-to-end selection latency by 47.37-51.19%.
[CV-75] LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
链接: https://arxiv.org/abs/2608.05706
作者: Jiarui Yang,Jiale Zhange,Jiawei Li,Hang Guo,Wen Huang,Jinpeng Wang,Peidong Liu,Shu-Tao Xia
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms. Recently, latent action models (LAMs) have alleviated this bottleneck by learning action representations directly from unlabeled human videos in a self-supervised manner. Nevertheless, most existing LAMs rely on single-view inputs and operate primarily in 2D pixel space, raising a fundamental question: can simply incorporating multi-view videos into LAM training endow the learned latent actions with 3D-aware perception? Our study shows that the answer is negative. The primary reasons lie in future-frame appearance leakage as well as inter-camera appearance discrepancies and viewpoint variations. To address these issues, we propose LAWM-3D, which introduces three tightly coupled key designs: (1) a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions; (2) a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, thereby explicitly providing cross-view geometric correspondences; and (3) a non-injective RGB-D joint reconstruction objective that prevents shortcut learning from future-frame appearance information, forcing the LAM to focus supervision on motion cues with geometric significance. Importantly, these components are not simply stacked but are tightly coupled through a unified motivation. Built upon a two-stage paradigm of large-scale human video pretraining followed by robot fine-tuning, extensive experiments demonstrate that the proposed 3D-aware latent actions significantly improve world model performance, achieving SOTA results in generation quality, physical consistency, and generalization ability.
[CV-76] G2ARD-GS: Geometry-Guided Anchor-Regularized Gaussian Splatting Distillation
链接: https://arxiv.org/abs/2608.05704
作者: Puyuan Zhang,Jianming Huang,Wenkai Ye,Wei Dong
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Dense colored LiDAR maps provide accurate city-scale geometry, but lifting them into 3D Gaussian Splatting (3DGS) retains millions of primitives, making the resulting models costly to store, transmit, render, and adapt. Aggressive primitive reduction alleviates this burden, but can remove the local surface support needed for stable novel-view synthesis and downstream geometric use. We introduce G ^2 ARD-GS, a geometry-guided distillation method that converts a dense Gaussian prior instantiated either as a training-free point-cloud lift or a trained GS model into a compact, reusable representation. G ^2 ARD-GS progressively consolidates the prior into surface-aware representatives, then recovers appearance on the resulting fixed topology under construction-time anchor constraints, with no primitives added or removed during recovery. Under limited supervision, geometry-aware view selection allocates the available view budget. On MatrixCity, G ^2 ARD-GS achieves the best PSNR, SSIM, and LPIPS across matched 5\times – 30\times compression budgets, outperforming PUP by 3.2 – 6.8 ,dB in PSNR. When reused as frozen geometry, the compact model improves off-trajectory appearance adaptation by 3.7 – 4.9 ,dB over PUP 3D-GS and preserves image-to-model registration accuracy on Cambridge KingsCollege at 30\times compression. Project page: this https URL.
[CV-77] StreamArena: Toward Continuous Interactive and Long-Horizon Agent ic Streaming Video Understanding
链接: https://arxiv.org/abs/2608.05703
作者: Xichen Zhang,Guankai Li,Yinghao Zhu,Shijian Wang,Sitong Wu,Shaozuo Yu,Meng Chu,Yuan Lu,Jiaya Jia
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predominantly rely on brief clips and multiple-choice formats. This design allows minimal baselines that process only the last four frames to match or surpass complex streaming models, while answer options also expose language shortcuts. We introduce StreamArena, a benchmark for hour-scale, interactive streaming video understanding. StreamArena contains 243 full-length videos averaging 88.8 minutes and 3,646 rigorously annotated, open-ended question-answer pairs that evaluate real-time perception, historical retrospection, proactive interaction, and multimodal tool utilization. Evaluation across diverse systems exposes a tension between continuous interaction and long-horizon multimodal comprehension. Methods that retain only recent frames cannot recover distant events, methods that convert past observations into text lose visual evidence, and methods that repeatedly compress visual memory struggle to preserve fine-grained details over time. We address this tension with StreamMind, a two-tier architecture that assigns latency-critical interaction and proactive monitoring to independently scheduled frontend workers, while backend workers asynchronously construct persistent multimodal memory and perform historical recall and external search. StreamMind outperforms existing streaming baselines across all four capabilities and reduces query-to-answer latency by reusing persistent state.
[CV-78] AU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding
链接: https://arxiv.org/abs/2608.05699
作者: Kepeng Yang,Dongxuan Liu,Rongxin Gao,Zixin Su,Rui Wu,Shuzhao Xie,Chenxin Li,Panwang Pan,Yuzhi Huang,Yue Huang,Jingyan Jiang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unfolds, and interpret why it violates the expectations of the surrounding scene. Video anomaly understanding (VAU) seeks to endow models with a similar capability, moving beyond deciding whether a video is anomalous toward explaining how the event develops and why it matters. Although recent vision–language models (VLMs) can generate detailed and plausible anomaly descriptions, their semantic fluency does not ensure that these interpretations remain grounded in the correct anomaly instance over time. Existing benchmarks typically evaluate tracking and semantic understanding through separate protocols, leaving such instance–semantic inconsistency largely unmeasured. We therefore introduce TAU-Bench, a track-centric benchmark for jointly evaluating anomaly instance tracking and fine-grained anomaly understanding. TAU-Bench contains 1,118 videos, 1,454 tracks, and 202,438 pixel-level masks spanning 49 event and 45 scene categories, together with track-centric annotations that connect instance-level identification, event-level understanding, and scene-level reasoning. To build TAU-Bench at scale, we developed an automated data engine integrating anomaly suitability filtering, anomaly instance track construction, hierarchical caption annotation, and human quality control. Evaluations across representative VLM families show that models producing plausible anomaly interpretations may still fail to localize and track the correct instance reliably, revealing a persistent gap between semantic reasoning and visual grounding. These findings therefore highlight instance-grounded evaluation as an important step toward more faithful and reliable VAU systems.
[CV-79] SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment
链接: https://arxiv.org/abs/2608.05691
作者: Yin-Loon Khor,Yi-Jie Wong,Jing Jie Tan,Ming Jie Lee
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Scientific images are essential for communicating experimental observations, quantitative evidence and conceptual knowledge. Unlike natural images, their quality depends on both visual clarity and scientific informativeness, making assessment challenging. In this work, we present SciQNet, a two-stage multimodal adaptation framework for scientific image quality assessment. The first stage performs domain-adaptive pretraining on scientific document images and the second stage conducts task-specific fine-tuning with joint scoring and understanding supervision. For scoring-oriented supervision, we combine instruction tuning with a Huber loss derived from rating-word logits, while understanding-oriented supervision is formulated as multiple-choice visual question answering. Experiments show that using a 40% stratified subset of the domain-adaptive data gives the best performance among the evaluated pretraining fractions, suggesting that pretraining-data relevance may be as important as pretraining-data scale. The final model achieves an SIQA-S score of 92.21, an SIQA-U score of 47.38 and a combined score of 69.80. This work presents our solution to the ICME 2026 Scientific Image Quality Assessment Challenge, which ranked 2nd in the scoring track.
[CV-80] DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation
链接: https://arxiv.org/abs/2608.05683
作者: Jiaxuan Li,Qing Xu,Xiangjian He,Yue Li,Daokun Zhang,Fiseha B. Tesema,Rong Qu
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: 10 pages, 5 figures
Abstract:Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered by uncertainty in both modalities under real-world clinical conditions. Existing vision-language segmentation methods rely on deterministic cross-modal matching, which overlooks aleatoric uncertainty from ambiguous boundaries and epistemic uncertainty from limited training data, leading to fragile performance under domain shift. To address this issue, we propose DistMedVL, a probabilistic vision-language framework that introduces a lightweight Probabilistic Cross-Modal Adapter (PCM-Adapter) upon frozen encoders to explicitly model representational uncertainty. Specifically, the PCM-Adapter comprises two sequential modules for progressive probabilistic alignment. We first devise a Mahalanobis Alignment Module (MAM) that models textual tokens as Gaussian distributions and computes patch-text compatibility via Mahalanobis distance, yielding variance-conditioned matching that downweights unreliable feature dimensions. Moreover, we devise a Distribution Flow Module (DFM) that estimates modality-wise confidence parameters and performs vision-guided refinement of textual distributions, accommodating distributional variation across imaging modalities. Extensive experiments across eight medical segmentation benchmarks demonstrate that DistMedVL outperforms state-of-the-art methods with only 6.3M trainable parameters, exhibiting superior data efficiency, perturbation robustness and cross-dataset generalization.
[CV-81] A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition ACM-MM2026
链接: https://arxiv.org/abs/2608.05673
作者: Jiaheng Chen,Jiaxing Li,Tinghe Zhang,Chaopeng Guo
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: Accepted by ACM MM 2026
Abstract:Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately distinguish the functional roles of social influence in trajectory planning. Observing that agents typically form motion plans by anticipating others’ future behaviors before making local reactive adjustments, we identify social interactions as playing staged roles, namely planning precedes reaction. We propose INTraJ, a unified framework that decomposes social influence into two stages: a planning stage constructs reference trajectories using future social information, and a reaction stage recovers local adjustments from the residual between full-context prediction and the reference. INTraJ supports both multi-target and single-target paradigms. Extensive experiments on four standard benchmarks, including Argoverse 2, Argoverse 2-ped, ETH/UCY, and SDD, demonstrate consistent improvements, particularly in FDE and long-horizon consistency, with state-of-the-art performance achieved in several settings. INTraJ reframes trajectory prediction as a planning-driven two-stage process, validating that staged social modeling is critical for stable predictions. The code is publicly available at this https URL.
[CV-82] URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation ACM-MM2026
链接: https://arxiv.org/abs/2608.05671
作者: Guoan Xu,Zhengxue Wang,Yang Xiao,Ligeng Chen,Guangwei Gao,Dongchen Zhu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: ACM MM 2026
Abstract:Previous RGB-D semantic segmentation methods commonly employ dual encoders to separately process RGB and depth inputs, followed by dedicated modules for cross-modal feature fusion. However, such designs often inadequately capture depth representations and consequently limit effective cross-modal interaction, while the additional encoder branch introduces redundant computation that hinders lightweight execution. To tackle these challenges, we propose URNet, a Unified Reparameterized RGB-D Network that performs simultaneous multi-modal feature extraction and cross-modal fusion within a single encoder. Specifically, we adopt a reparameterization strategy to compact the network architecture and facilitate fast inference. Within each Reparameterized Block (RepBlock), a Linear Gated Attention (LGA) module is introduced to fully exploit complementary RGB and depth cues across different feature scales. Furthermore, considering that decoder design has been relatively underexplored in existing RGB-D segmentation models, we develop a concise yet effective universal decoder, termed the Pyramid Merging Decoder (PMD). Extensive experiments on multiple RGB-D segmentation benchmarks demonstrate that URNet achieves state-of-the-art performance while maintaining high efficiency. Code will be available at this https URL.
[CV-83] Dual-Attention and Adversarial Transfer Networks for Sim-to-Real Cross-Orientation Wireless Sensing
链接: https://arxiv.org/abs/2608.05664
作者: Linfeng Du,Kehan Wu,Tong Zhang,Rui Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Millimeter-wave human activity recognition suffers significant performance degradation when the user’s orientation changes relative to the sensing system, yet collecting labeled multi-orientation data is labor-intensive and costly. To eliminate the need for exhaustive multi-orientation measured data, we develop a physics-guided simulator that synthesizes orientation-diverse wireless training data from single-orientation motion. Specifically, to suppress orientation-induced feature variations, we propose a dual-attention network that extracts activity-discriminative and orientation-robust representations from dual-link Doppler spectrograms. To bridge the simulation-to-reality gap, we introduce an adversarial unsupervised transfer learning mechanism that aligns feature distributions using only a small number of unlabeled target-domain samples. The S2M-Sense platform shows high fidelity in reproducing real-world signatures, validated against 60.48 GHz mmWave measured data with an average structural similarity index measure (SSIM) of 0.84 between simulated and measured Doppler spectrograms across all 4 activities and 4 orientations. Experimental results show that S2M-Sense achieves 88.33% recognition accuracy using only the dual-link multi-orientation simulated dataset, which improves to 95% after simulation-to-reality transfer learning with as few as 16 unlabeled measured samples. Both cases with and without transfer learning outperform state-of-the-art cross-domain sensing methods.
[CV-84] Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
链接: https://arxiv.org/abs/2608.05663
作者: Menglin Han,Yang Ding,Yulei Lu,Haoran Yu,Xin Ma,Junyi Chen,Zhangkai Ni,Lin Ma,Yaohui Wang
类目: Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)
备注: Project page: this https URL
Abstract:Real-time long-form avatar audio–video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio–video context is available. We present \textbfVorch-Streamer, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12–21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio–lip synchronization and strong identity preservation over long-form generation.
[CV-85] Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
链接: https://arxiv.org/abs/2608.05648
作者: Yaole Wang,Xiaoyu Chen,Xin Ma,Yang Ding,Gang Yue,Jingjing Chen,Lin Ma,Yaohui Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project page: this https URL
Abstract:Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.
[CV-86] ChronoVision: Temporal Reasoning via Latent State Reconstruction
链接: https://arxiv.org/abs/2608.05631
作者: Yifan Shen,Jian Xu,Boyi Li,Yuner Zhang,Tianjiao Yu,Bingxuan Li,Houze Yang,Rushi Wang,Xu Cao
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.
[CV-87] SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation
链接: https://arxiv.org/abs/2608.05627
作者: Mohamad Zamini,Diksha Shukla
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Training-free open-vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision-language features are produced at patch level, yet dense prediction requires a unit that simultaneously governs feature interaction, spatial support, contextual recovery, and retrieval-based correction. We present SCI-CLIP, a segment-centric inference framework built around the principle that the same region abstraction should organize all stages of dense open-vocabulary prediction. SCI-CLIP first induces a region-consistent interaction graph over frozen visual tokens, then reconstructs dense features by propagating values over this graph, augmenting them with selective cross-window support only where local evidence is insufficient. The same segment abstraction is subsequently used to construct and query an offline reference memory, aligning exemplar retrieval with the units on which prediction is made. SCI-CLIP turns frozen CLIP-style features into spatially coherent, context-aware, and retrieval-compatible dense predictions without any training. SCI-CLIP consistently improves the structural quality of dense predictions, the robustness of contextual reasoning, and the alignment of exemplar-based correction, yielding stronger open-vocabulary segmentation across eight benchmarks. Project code is available at: this https URL.
[CV-88] Dual-Output Multi-Exposure HDR Reconstruction via SDR Fusion and Gain Map Inverse Tone Mapping ECCV2026
链接: https://arxiv.org/abs/2608.05626
作者: Jinho Kim,Jinwoo Kim,Seon Joo Kim
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted to ECCV 2026
Abstract:We propose DOME-HDR, a dual-output multi-exposure HDR reconstruction framework that jointly produces a perceptually balanced SDR image and a consistent HDR image via gain map inverse tone mapping. Given three bracketed LDR inputs, DOME-HDR first synthesizes a base SDR using a LoRA-adapted latent diffusion model. A dual cross-attention fusion module injects complementary structural and color cues from the under- and over-exposed images while anchoring on the mid exposure for stability. The synthesized SDR then guides HPGM, our HDR Prior-guided Gain Map network, to predict a spatially varying gain map for reliable dynamic-range expansion. We evaluate on Kalantari, Tel, and Challenge123 using both full-reference and no-reference metrics, where DOME-HDR achieves state-of-the-art HDR reconstruction quality; ablations further confirm the effectiveness of dual cross-attention and SDR-guided gain map estimation.
[CV-89] ruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs ECCV2026
链接: https://arxiv.org/abs/2608.05616
作者: Yanqi Wu,Runhe Lai,Xinhua Lu,Qichao Chen,Zhiping Zhou,Jia-Xin Zhuang,Weijiang Yu,Ruixuan Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by ECCV 2026
Abstract:Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustworthy deployment. A key finding motivates our work: real and hallucinated object tokens are clearly separable in hidden representations, yet this separability is largely lost at the language-modeling (LM) head. We propose TruthLens, a self-evaluation framework that teaches the LM head to expose a per-object truthfulness signal without any auxiliary model or additional inference cost. Concretely, a rarely-used special token is repurposed as a reference token. For each object-token position, we extract the log-probability assigned to this special token by the LM head, and define its difference from a predefined constant as the truthfulness score. The model is then fine-tuned with an MSE objective that drives scores toward 1 for real objects and 0 for hallucinated ones, while a divergence constraint preserves the original generation capability. Despite being trained on only a limited set of object categories, TruthLens generalizes effectively to benchmarks with substantially larger label spaces. Extensive experiments across multiple LVLMs demonstrate state-of-the-art performance; notably, on Qwen2.5-VL-7B, TruthLens outperforms the previous best method on MS-COCO by over 17% in AUROC. Our code is available at this https URL.
[CV-90] ALTER: Modeling Longitudinal Changes via Regional Differencing for 3D CT Report Generation
链接: https://arxiv.org/abs/2608.05615
作者: Dongchen Li,Jitao Liang,Wei Li
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Computed tomography (CT) is widely used for clinical diagnosis and longitudinal follow-up, yet automatically generating accurate and complete radiology reports from three-dimensional (3D) CT remains challenging. Existing methods improve fine-grained correspondence between images and text by modeling anatomical regions, but remain centered on the current examination. Consequently, patient-specific longitudinal changes within individual regions remain insufficiently modeled. Meanwhile, interval changes are often distributed across multiple anatomical regions, complicating a coherent assessment of the overall longitudinal state. We propose Anatomically Localized Temporal Evidence Representation (ALTER) to address these limitations. Global Prior Integration (GPI) incorporates the prior CT and report to establish historical context for the current examination. Regional Proxy Differencing (RPD) enables each current anatomical region to retrieve a historical proxy from a single shared encoding of the prior volume and to derive localized interval evidence. Interval Change Fusion (ICF) further combines current abnormality states with region-distributed differences, converting their joint representation into change-aware soft prompts that guide report generation. ALTER achieves state-of-the-art results on most evaluation metrics across the RadGenome-ChestCT validation and CTRG-Chest-548K test sets. Code and data preprocessing details are available at this https URL.
[CV-91] LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
链接: https://arxiv.org/abs/2608.05600
作者: Yingqing Guo,Hui Yuan,Zijian He,Mengdi Wang,Zheng Ding
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training. Although the ODE and SDE share the same marginal distributions in continuous time, their finite-step discretizations can differ substantially. In particular, SDE rollouts often become blurry as the exploration noise increases, creating a mismatch between the samples used for reinforcement learning and those generated by the test-time ODE sampler. We introduce LC-GRPO, a flow-based GRPO framework with Langevin correction. Each rollout transition first takes an inference-aligned ODE Euler step and then applies a stochastic Langevin correction targeting the marginal distribution at the resulting timestep. The required score is recovered directly from the flow velocity, requiring no additional score model, while the resulting transition remains an isotropic Gaussian with a tractable likelihood for policy optimization. We theoretically show that, under suitable conditions, one Langevin correction step reduces the Wasserstein error of an imperfect ODE Euler step. At a matched randomness level, we further show that the proposed transition can be more accurate than the standard Euler–Maruyama discretization of the reverse SDE. Experiments on SD3.5-Medium, FLUX.1-Dev, and HunyuanVideo demonstrate that LC-GRPO consistently improves reward optimization across text-to-image and text-to-video tasks, preserves generation quality, and substantially narrows the gap between stochastic training rollouts and deterministic test-time ODE inference.
[CV-92] Uncertainty-Aware World Model for Aerial Image-Goal Navigation
链接: https://arxiv.org/abs/2608.05597
作者: Deyi Zhu,Haoyu Fan,Yinan Zhu,Weichen Zhang,Shilin Ma,Xinlei Chen,Yansong Tang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based methods rank candidate trajectories using predicted futures, but typically rely on only one or a few point predictions, which is inadequate for large-scale outdoor environments with substantial future-state uncertainty. To address this limitation, we propose the Uncertainty-Aware Navigation World Model (UA-NWM), an efficient latent world model for aerial image-goal navigation, which formulates trajectory scoring as conditional out-of-distribution detection. UA-NWM represents plausible futures with an uncertainty subspace and decomposes the prediction–goal discrepancy into uncertainty-explainable and unexplainable components. Only the unexplainable residual is used for scoring, enabling robust selection without multiple future samples. Extensive experiments demonstrate that UA-NWM consistently outperforms existing navigation world models while maintaining low inference latency. Real-world UAV experiments further validate its practical applicability. Project page: this https URL
[CV-93] Beyond Frame Selection: Rethinking Long-Video Understanding with MLLM s
链接: https://arxiv.org/abs/2608.05592
作者: Ziling Huang,Shin’ichi Satoh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multimodal Large Language Models (MLLMs) have achieved strong progress in video understanding, yet it remains challenging because the token limitation makes MLLMs difficult to capture temporally sparse evidence. Existing methods typically rely on uniform sampling, or frame selection, but these strategies usually optimize either broad temporal coverage or local relevance, making it difficult to preserve both global storyline context and fine-grained evidence. We propose VideoRouter(VR) that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames. It first organizes each video into a question-agnostic temporal hierarchy, which partition the video into coarse-to-fine temporally coherent segments. In this hierarchy, upper-level nodes capture broad storyline context and event progression, while lower-level nodes preserve fine-grained local details and evidence-bearing moments. This naturally gives rise to two complementary views: a global view for coverage-oriented reasoning and a local view for detail-oriented evidence recovery. We further introduce a verification-guided router to determine which view is better supported by the selected evidence and select the final answer. We validate the effectiveness of the proposed design through extensive experiments, showing that the verification-guided router effectively coordinates global and local reasoning, and that, on VideoMME, our method outperforms state-of-the-art frame selection methods by 2.9 points, respectively, under the LLaVA-Video-7B backbone. We will release the code.
[CV-94] CoordRefer: Coordinate-Aware 3D Visual Grounding from Multiview Images
链接: https://arxiv.org/abs/2608.05569
作者: Haijie Li,Jiaxin Zhang,Dave Zhenyu Chen,Youyu Chen,Yanmin Wu,Jian Zhang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Multiview image-based 3D visual grounding predicts a coordinate frame to define a coordinate system and then regresses a 3D bounding box for localization. However, existing methods jointly optimize coordinate frame selection and box regression, leading to coordinate-relative box ambiguity and degraded grounding performance. This ambiguity arises because the same box admits different numerical representations across coordinate frames, creating multiple optimization targets and yielding invalid compromise predictions. To tackle this challenge, we propose CoordRefer, a coordinate-aware framework that decouples coordinate frame selection from coordinate-conditioned grounding. CoordRefer first selects a reference frame to define the coordinate system and then conditions 3D box prediction on the coordinate system. We perform coordinate-aware supervised fine-tuning to establish coordinate frame selection and coordinate-conditioned box regression, followed by Group Relative Policy Optimization with 3D IoU-based rewards to align both stages with downstream grounding quality. On ScanRefer with Qwen3-VL-2B, CoordRefer achieves gains of 11% in Acc@0.25 and 7% in Acc@0.5 over the coordinate-agnostic baseline, while its geometrically refined variant surpasses methods using explicit 3D inputs.
[CV-95] EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
链接: https://arxiv.org/abs/2608.05565
作者: Feier Wu,Wanke Xia,Xu He,Zilang Zhou,Si Chen,Dongxia Liu,Liyang Chen,Qimeng Wu,Zhengbo Zhang,Wenming Yang,Zhiyong Wu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Project: this https URL
Abstract:Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
[CV-96] Hierarchical Flow Matching for 3D Point Cloud Generation
链接: https://arxiv.org/abs/2608.05557
作者: Linhao Wang,Qichang Zhang,Ye Su,Hao Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Generating high-quality 3D point clouds requires capturing both global shape topology and local geometric details. Existing flow-based methods rely on continuous normalizing flows (CNFs) that demand expensive ODE solving and trace estimation during training, while diffusion models require hundreds of iterative denoising steps. Moreover, most approaches adopt single-level generation directly in point space, disregarding the hierarchical structure natural to 3D shapes. We propose Hierarchical Flow Matching (HFM) that extends flow matching to bilevel structure for unconditional 3D point cloud generation. HFM decomposes the task into two levels via optimal-transport flow matching: a \textitLatent Flow Matching models the global shape manifold in a compact latent space, and a \textitConditional Point Flow Matching reconstructs detailed point clouds conditioned on the latent code. Both flows are trained with simple MSE regression losses. The resulting straight OT paths enable efficient sampling with as few as 15 Euler steps per flow, while the structured latent space supports downstream tasks including classification. Extensive experiments on ShapeNet and ModelNet benchmarks demonstrate that HFM achieves competitive or even best performance compared with prior state-of-the-art methods.
[CV-97] OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction
链接: https://arxiv.org/abs/2608.05539
作者: Taiting Lu,Runze Liu,Ziwei Dong,Sisong Bei,Jingying Zeng,Mingjia Wang,Zhenghao Li,Kaiyuan Lin,Yi-Shan Wu,Yangshoudu Zheng,Hongxing Pan,Kai Zhang,Guoliang Shi,Ling Ma,Yifan Yang,Jiaying Lu,Qi He,Sung-Liang Chen,Yi-Chao Chen,Yincheng Jin,Mahanth Gowda
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.
[CV-98] HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
链接: https://arxiv.org/abs/2608.05523
作者: Yuanruyi,Yue Cao,Haojia Gao,Guanqiu Guo,Ziyuezhang,Shangqin,Junbo Tan,Bokui Chen,Zhuo Zou,Xueqian Wang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain accessible when it becomes relevant to a subsequent prediction. Existing approaches mainly enlarge the temporal context, cache generic video features, or impose explicit object-centric states, thereby improving the capacity or structure of retained history. However, they do not directly address how relevant historical evidence can be selectively retrieved and integrated into a pretrained predictor without interfering with its native latent workspace. Accordingly, we introduce HERA (Historical Evidence Routing Adapter), a framework for routing retained historical evidence into a frozen latent predictor, and instantiate it with Register-Routed Patch Memory (RRPM), a lightweight adapter comprising a Structured Memory Bank, Memory Registers, and Workspace Registers. On the IntPhys2 Main split, HERA with RRPM improves the pairwise AvgSurprise accuracy of V-JEPA 2-G from 52.57% to 54.35%. Subgroup analysis shows particularly strong improvements on fixed-camera continuity, from 46.15% to 57.69%, and fixed-camera immutability, from 46.15% to 63.46%. These results support historical evidence routing as a practical adaptation strategy for physical prediction in latent world models.
[CV-99] DynaPix: Can Vision-Language Models Identify the Exact Future?
链接: https://arxiv.org/abs/2608.05505
作者: Thong Nguyen,Vinh-Hien Do,Quynh Vo,Cong-Duy Nguyen,See-Kiong Ng
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Work in progress
Abstract:Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable. Given a video clip that stops before a key event and a question about a later moment, a model must pick the true future image from close candidates or a large gallery. The scenes come from a physics simulator, so the correct image and its time are known exactly and the wrong options are deliberately similar. Models often succeed when a visible event marks the target moment, but are near chance when only elapsed time marks it. Gallery search is harder still, as the true image rarely ranks first. People handle the elapsed-time items well, so the difficulty lies with the models, not the questions. Training on scene accounts drawn from the simulator’s true record, not a teacher’s guess, repairs much of this but not the longer elapsed time case. DynaPix thus exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.
[CV-100] APQF: Agent ic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
链接: https://arxiv.org/abs/2608.05499
作者: Sadegh Jafari,Mohiuddin Bilwal,Fan Zhou,Brian Gelder,Ali Jannesari
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 22 pages, 7 figures
Abstract:Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge devices. Pruning and quantization address this, but rely on manual, expert choices and on algorithms that are hard to apply across architectures. Uniform settings also ignore how differently individual layers respond to compression, which costs accuracy. We introduce APQF, an agentic profiling-guided framework that combines structured pruning, mixed-precision quantization-aware training, and accuracy recovery in one automated pipeline. A profiling agent measures how cost is distributed across the model and how sensitive each part is to pruning, and this evidence drives per-layer pruning ratios, per-layer bit-widths, and the recovery strategy, all proposed by LLM planners and validated before execution. To our knowledge, APQF is the first framework to combine LLM-guided, profiling-grounded decisions with a fully training-aware pruning and quantization pipeline for both CNNs and vision transformers. We evaluate APQF on ResNet, VGG7, ViT, DeiT, and Swin using ImageNet-1k and CIFAR-10. On ImageNet it cuts compute to 5.6-7.7 percent of the original bit-operations, a 13-18x reduction, while keeping accuracy close to the baseline, and under a 200K-image budget it stays roughly 17 points higher in Top-1 than existing joint pruning and quantization methods. On CIFAR-10 it compresses further than that method on four of five architectures. On VGG7 it reaches 93.15 percent using only 0.41 percent of baseline bit-operations, the only method at that compression level to improve on its full-precision baseline. Ablations show that uniform compression loses the most accuracy at matched compute, and that withholding profiling data from the planner hurts every model. Six LLM planners, including free open-weight ones, all reach 97.4-97.9 percent on Swin-Tiny.
[CV-101] VideoArgus: Agent ic Rubric-Grounded Unified Evaluation for Video Generation and Editing
链接: https://arxiv.org/abs/2608.05485
作者: Ziyun Zeng,Zixuan Wang,Yongsheng Yu,Hang Hua,Jiebo Luo
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Preprint
Abstract:Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: this https URL
[CV-102] CDSeg: A Renderable Gaussian Carrier for Image-to-3D Label Transfer
链接: https://arxiv.org/abs/2608.05482
作者: Wentao Sun,Yiping Chen,Zhengsen Xu,Jonathan Li,John S. Zelek
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 15 pages, 8 figures
Abstract:Modern image models provide strong cues about \emphwhat should be segmented in each view, but their masks do not by themselves determine \emphwhere those labels should persist in 3D. We present Cross-Domain Segmentation via Gaussian Splatting (CDSeg), a label-transfer interface that requires no task-specific 3D segmentation training and uses Gaussian primitives as a renderable label carrier. An external mask source supplies the labels, while renderer-derived visibility determines which 3D primitives receive them. The carrier is instantiated either by completing each input point into one Gaussian, preserving its index, or by reusing the native primitives of an optimized Gaussian scene. CDSeg records pixel–primitive associations during rendering and fuses multi-view masks through voting and a local filter. The resulting labels can be returned to the original points, retained on the native Gaussian scene, or rendered into other views. CDSeg covers promptable, automatic instance, semantic, and LiDAR settings and processes scenes with millions of primitives in seconds. It obtains 92.35% mIoU on DesktopObjects-360, 95.89% on NeRDS-360, and 65.77% on the full ScanNet-v2 validation split using the provided 2D semantic annotations. CDSeg thereby provides one interface for reusing 2D masks across point clouds, Gaussian scenes, and image views without a task-specific 3D segmentation network.
[CV-103] MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
链接: https://arxiv.org/abs/2608.05450
作者: Mohammadreza Hami,Mohammadreza Samadi,Chao Gao,Negar Hassanpour
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count makes generation expensive due to the quadratic complexity of self-attention. Several existing efficiency methods reduce this cost by using larger patches at selected denoising steps, thereby representing the image with fewer tokens. Yet, each step still uses a single patch size uniformly across the entire image, overlooking that different regions suffer different fidelity losses when coarsened. We introduce MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps. MOSAIK adapts the PixelDiT backbone to generate arbitrary heterogeneous patch layouts, and a lightweight predictor uses intermediate denoising features to estimate the fidelity loss caused by coarsening each region. Given a token budget, our damage-guided layout predictor assigns fine patches to sensitive regions and coarse patches elsewhere. Remarkably, while reducing FLOPs by 70% and token count by 83%, MOSAIK matches the full-compute PixelDiT on GenEval and its DPG-Bench score drops by only 1.0 point. Compared to diverse efficiency paradigms, including temporal patch scheduling and feature caching, our approach delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.
[CV-104] Invisible Shortcuts: Why Vision Encoders Know Your Camera ECCV2026
链接: https://arxiv.org/abs/2608.05424
作者: Vladan Stojnić,Ryan Ramos,Giorgos Kordopatis-Zilos,Noa Garcia,Giorgos Tolias
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注: ECCV 2026
Abstract:Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning: invisible metadata traces embedded at the pixel level, for metadata such as image processing and photo acquisition. We hypothesize that large-scale semantic supervision, whether through categorical labels (ImageNet) or billion-scale captions (LAION), naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features. By introducing controlled metadata-semantics correlations, we show that stronger ones produce systematically higher sensitivity to metadata traces and larger performance degradation under metadata distribution shifts. We further explore mitigation strategies applied during and after pretraining that reduce sensitivity not only to targeted metadata but also to unseen ones, without sacrificing performance on downstream tasks. Metadata sensitivity also has a positive side: it partly explains the strong generated-image detection ability of some encoders, while its mitigation can improve out-of-distribution generalization. Code: this https URL
[CV-105] Adapting Vision Foundation Models with Cascaded Semantics
链接: https://arxiv.org/abs/2608.05393
作者: Xi Xiao,Xingjian Li,Cheng Han,Tianyang Wang,Lin Zhao,Yunbei Zhang,Guosheng Hu,Runmin Jiang,Xi Li,Xiao Wang,Min Xu
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted by Transactions on Machine Learning Research (TMLR), 2026. Project page: this https URL
Abstract:Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transformers (ViTs) by updating a small set of additional prompt parameters. However, existing visual prompts are randomly initialized and do not exploit prior knowledge, such as instructions in NLP. We address this gap by injecting two complementary semantic priors into VPT. Fundamental image priors, including color, texture, and shape, are extracted with classical hand-crafted operators and injected into the input space, while self-attention maps provide instance-aware semantics in the feature space. We further propose a cascaded scheme that integrates both priors throughout ViT adaptation. Experiments on 34 challenging image classification datasets demonstrate superior downstream adaptation while tuning only 0.74% of ViT parameters. Project page: this https URL.
[CV-106] xt-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model
链接: https://arxiv.org/abs/2608.05389
作者: Zach Eidex,Yu-nong Lin,Mojtaba Safari,Sean Pitroda,Ralph Weichselbaum,Zhen Tian,Xiaofeng Yang
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such as nnU-Net may generalize imperfectly and lack clinician-directed text correction. Purpose: We investigated adapting a three-dimensional (3D) vision-language foundation model for text-guided brain tumor segmentation refinement. Methods: We developed a lightweight VoxTell-based framework. Pretrained VoxTell generated initial masks. Oracle prompts derived from segmentation errors encoded target, action, location, imaging evidence, edit size, and preservation constraints. Frozen Qwen/VoxTell prompt embeddings were injected through trainable projections into its multiscale decoder conditioning; other weights remained frozen. Training, validation, and testing used 901, 100, and 250 BraTS-GLI cases. Cross-dataset transfer was evaluated on 100 meningioma, metastasis, pediatric tumor, and UPENN-GBM cases. Results: On the internal test set using post-contrast T1-weighted input, correct instructions improved subregion Dice similarity coefficient (DSC; enhancing tumor, edema, and necrotic/non-enhancing core) from 0.774\pm0.158 to 0.796\pm0.137 . They outperformed blank prompts ( 0.762\pm0.155 ; Holm-adjusted p0.001 , d_z=0.71 ) and contradictory prompts ( 0.770\pm0.163 ; p0.001 , d_z=0.48 ). In cross-dataset testing, correct instructions improved DSC from 0.527\pm0.287 to 0.550\pm0.278 and outperformed contradictory instructions ( 0.504\pm0.275 ; p0.001 , d_z=0.43 ). Conclusion: A 3D vision-language foundation model can perform instruction-guided refinement of glioma subregion segmentations. Sensitivity to correct, blank, and contradictory prompts suggests text-dependent contour editing rather than nonspecific post-processing, supporting further evaluation as a clinician-in-the-loop tool. Subjects: Computer Vision and Pattern Recognition (cs.CV) Cite as: arXiv:2608.05389 [cs.CV] (or arXiv:2608.05389v1 [cs.CV] for this version) https://doi.org/10.48550/arXiv.2608.05389 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Zach Eidex [view email] [v1] Wed, 5 Aug 2026 20:17:02 UTC (1,853 KB) Full-text links: Access Paper: View a PDF of the paper titled Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model, by Zach Eidex and 5 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.CV prev | next new | recent | 2026-08 Change to browse by: cs References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[CV-107] World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
链接: https://arxiv.org/abs/2608.05369
作者: Yuhao Pan,Haosong Peng,Zhengshen Zhang,Zhengyang Yan,Yalun Dai,Fushuo Huo,Chujie Wang,Tianyu Qi,Xiucheng Wang,Nan Cheng,Wenchao Xu
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.
[CV-108] LoDA: A Level of Detection Aware Method and a Multimodal Sensing Benchmark for Object Level Change Detection ACM-MM2026
链接: https://arxiv.org/abs/2608.05356
作者: Haitian Wang,Xinyu Wang,Sheldon Fung,Xian Zhang,Zichen Geng
类目: Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
备注: 10 pages, 5 figures, 5 tables. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026) Main Track
Abstract:High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in multi-temporal urban LiDAR to keep digital maps aligned with the physical world. Existing approaches from raster height differencing to depth image and point-cloud networks often remain tile-based and threshold-driven, yielding per-point scores without explicit detection limits or consistent object-level labels. We propose an object-level 3D change-detection pipeline that integrates detection-limit-aware registration, geometry-driven object proxies with rule-based semantic and instance segmentation, and displacement cues in height, volume, and surface-normal direction to assign five change labels with confidence. By decoupling registration, geometry, and semantics, the pipeline propagates pose uncertainty into spatially varying detection limits, stabilizes cross-epoch correspondences, and suppresses false changes caused by residual misalignment and density variation. We also present LoDA, a level-of-detection (LoD) aware benchmark for the Subiaco district with fused multi-temporal vehicle-LiDAR maps constructed with LiDAR, GNSS, and IMU support, semantic instances, and object-level annotations. On this benchmark, our method achieves 95.0% accuracy, 90.8% macro F1, and 83.0% macro IoU, exceeding the best baseline by 8.7 IoU points and 4.4 F1 points. On the public Urb3DCD-V2 benchmark evaluated under the official point-wise protocol, it reaches 96.81% mean accuracy and 89.52% mean change IoU, improving over the strongest reported baselines by 1.36 points in mAcc and 3.18 points in mIoUch.
[CV-109] Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
链接: https://arxiv.org/abs/2608.05341
作者: Yuta Kobayashi,Pradyun Ramesh,Muhammad Ahmed Chaudhry,Vincent Jeanselme,Judy Wawira Gichoya,Sanmi Koyejo,Kathleen Capaccione,Shalmali Joshi
类目: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
备注:
Abstract:Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused on monitoring support device placement. As a result, models trained with standard approaches inherit these omissions, learning to under-report findings themselves. We propose PU-DPO, a preference optimization framework to prevent omission noise from corrupting the preference signal. We reformulate the objective under a positive-unlabeled (PU) learning framework, treating absent mentions as unlabeled rather than truly negative. Our framework provides preference supervision using constructed contrastive pairs, generated using edits to model responses, producing variants that explicitly mention or omit a specific finding. Generated responses that mention the finding are naturally preferred in the context of visual evidence. Across semi-synthetic experiments and analyses on real-world chest radiograph benchmarks where adjudicated labels are available, PU-DPO yields consistent gains in detection rates and recovery of hidden positives across multiple pathologies, and is more robust to omission noise than prior approaches.
[CV-110] Context Matters: Support Set Selection and Failure Detection for In-Context Medical Image Segmentation
链接: https://arxiv.org/abs/2608.05333
作者: Youssef Gehad,Emmanuel Zerefa,Krish Kabra,Guha Balakrishnan
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: 10 pages
Abstract:In-context learning (ICL) adapts medical image segmentation models to unseen structures and modalities without retraining by conditioning on a task-specific support set of image-mask exemplars. Because this support set is the model’s only task-specific signal, its composition directly influences segmentation performance. In this work, we investigate the support set as a controllable determinant of ICL reliability. First, we compare random sampling against similarity-based selection, where exemplars are retrieved based on their visual similarity to the query image. Second, we train a transformer-based classifier to predict, from the query and support images alone, whether a segmentation will fall below a specified Intersection-over-Union (IoU) threshold. Using MultiverSeg with DINOv3 embeddings across four benchmarks and three imaging modalities, we show that similarity-based selection consistently matches or outperforms random sampling, with the largest gains at the smallest support set sizes. Furthermore, our classifier predicts segmentation failure above chance on all four benchmarks. Ultimately, these results demonstrate that the reliability of in-context segmentation can be both improved via informed support selection and anticipated before use, providing practical mechanisms for safer clinical deployment.
[CV-111] Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction
链接: https://arxiv.org/abs/2608.05265
作者: Quinn Ledingham,Zhengsen Xu,Yimin Zhu,Zack Dewis,Mabel Heffring,Saeid Taleghanidoozdoozan,Motasem Alkayid,Megan Greenwood,Lincoln Linlin Xu
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas. However, identifying reliable machine learning models is complicated by overlapping debris-flow and non-debris-flow events in feature space, the need for model interpretability, and limited training data. This paper addresses these challenges through a systematic evaluation of machine learning models in terms of predictive performance, feature importance, and synthetic data augmentation. Using basin-scale observations of post-wildfire debris-flow events across the western United States, we compare 15 models, including the Tabular Prior-Data Fitted Network (TabPFN). Repeated stratified cross-validation shows that TabPFN achieves the highest unaugmented performance with a threat score of 0.637, closely followed by the best tree-based models. SHapley Additive exPlanations (SHAP) are used to identify the features driving predictions, revealing that short-duration rainfall intensity and storm accumulation consistently rank highest, while burn severity and terrain features contribute less. We further evaluate synthetic data augmentation using TabPFN-generated samples to address the scarcity of debris-flow observations. Synthetic augmentation improves the performance of all models except CNN, with the largest mean threat score increase of +0.041 among the deep learning models. By combining rigorous model benchmarking, interpretable feature analysis, and synthetic data augmentation, this work provides a comprehensive framework for improving post-wildfire debris-flow prediction.
[CV-112] A Parag raph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval ECCV2026
链接: https://arxiv.org/abs/2608.05260
作者: Mahyar Ghazanfari,Amin Tabrizian,Arsyi Aziz,Binshuai Wang,Peng Wei
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at the MUCG workshop, ECCV 2026
Abstract:Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions. While methods such as Long-CLIP extend the token limit through positional embedding interpolation, we ask a simpler question: does training text granularity alone determine long-text retrieval performance? We present a systematic study of supervision ranging from single captions to multi-sentence paragraphs for contrastive image-text retrieval. Using a synthetic pipeline based on Qwen2-VL and Llama 3.2 Vision, we generate diverse captions, hard negatives, and quality-scored paragraphs for 500K CC3M images. To isolate the effect of text granularity, we fine-tune only the BLIP text encoder while keeping the vision encoder frozen across 10 training configurations. Our paragraph-supervised models match Long-CLIP-L on ShareGPT4V and outperform it by more than 14 points on DOCCI for image-to-text retrieval, without architectural changes. We further show that paragraph supervision enables effective use of long token sequences, whereas caption-only training degrades beyond 60 tokens. Increasing caption diversity improves short-caption retrieval with diminishing returns, while paragraph supervision consistently benefits long-description benchmarks and hard negatives prove detrimental in text-only fine-tuning. Evaluations on Flickr30k, COCO, ShareGPT4V, and DOCCI provide a comprehensive analysis of the trade-offs between text granularity, retrieval direction, and description length.
[CV-113] Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI
链接: https://arxiv.org/abs/2608.05258
作者: Casey Wall,Longwei Wang,Rodrigue Rizk,KC Santosh
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Gradient-weighted Class Activation Mapping (Grad-CAM) is widely used to visualize model decisions, but it was originally formulated for convolutional neural networks, where spatial feature maps and channel dimensions have clear architectural meanings. Vision Transformers (ViTs) do not provide the same structure, instead representing images through tokens, attention, residual streams, and multimodal interactions. This paper presents a systematic taxonomy and literature audit of how Grad-CAM and related methods are adapted, justified, and reported for ViT-based architectures. From an initial search of more than 550 papers, we identify 175 papers that apply Grad-CAM or Grad-CAM-adjacent methods to ViTs. We find that most papers do not provide a full mathematical or implementation-level account of how Grad-CAM is adapted to transformer representations. To characterize this gap, we introduce a descriptive taxonomy of ViT Grad-CAM adaptations that makes explicit the feature locations, gradient targets, spatial reconstruction steps, and aggregation choices that are often left implicit. This taxonomy is not intended to prescribe a single correct adaptation, but to clarify the range of methodological choices being made. The study shows that Grad-CAM on ViTs is often treated as a trivial extension of CNN-based Grad-CAM, despite requiring nontrivial choices that affect rigor, reproducibility, and interpretation.
[CV-114] WorldClaw: Agent ic 3D Open-World Generation at Scale
链接: https://arxiv.org/abs/2608.05248
作者: Chunchao Guo,Jinpeng Li,Yang Li,Zilong Huang
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: Authors are listed in alphabetical order by given name
Abstract:Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.
[CV-115] Disentangling 3D Modeling from Spatial Reasoning
链接: https://arxiv.org/abs/2608.05242
作者: Haoze Sun,Jiequan Cui,Qingshan Xu,Richang Hong
类目: Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:In this work, we explore an alternative paradigm for spatial reasoning by explicitly disentangling 3D perception from reasoning, rather than jointly acquiring implicit 3D perception and reasoning through large-scale training. Our key observation is that modern perception models excel at estimating continuous 3D geometry, whereas large language models (LLMs) are particularly effective at compositional and symbolic reasoning. Motivated by these complementary strengths, we propose the Disentangled Spatial Reasoner (DiSR), a simple yet effective framework that reconstructs the physical world into structured 3D evidence using off-the-shelf expert perception models and fine-tunes an LLM with LoRA to perform reasoning solely over this explicit geometric evidence. Without large-scale 3D VQA training or complex tool-use policies, DiSR achieves competitive performance on popular spatial reasoning benchmarks. Beyond its strong performance, DiSR offers improved interpretability, modularity, and computational efficiency, demonstrating that explicit separation of perception and reasoning is a scalable and effective alternative paradigm to end-to-end modeling for spatial intelligence.
[CV-116] In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
链接: https://arxiv.org/abs/2608.05237
作者: Lingxiao Yang,Liu Liu,Moran Li,Han Feng,Wenjian Cao,Jiangning Zhang,Ye Shi
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注:
Abstract:Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes the model to take shortcuts, resulting in compromised temporal semantics and dynamics. Inspired by the perspective of diffusion as masking, we explore the impact of noisy contexts on few-step autoregressive generation. Yet, simply applying contexts with the same noise levels provides insufficient guidance, leading to poor temporal consistency. To resolve this dilemma, we introduce In-Context Forcing, a progressive autoregressive paradigm that utilizes contexts with decreasing noise levels. By applying less masking to distant frames and more masking to adjacent ones, this approach provides adaptive guidance, effectively ensuring both robust temporal consistency and high inter-frame dynamics. Furthermore, by decoupling the strict dependence on previous clean frames, our paradigm enables cross-frame parallel denoising, achieving substantial inference acceleration without sacrificing performance. Extensive experiments on VBench demonstrate that our method significantly outperforms state-of-the-art approaches in both visual fidelity and inference speed.
[CV-117] Coherence-Oriented Dream Scene Visualisation
链接: https://arxiv.org/abs/2608.05233
作者: Azra Açıl,Simon Colton
类目: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注: short paper accepted at ICCC 2026
Abstract:Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dream descriptions into a temporal sequence of four panel images visualising the dream. This starts with a large language model prompted to split a dream description into four chronological parts. Then a text-to-image model produces images for each part with visual coherence maintained across the sequence, and DSV regenerates any image not suitably matching the text. We evaluate DSV over 50 visualisations from dream descriptions in DreamBank, and report quality, fidelity and coherence results via objective measures employing the CLIP, DINOv2 and Qwen2-VL vision-language models.
[CV-118] NeuroAdaptTrainer: A Fiji/ImageJ Plugin for YOLO-Based Neuron Segmentation InteractiveCorrection and Transfer Learning
链接: https://arxiv.org/abs/2608.05226
作者: Daniela Eraso-Casas,Gerard Villarroya-Pique,Esther Serrano-Pertierra,M. Teresa Fernández-Sánchez,Antonello Novellie,Angel Rio-Alvarez,Víctor M. González
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Neuron counting and segmentation in microscopy images of neuronal cultures is a routine and time-consuming task in neuroscience research, traditionally performed through manual inspection or semi-automatic tools. We present NeuroAdaptTrainer, an open-source Fiji/ImageJ plugin that integrates a YOLO instance-segmentation model directly into the microscopist’s workflow. The plugin allows a user to run automatic neuron detection on a single image or a batch of images, manually correct the resulting detections from within Fiji, and use those corrections to adapt the model to new imaging conditions via transfer learning. A built-in external validation module allows the base and adapted models to be compared quantitatively on a held-out annotated set. NeuroAdaptTrainer lowers the barrier for non-specialist users to benefit from deep-learning-based segmentation while keeping expert supervision at the center of the workflow.
[CV-119] A Survey of Adversarial Efficiency Degradation for Vision Transformer by Exploiting Input-adaptive Optimization
链接: https://arxiv.org/abs/2608.05217
作者: Anadi Goyal,Nandish Chattopadhyay,Anupam Chattopadhyay,Chandan Karfa
类目: Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted for publication at AIIoT 2026
Abstract:Vision Transformers (ViTs) increasingly rely on input-adaptive inference, such as token pruning and early halting, to meet energy and latency budgets. This survey examines a recent class of adversarial efficiency degradation attacks that target these mechanisms to increase computation without necessarily degrading accuracy. We unify and compare two representative attacks, SlowFormer (a universal adversarial patch) and DeSparsify (per-image perturbations), across three popular token-pruning frameworks: A-ViT, ATS, and AdaViT. We standardize reporting using GFLOPs, accuracy loss, and an Attack Success (AS) metric that measures how much of the model’s compute savings the attack takes away. Understanding these attacks is crucial for designing countermeasures that not only mitigate risk but also remain lightweight, since deployment often occurs in low-power settings such as mobile or embedded devices. To organize our analysis, we focus on three questions: how input-adaptive optimizations (e.g., token pruning and early halting) create attack surfaces for efficiency degradation; how such attacks operate in practice and which optimizations are most vulnerable; and which defenses exist today and whether they meaningfully restore efficiency under attack.
[CV-120] VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances IROS2026
链接: https://arxiv.org/abs/2608.05215
作者: Jihoon Oh,Kento Kawaharazuka,Kei Okada
类目: Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
备注: 8 pages, 5 figures. Accepted to IEEE/RSJ IROS 2026. Project page: this https URL
Abstract:Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots makes this challenging. One promising solution is to learn object-centric actionable affordances that are embodiment-agnostic. In this work, we propose a framework that leverages egocentric human videos with state-of-the-art 3D Structure-from-Motion and hand mesh reconstruction to extract actionable affordances such as visual, grasp, and trajectory affordances that explicitly encode where to interact, how to grasp, and how to move. We construct EgoAffordance, a large-scale dataset comprising 204K episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances. Building on this, we introduce VLAff, a large vision-language model-based unified foundation model that learns cross-modal correlations across all actionable affordances. Given a visual observation and instruction, VLAff generates visual affordance heatmaps, grasp poses, and trajectories, which are then converted into directly executable actions by utilizing 3D scene information. Through extensive experiments, we demonstrate that VLAff not only achieves state-of-the-art performance on visual affordance prediction, but can also be effectively applied to real robot applications such as zero-shot manipulation and affordance-guided robot learning.
[CV-121] StyleComposer: Training-Free Multi-Reference Style Composition
链接: https://arxiv.org/abs/2608.05213
作者: Sanghyeok Lee,Jihye Kang,Namhyuk Ahn
类目: Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer them as one style signal, leaving each attribute’s source and strength outside the user’s control. We ask where in a diffusion model one attribute can change while the others hold, and find that no single representation isolates all three. The proposed StyleComposer therefore routes each style attribute through the representation where it separates best and coordinates the routes over denoising time. Without training or inversion, it satisfies three references and the prompt jointly more closely than prior methods, and exposes one strength slider per attribute. Project page: this https URL
[CV-122] Innocent Panels Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
链接: https://arxiv.org/abs/2608.05210
作者: Ye Leng,Junjie Chu,Yiting Qu,Mingjie Li,Yun Shen,Yang Zhang
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 16 pages, 5 figures, 10 tables
Abstract:Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exemplified by the notorious Nazi propaganda picture book \emphDer Giftpilz. Recently, frontier text-to-image (T2I) systems such as Gemini and GPT-Image have enabled conversational generation with consistent characters and scenes across turns, making hateful visual stories, namely ordered image groups that collectively convey hateful narratives, cheap and scalable to produce. Although prior work has studied hateful content generation by T2I systems, it focuses on individual images, leaving group-level hateful meaning largely unexplored. We aim to address the gap. Concretely, we introduce \textttHatefulStoryPrompts, comprising 330 multi-turn configurations from 55 hateful stories across two languages and three visual styles, and evaluate five frontier models over 4,950 attempts. Every model completes over 80% of the stories, with the strongest reaching 99.0%. We further evaluate existing moderation systems on \textttHatefulVisualStory, a human-labeled dataset of 969 hateful image sets and 990 benign controls, and find that they frequently miss group-level hateful meaning: dedicated safety models achieve at most 34.9% recall, while a strong vision-language model reaches 67.5%. Finally, we propose complementary proactive and post-generation defenses. An interaction-aware monitor achieves 97.3% recall for prompt-only sessions and 92.6% when the user supplies the first image, while post-generation methods jointly analyzing completed image groups reach 80.2%. Our work shows that, as image generation evolves from isolated outputs to coherent visual narratives, safety must evolve accordingly, from per-image moderation to stateful reasoning over interactions and image relationships.
[CV-123] MapTCL: Temporal Consistency Learning via Bidirectional Alignment for Vectorized HD Map Construction IROS
链接: https://arxiv.org/abs/2608.05209
作者: Hyeonseo Kim,Juyeb Shin,Hyeonjun Jeong,Hiwon Shin,Dongsuk Kum
类目: Computer Vision and Pattern Recognition (cs.CV)
备注: Accepted at 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Abstract:Constructing reliable online HD maps remains challenging in dynamic urban environments due to moving objects and occlusions. While recent works employ feature-level temporal fusion to address this, they rely solely on per-frame ground truth supervision. Consequently, they lack an explicit objective to directly penalize the geometric noise and temporal jitter between consecutive online HD maps. To address this, we propose MapTCL, an auxiliary training strategy that formulates temporal consistency loss between current and past frames via bidirectional alignment. Specifically, Bidirectional Vector Consistency Learning (BVCL) models the geometric and semantic discrepancies between associated past and current vector instances as an auxiliary loss. We also employ Raster map Consistency Learning (RCL) as an additional loss to stabilize dense BEV features. By jointly training with these dual losses, MapTCL improves the temporal stability of generated HD maps. Extensive experiments on two standard benchmarks demonstrate the effectiveness of our approach. As a versatile plug-and-play module, MapTCL consistently enhances existing baseline models, achieving gains of +3.7 mAP +2.8 C-mAP on nuScenes and +3.1 mAP +2.5 C-mAP on Argoverse 2 without additional inference overhead.
[CV-124] HoloCount: A Holistic Visual Counting Benchmark for MLLM s
链接: https://arxiv.org/abs/2607.06420
作者: Jinhong Deng,Limeng Qiao,Guanglu Wan
类目: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
备注: Technical report
Abstract:Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by persistent numerical hallucinations. Existing counting benchmarks primarily focus on basic perception in simplified contexts, failing to capture the complex failure modes that emerge under logical constraints or adversarial conditions. To address these limitations, we introduce HoloCount, a holistic and diagnostically rich benchmark structured around a three-level hierarchical taxonomy. HoloCount evaluates MLLMs across: (1) Semantic Counting, focusing on atomic and property-based enumeration; (2) Analytical Counting, assessing logical composition through spatial and set-based reasoning; and (3) Robustness Testing, probing model integrity against adverse scenarios and grounded counter-priors, such as high-density scenes and linguistic biases. Through an exhaustive evaluation of over 20 state-of-the-art MLLMs, we reveal a critical performance gap: even top-tier models degrade significantly as tasks transition from perception to complex analytical reasoning and adverse scenarios. Our findings provide a systematic landscape of current MLLM counting capabilities and offer a roadmap for developing more grounded and reliable multimodal systems. The dataset is available at this https URL.
[CV-125] SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
链接: https://arxiv.org/abs/2608.05669
作者: Hao Si,Zehua Chen,Qingquan Yang,Xiao Wang,Dengdi Sun,Wanli Lyu,Gaoting Chen,Guosheng Xu,Hang Su,Jin Tang,Jun Zhu
类目: Plasma Physics (physics.plasm-ph); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
备注:
Abstract:Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-confinement fusion devices, while conventional infrared-based inversion is usually performed after discharge and requires heat-conduction modeling with device-specific material properties, divertor geometry, and boundary conditions. Rather than accelerating this conventional infrared-based inversion paradigm, we introduce a new online-oriented signal-based reconstruction paradigm that directly reconstructs time-resolved radial heat-flux profiles from multi-source macroscopic plasma-state signals available during discharge. To enable systematic study of this task, we construct \textbfDivMPS2HF, a multi-source discharge dataset that provides the data foundation and benchmark for signal-based divertor heat-flux reconstruction. We further propose \textbfSafeDivertor, a task-driven framework designed to address the key challenges of signal-based heat-flux reconstruction. It employs physical prior-aware initialization to provide radial-distribution guidance for target channels, input perturbation to reduce over-reliance on specific heterogeneous signals, spectral-aware reconstruction optimization to exploit time-frequency priors and preserve transient dynamics, and progressive training to stabilize the optimization of these complementary objectives. Experiments on DivMPS2HF demonstrate that SafeDivertor achieves the best overall performance among the evaluated time-series baselines across all five metrics, establishing a new performance benchmark for signal-based divertor heat-flux reconstruction. The source code will be released on this https URL
[CV-126] A Foundational EDM2-Based Generative Model for High-Resolution Synthetic Fetal Ultrasound Imaging from Open Datasets
链接: https://arxiv.org/abs/2608.05471
作者: Harvey Mannering,Yilin Zhang,Ziao Liu,Zhiwu Huang,Jacqueline Matthew,Miguel Xochicale
类目: Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Medical Physics (physics.med-ph)
备注: Short paper for the 30th Conference on Medical Image Understanding and Analysis
Abstract:Prenatal ultrasound imaging is key for assessing fetal health, but AI progress is limited by scarce, privacy-restricted, and hard-to-annotate datasets. We propose a high-resolution fetal ultrasound synthesis framework based on the EDM2 diffusion architecture, trained on multiple public datasets to generate 512x512 images across six anatomical classes. Our method achieved improved image quality with lower FID scores and enhanced downstream fetal plane classification, reaching 93.36% ensemble accuracy after fine-tuning, surpassing real-data-only training. Clinical evaluation by an experienced fetal ultrasound specialist (10+ years) on 100 images yielded a mean realism score of 2.67/5, with real images rated higher than synthetic. Artefacts included smoothing, speckle irregularities, and anatomical inconsistencies. Code, data, models and other resources to reproduce this work are available at this https URL.
人工智能
[AI-0] racing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
链接: https://arxiv.org/abs/2608.06366
作者: Soorya Ram Shimgekar,Michelle Hu,Dorisa Shehi,Daniel Kang,Roy Ka-Wei Lee,Koustuv Saha,Christian Poellabauer,Christopher Lee,Sajeev Singh,Piyum Zonooz,Navin Kumar,Zeeshan Ahmed,Priyadarshini Kachroo
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists’ workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data with disease-specific, guideline-based clinical reasoning. Existing rule-based and large language model (LLM)-based approaches offer only partial automation with limited maintainability and evidence traceability. We developed the Nimblemind Multi-Agent System (nMAS), an evidence-linked, rubric-grounded pipeline for automated heart-failure feature engineering, and evaluated it on 500 dummy patient records from nine EHR source tables. nMAS generated 132 structured and 70 rubric-scored aggregated features, verified for structural integrity, rubric compliance, and provenance, and audited by a restricted LLM. Adding the aggregated features improved held-out AUROC from 0.895 to 0.963 for HFrEF and 0.870 to 0.910 for HFpEF phenotyping, and an independent LLM-based rubric assessment of evidence support and methodological soundness scored the features at 81.5% of maximum points. These results demonstrate the feasibility of automated, auditable feature engineering for complex cardiovascular EHR data, though evaluation was limited to a single-institution cohort and external validation is needed.
[AI-1] Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
链接: https://arxiv.org/abs/2608.06364
作者: George Grispos,Sajda Qureshi
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注: Paper presented at Thirty-second Americas Conference on Information Systems, Reno, USA
Abstract:The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applications affects digital sovereignty, examined through platform transparency as a key indicator of user awareness and control. Using an interpretive approach, the research combines the forensic analysis of selected Android applications with contextual document analysis to identify AI features and evaluate disclosure practices. The findings show that AI is widely implemented in the applications, yet transparency about its use remains limited. A socio-economic analysis of Nigeria further shows an increasing dependence on consumer digital platforms, moderate AI awareness, and uneven patterns of interaction. By providing empirical evidence on AI transparency and platform practices, this study advances understanding of individual digital sovereignty and highlights challenges for protecting user control in AI-driven digital environments.
[AI-2] An Optimal Agnostic PAC Algorithm
链接: https://arxiv.org/abs/2608.06363
作者: Markus Engelund Mathiasen,Jian Qian,Nikita Zhivotovskiy
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST)
备注: 18 pages
Abstract:Let H\subseteq-1,+1^X be a class of finite VC dimension d\ge1 . Writing L for the binary risk and L^=\min_h\in HL(h) , we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size n , for every 0\delta\le 1/2 , with probability at least 1-\delta , [ L(\widehat h) \le L^+ 7\cdot10^8\left( \sqrt\fracL^(d+\log(1/\delta))n +\fracd+\log(1/\delta)n \right). ] This settles the sample complexity of agnostic PAC learning up to universal constants at every fixed L^ , matching the lower bounds of Devroye, Györfi, and Lugosi [A Probabilistic Theory of Pattern Recognition, Springer, 1996]. Comments: 18 pages Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Data Structures and Algorithms (cs.DS); Statistics Theory (math.ST) Cite as: arXiv:2608.06363 [cs.LG] (or arXiv:2608.06363v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.06363 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Nikita Zhivotovskiy [view email] [v1] Thu, 6 Aug 2026 17:57:25 UTC (20 KB) Full-text links: Access Paper: View a PDF of the paper titled An Optimal Agnostic PAC Algorithm, by Markus Engelund Mathiasen and 2 other authorsView PDFHTML (experimental)TeX Source view license Current browse context: cs.LG prev | next new | recent | 2026-08 Change to browse by: cs cs.AI cs.DS math math.ST stat stat.TH References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) IArxiv recommender toggle IArxiv Recommender (What is IArxiv?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[AI-3] he Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
链接: https://arxiv.org/abs/2608.06361
作者: Sarvesh Baskar,Zikui Cai,Shayan Shabihi,Anirudh Satheesh,Muhammad R. Islam,Udari Madhushani Sehwag,Tom Goldstein,Furong Huang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence – a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.
[AI-4] Challenges in Evaluating Explanation Methods for Static and Evolving Data IJCAI ECAI2026
链接: https://arxiv.org/abs/2608.06351
作者: Jerzy Stefanowski
类目: Artificial Intelligence (cs.AI)
备注: 13 pages, 1 figure = this paper is a preprint of the workshop [Explainable AI in Space] paper for IJCAI ECAI 2026 conference
Abstract:This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for explaining image classification is presented. The paper further explores methods for adapting explanations to evolving data streams with concept drift. Experiences with adapting counterfactuals for this problem are discussed. Finally it is related to the challenges of tracking the co-evolution of data, models, and explanations.\footnoteThis paper has been accepted for a publication in this http URL (ed) Explainable AI in Space. Proceedings of EASi 2026 Workshop at IJCAI-ECAI 2026 Bremen, Springer CCIS vol 3107 (2016).
[AI-5] RAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
链接: https://arxiv.org/abs/2608.06346
作者: Yunjia Qi,Zehua Yin,Xintong Shi,Hao Peng,Songyuanyi Lu,Yixian Liu,Richeng Xuan,Yuhong Liu,Zhichao Hu,Xiaozhi Wang,Lei Hou,Bin Xu,Juanzi Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress faces two main challenges. First, long trajectories make it difficult to identify individual errors, since the evidence for judging a step may be scattered across distant instructions, observations, and prior context. Second, failed trajectories often contain multiple local errors with different downstream effects, only some of which remain responsible for the final failure. In this work, we propose TrajDebug, an error-lifecycle tracing framework that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error’s resolution status and terminal impact. We further construct TrajErrBench, a benchmark of 486 manually annotated failed trajectories from Tau2Bench and SWE-Bench Pro, covering realistic tool-use and coding scenarios. Experiments across diverse agent benchmarks show that TrajDebug achieves the best overall performance over existing baselines, and application studies further demonstrate that its diagnoses provide actionable feedback for improving downstream agent success. We will release the codes and data to facilitate further research.
[AI-6] ytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data
链接: https://arxiv.org/abs/2608.06331
作者: Donna Hooshmand,Shubham Shahi,Cameron Barrie,Abhratanu Dutta,Marko Sterbentz,Harper Pack,Kristian J. Hammond
类目: Databases (cs.DB); Artificial Intelligence (cs.AI)
备注: 20 pages, 4 figures, 6 tables
Abstract:From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually written by hand. This is a knowledge-acquisition bottleneck that limits the scalability of analytic systems, keeps non-technical users dependent on experts, and is itself error-prone. We present TYTAN, a system for automatically constructing an analytic semantic schema from a relational database and, when available, a short user-provided description. TYTAN combines symbolic analysis of the database with LLM-based semantic inference for entity proposal, role assignment, and naming. When the evidence leaves a decision ambiguous, TYTAN asks the user a targeted natural-language question. We evaluate TYTAN on eight databases spanning real-world and benchmark domains along the three axes that define a schema’s functional utility: (i) coverage, are all important entities and features captured?; (ii) retrieval correctness, do the schema’s instructions actually reach the data; and (iii) characterization accuracy, are semantic types correct? Across the seven reference domains, TYTAN reaches every entity, attribute, and aggregable feature of the expert-corrected reference schemas (100% coverage). Additionally, 100% of its retrieval instructions execute correctly (1,678 of 1,678 self-generated claims), and semantic roles agree with the reference on 92-100% of matched attributes. Checking the underlying data showed the small disagreement is in the reference, not in TYTAN. On a held-out blind test (a live, ten-table database with no declared keys), TYTAN recovers the full entity structure with verified keys and satisfies 100% of the satisfiable expectations of five independent blind annotators.
[AI-7] Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
链接: https://arxiv.org/abs/2608.06300
作者: Arya Labroo,Mengjie Qian,Kate Knill
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners’ speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts’) in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model’s activation space, allowing us to distinguish between whether a concept is encoded in a model’s internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.
[AI-8] QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agent ic AI for Cardiac Arrest Mortality Prediction
链接: https://arxiv.org/abs/2608.06294
作者: Mutasim Fuad Sarker,Adiba Rahman Namira,Wafa Binte Alam,Md Adnan Arefeen,Mahzabeen Emu,Sumaiya Tabassum Nimi
类目: Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET)
备注: Submitted for review
Abstract:Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches ignore the temporal progression of physiological deterioration and recovery that unfolds throughout a patient’s ICU stay. To address this limitation, we introduce QuanTiMedAI, a quantum-agentic framework developed for cardiac arrest mortality prediction using agentic AI guided quantum enhancement time series model. The proposed system combines an agentic large language model (LLM) for clinically informed feature discovery with a compact quantum recurrent network for temporality aware mortality prediction. Our findings demonstrate that agentic LLM-guided feature selection consistently outperforms conventional feature selection approaches, and the proposed quantum architecture achieves competitive predictive performance through nonlinear feature enhancement while keeping the number of parameters very low. Through extensive experimentation on a MIMIC-IV cohort of cardiac arrest patients, QuanTiMedAI’s quantum-enhanced architecture attains an AUROC of 0.852 using only 605 parameters, an improvement of approximately 2.9% over a current state-of-the-art baseline for this task. A structured ablation study systematically validates the contribution of each architectural design choice. These results show that quantum-enhanced sequential modeling can exceed classical recurrent networks while using substantially fewer parameters.
[AI-9] BaKron: Efficient Quantization with Kronecker-Factored Hessians
链接: https://arxiv.org/abs/2608.06291
作者: Johann Birnick,Rayan Saab
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian. GPTQ-style adaptive rounding typically uses one-sided information derived from input activations. Two-sided Kronecker-factored Hessian approximations can additionally capture correlations across output coordinates, but applying GPTQ directly in the vectorized weight domain is computationally expensive. Building on the two-sided adaptive-rounding formulation used by BoA and YAQA, we introduce BaKron, an efficient solver that combines anti-diagonal parallelism with a recursive divide-and-conquer construction. For an m\times n weight matrix, BaKron uses O(m+n) sequential steps while reducing the total work from O(m^2n^2) to O(mn(m+n)) . Thus, it matches the cubic scaling of GPTQ while exploiting richer curvature information. Moreover, BaKron is modular with respect to both the base quantizer and the Hessian estimator. We also provide practical benchmarks, consider a range of Hessians that BaKron can be called with, find an efficient technique to compute these Hessians, and evaluate the algorithm experimentally.
[AI-10] he Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
链接: https://arxiv.org/abs/2608.06270
作者: Zhiheng Wang,Bo Peng,Lai Wei,Chaochao Lu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:The “thinking-with-images” paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a causal graph that separates observation-mediated paths from action-induced shortcuts. We then audit it through interventions at the three levels: policy (comparing tool-use with direct inference), trajectory (corrupting all observations during rollout), and step (counterfactually replacing one individual observation under a fixed prefix). Our step-level estimand, Visual Evidence Gain, isolates the contribution of each returned observation. Across six representative models and five fine-grained perception benchmarks, we uncover policy miscalibration with two failure modes. In Calling Without Looking, returned observations have no causal effect on the answer. In Looking Without Planning, observations are informative but the call schedule is incoherent. A trajectory-level diagnostic decomposes the policy-level accuracy gain and shows that the gain is concentrated in a Calibrated minority. We term this discrepancy the illusion of visual tool-use: despite aggregate accuracy gains, visual tool-use is not causally effective across a broad range of rollouts. The code is available at this https URL.
[AI-11] Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
链接: https://arxiv.org/abs/2608.06265
作者: Omid Bazgir,Md Nasir,Jacob Hoffman,Yang Yang,Manu Agrawal,Anusua Trivedi,Vinay Rao Dandin,Chris Gibbons,Christine Swisher
类目: Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG)
备注:
Abstract:Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism. Subjects: Artificial Intelligence (cs.AI); Databases (cs.DB); Machine Learning (cs.LG) Cite as: arXiv:2608.06265 [cs.AI] (or arXiv:2608.06265v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.06265 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-12] DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
链接: https://arxiv.org/abs/2608.06243
作者: ZhiYan Hou,Xinyu Tang,Hongyan An,Jianjin Zhang,Weizhen Wang,Yunyun Han,Gengsheng Li,Xiangzhao Hao,Haiyun Guo,Wenbin Hu,Jinqiao Wang,Yafeng Deng
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 4 figures, 9 tables. Code at this https URL
Abstract:Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals are typically sparse and at the sequence-level. On-policy self-distillation (OPSD) mitigates this sparsity by querying a privileged teacher at student-visited prefixes and providing dense token-level distributional supervision. Although this dense supervision alleviates signal sparsity, we find that standard OPSD still underexploits the temporal structure of the rollout. It assigns every local divergence the same coefficient, regardless of its position or the divergence sequence in which it occurs. In on-policy autoregressive generation, the same divergence magnitude can follow different discrepancy histories, reflecting different evolutions of the mismatch between the teacher and student. Since the local scalar alone cannot distinguish these temporal contexts, standard OPSD cannot adapt its token-level weights to the realized discrepancy sequence. To address this limitation, we propose Divergence-Adaptive Supervision Horizons (DASH). DASH maps the gap between each local distillation signal and the sequence-level mean to an adaptive propagation gate and then uses these gates to control backward multi-step aggregation. By doing so, DASH adjusts token-level supervision weights according to how local divergences evolve during generation. Experiments on three mathematical reasoning benchmarks across three model scales show that DASH improves over our matched vanilla OPSD reruns on every benchmark at all three scales. DASH reuses the teacher and student distributions that OPSD already computes, so the gains require no additional teacher or student forward pass. Code: this https URL Comments: 17 pages, 4 figures, 9 tables. Code at this https URL Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.06243 [cs.AI] (or arXiv:2608.06243v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.06243 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-13] From Passive Mirrors to Active Agents : Holonic Digital Twins for Physical AI over Networks
链接: https://arxiv.org/abs/2608.06227
作者: Christo Kurisummoottil Thomas,Omar Hashash,Walid Saad
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI); Information Theory (cs.IT); Systems and Control (eess.SY)
备注:
Abstract:Despite advances in artificial intelligence (AI) across multiple sectors, today’s AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots and vehicles operating under real-world physical laws. This stems from their inability to maintain reliable world models for long-horizon planning under uncertainty and generalize to unseen scenarios. In this context, wireless networks, through pervasive sensing and communication, can orchestrate physical intelligence. However, current architectures optimize throughput, latency, and reliability and cannot support real-time physical AI coordination, requiring agents to maintain shared spatiotemporal context. To address these challenges, a network of holonic digital twins (HDT-Nets) framework is proposed to deliver real-time physical AI inference through holonic agents that actively reason about their environment rather than passively mirror physical assets. Each HDT is realized as a hierarchical structure spanning the physical agent and network edge, reasoning autonomously at the local level while cooperating with neighboring HDTs to form collectively intelligent units. In HDT-Net, causal Markov blankets spanning sensing, communication, and control determine which agents must coordinate and enable counterfactual reasoning over multi-domain interventions. Active inference within these boundaries unifies perception, action, and learning by minimizing expected free energy while deciding which beliefs to transmit based on their cognitive value to the receiver. Category theory ensures that transmitted beliefs preserve semantic structure across heterogeneous agents with incompatible representations. Finally, integrated information theory quantifies when collective intelligence exceeds independent operation and how network intelligence evolves through coordinated learning and information exchange.
[AI-14] S-RAG : Retrieval Augmented Generation for Time Series Forecasting
链接: https://arxiv.org/abs/2608.06223
作者: Yixiong Xiao,Congxi Xiao,Jingbo Zhou
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the application of retrieval-augmented generation (RAG) in this domain remains limited. Since RAG has proven effective in enhancing the capabilities of large language models by incorporating relevant external information, retrieving similar time series sequences as references might also improve accuracy in time series forecasting tasks. However, most time series models are constrained by limited training data, smaller parameter scales, and a lack of the extensive generative capabilities found in large language models. Simply concatenating reference sequences into the prompt, as done in language models, may not yield the expected results. To address these challenges, we propose a novel approach, TS-RAG, which leverages RAG to enhance forecasting performance. The framework introduces specially designed reference tokens to effectively fuse information from the input sequence with that from retrieved similar sequences, enabling a more robust capture of complex temporal dynamics. Experimental results demonstrate that TS-RAG achieves consistent state-of-the-art performance across several real-world forecasting benchmarks.
[AI-15] Continual Learning in Transition
链接: https://arxiv.org/abs/2608.06216
作者: Zhiyan Hou,Dan Zhang,Tao Feng,Liyuan Wang,Wei Li,Xiangzhao Hao,Hongyan An,Junfeng Fang,Haokai Ma,Zhaohui Xu,Haiyun Guo,Jinqiao Wang,Tat-Seng Chua
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 23 pages, 6 figures, 1 table. Survey on continual learning in the LLM and agentic-AI era
Abstract:Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging paradigms are reshaping the scope of CL beyond this traditional model adaptation view. For instance, on-policy learning broadens the space of update mechanisms; test-time training extends CL from the training phase to inference; and external harness components such as memory, skill libraries, and interaction protocols extend the evolutionary boundaries of model capabilities far beyond the static parameter space. Collectively, these developments indicate a transition from parameter-centric learning toward system-level adaptation. To characterize this transition, we examine the evolution of continual learning through three dimensions: When, How, and Where learning occurs. The How dimension encompasses off-policy, on-policy, and beyond-gradient optimization mechanics. The When dimension captures evolution across pre-training, post-training, and inference-time stages. The Where dimension delineates updates occurring within internal parameters versus external structural constraints. Anchored by this tri-axial framework, we systematically survey representative methods, trace the ongoing transition of continual learning, and discuss the key challenges, broader implications, and future directions arising from this paradigm shift.
[AI-16] EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2608.06197
作者: Zishan Xu,Zhiyuan Yao,Yuxin Chen,Yifu Guo,Zhengxi Lu,Yuquan Lu,Jinyang Huang,Yan Xu,Yasheng Wang,Weinan Zhang,Xingshan Zeng,Weiwen Liu
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at this https URL.
[AI-17] Comparative Approaches to Agent Retrieval over Large Skill Libraries
链接: https://arxiv.org/abs/2608.06196
作者: Indivara Kolluru,Nathan Sportsman
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 6 figures
Abstract:Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expensive and provides no structure for autonomous sequencing. We study two systems for this problem over a corpus of 690 skills: a hybrid ranker combining lexical and dense-embedding retrieval for sparse, on-demand loading, and a typed knowledge graph encoding workflow relations such as prerequisites, data flow, and ordering. On a set of 117 realistic, non-echoing queries, the hybrid ranker retrieves the correct skill within the top five in 73.5% +/- 8.0 of cases, leaving roughly a quarter of queries unserved. When used as the design intended (substituting graph neighbours for additional ranked results at matched token budget), the graph is significantly worse (-11.2 points, p = 0.0007). Its LLM-generated edge layer adds nothing over neighbours obtained free from a local embedding pass, and 73% of the queries the ranker misses are not reachable through the graph at all. We attribute this to a pre-filter topology bound. Because the graph’s candidate edges are drawn from the same embedding neighbourhood the ranker already searches, 98.6% of typed edges connect skills the ranker had already surfaced together. The graph can enrich relation semantics but cannot extend retrieval reach. We further show that evaluating on author-written queries overstates hit@5 by up to 44 points, which would have hidden these results entirely. Our contribution is a mechanistic account of why added structure does not improve retrieval over a strong ranker, and identify the conditions under which adding structural interdependence into the retrieval is optimal.
[AI-18] MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
链接: https://arxiv.org/abs/2608.06183
作者: Jia Xiong,Runkai Li,Chenxu Niu,Guangyuan Gao,Changwen Xing,Yifan Zhang,Xinlai Wan,Jieran Cui,Chen Bai,Yusheng Hua,Ying Wang,Ming Ling,Xi Wang,Tao Xie
类目: Artificial Intelligence (cs.AI)
备注: Accepted by ICCAD 2026
Abstract:Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulation budget for design decision-making. Existing methods perform blind search without considering microarchitectural dependencies and fail to learn from the iterative search effectively, leading to wasted evaluations and weak Pareto convergence. In this paper, we propose MicroEvo, a knowledge-guided framework that couples off-the-shelf LLMs with Monte Carlo Tree Search (MCTS) for multi-objective microarchitecture optimization. MicroEvo combines LLM-driven evolutionary operators, a Pareto-aware tree policy that balances Pareto contribution and diversity, an active knowledge accumulation mechanism that extracts and reuses optimization insights, and state-aware directives that adapt the search behavior online. Experiments show that MicroEvo improves Pareto-front quality by up to 36.2% over NSGA-II and achieves 10.6x higher search efficiency, and also demonstrates strong scalability to a complex industrial-scale core. The code repository is available at: this https URL.
[AI-19] Audio-to-Score Transcription using Pre-trained Features Data Augmentation and the New SheetSage-A2S Dataset
链接: https://arxiv.org/abs/2608.06165
作者: Eoin Cummins,Zhongyi Huang,Alexandre D’Hooge,Zhuoro Mo,Yaolong Ju
类目: ound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
备注: Accepted at the 34th ACM International Conference on Multimedia (MM '26)
Abstract:Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt**kern score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3% SER from the existing state-of-the-art \citealfaro-contrerasTransformer2024. Additionally, our model achieves 20.92% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: this https URL.
[AI-20] ARCS: Iterative Agent ic RL for Controllable 3D Scene Generation
链接: https://arxiv.org/abs/2608.06161
作者: Saugat Adhikari,Ashok Prasad Neupane,Pramish Paudel,Ajad Chhatkuli,Danda Pani Paudel
类目: Artificial Intelligence (cs.AI)
备注: 15 pages, 9 figures, 4 tables. Includes appendix
Abstract:Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints. This mismatch limits the usefulness of synthetic data for downstream training, where accessibility, traversability, and spatial rule compliance are often essential. We present iARCS, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to natural-language task requirements. iARCS uses a two-stage strategy: universal-reward pretraining to improve physical plausibility and layout quality, followed by task-specific fine-tuning with LLM-generated reward programs that are iteratively refined from training feedback. Experiments show improved constraint fidelity on walkability, reachability, and clearance-focused tasks, effective task-specific constraint optimization, and competitive scene diversity. We further show that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable scene editing method.
[AI-21] Learning Globally Reusable Skills for Coding Agents
链接: https://arxiv.org/abs/2608.06153
作者: Chen Yang,Jiashuo Tian,Ziqi Wang,Xinyin Liu,Meiru Ye,Junjie Chen
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing approaches typically treat skill evolution as a sequence of local updates, overlooking relationships among skills and often producing overfitted skill updates that fail to generalize across tasks. We propose GSE, a globalized skill evolution framework that jointly optimizes skill compatibility and skill generalization. To preserve consistency across the skill bank, GSE maintains a Skill Relation Graph (SRG) that explicitly models and co-evolves inter-skill relationships. To improve generalization, GSE performs cluster-based skill consolidation to abstract reusable capabilities from local updates and employs replay-driven verification to prevent overfitting and behavioral regressions. We evaluate GSE on two representative software engineering tasks: bug-revealing test generation and false-positive bug report filtering. Across two state-of-the-art coding agents, OpenHands and mini-SWE-agent, GSE consistently achieves the best precision, recall, and F1-score. Compared with existing evolution techniques, GSE improves precision and recall by 6.1%~34.1% and 31.8%~180.0% for test generation, and by 15.4%~96.4% and 13.1%~19.8% for false-positive filtering. Deployment on an internal industrial agent further yields a 61.4% improvement in F1-score, demonstrating the effectiveness and generalizability of GSE for evolving effective skills.
[AI-22] PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
链接: https://arxiv.org/abs/2608.06146
作者: Hao Yu,Jiabo Zhan,Kang Liu,Linnan Zhao,Dongxu Yue,Rui Chen,Jinglin Wang,Chong Sun,Chen Li,Jing Lyu,Chun Yuan
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at this https URL
[AI-23] FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
链接: https://arxiv.org/abs/2608.06144
作者: Bo Deng(1 and 2),Kang Zhou(2),Lifan Guo(2),Chongyang Tao(1),Xuanren Chen(1),Chenggang Xie(1),Renzhao Liang(1),Feng Chen(2),Chi Zhang(2) ((1) Beihang University, (2) Qwen DianJin Team, Alibaba Cloud Computing)
类目: Artificial Intelligence (cs.AI)
备注: 22 pages, 4 figures; includes appendices
Abstract:Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing self-evolution benchmarks do not jointly cover professional workflows, open-ended deliverables, and multi-aspect evaluation. We introduce FinEvo-Bench, a longitudinal benchmark with 120 real-case-grounded tasks, 20 business scenes across six financial domains. Institution-provided professional procedures define the required operations and constraints. Eligible institution-provided and publicly documented cases supply the task facts. Each scene contains six related but substantively distinct cases that share a professional procedure and a manually reviewed rubric for task quality and financial compliance. We compare four self-evolving agent scaffolds using the same Qwen3.7-Max backbone and three independently shuffled, globally interleaved task streams. Paired non-evolving controls estimate each scaffold’s self-evolution gain from retained experience, while an independent Claude Code scoring agent backed by Claude Opus 4.6 evaluates all outputs. Letta achieves the highest evolved score (91.65) and fewest compliance issues (0.09 per task); Codex achieves the largest self-evolution gain (+19.37). Across scaffolds, the evolving condition raises scores by 9.33-19.37 points and reduces compliance issues by 0.12-0.44 per task. Paired score gains at within-scene ranks 4-6 exceed those at ranks 1-3 by 6.10-8.70 points. In Claude Code, skill-only evolution produces higher task quality and fewer compliance issues than memory-only and combined memory-skill evolution. Across all four scaffolds, rubric feedback also yields higher scores and fewer compliance issues than reference-answer feedback. FinEvo-Bench measures both professional performance and self-evolution ability: how effectively an agent turns prior experience into later improvement.
[AI-24] Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture
链接: https://arxiv.org/abs/2608.06130
作者: Leo Sambrook,Sampo Sovio
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 11 pages, 2 figures. Accompanying code and artifacts available at: this https URL
Abstract:AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext files, environment variables, or container memory. Any process with sufficient read privileges can extract the raw key material. A recent production incident demonstrated the practical severity: private keys were exfiltrated from a widely deployed framework via email injection in under five minutes. We aim to enforce both key confidentiality and content-aware authorisation for key use. To that end, we replace software-resident keys with hardware-confined keys accessible through a vendor-neutral PKCS#11 interface. A hardware keystore (HSM, TPM, smart card) executes cryptographic operations on-device; the host receives only the result via opaque handles. Hardware confinement is the primary contribution; it is enabled by a surrounding five-layer Zero-Trust enforcement stack comprising session identity (SAGA), scope bounds (Smax), semantic validation (RAV), taint tracking, and the hardware execution boundary. We evaluate against 12 injection scenarios derived from AgentDojo’s ImportantInstructionsAttack template (Debenedetti et al., arXiv:2406.13352). We run four LLM models; three follow injections in baseline mode (gpt-oss-120b, Qwen2.5-72B, DeepSeek-V4-Flash, n=192 combined). Baseline Attack Success Rate (ASR): 19.3% [14.3%, 25.4%]; protected ASR: 0% (Wilson 95% CI upper bound 2.0%). Zero false positives across four benign task scenarios.
[AI-25] Contextual Information Policy Optimization for Search Agents
链接: https://arxiv.org/abs/2608.06128
作者: Xingyu Guo,Wei Chen,Linlin Yang,Baochang Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it to guide subsequent reasoning. However, existing methods primarily reward final-answer cor rectness or intermediate progress, without directly assessing whether post-retrieval actions are grounded in the retrieved evidence. This misalignment encourages prior-driven reason ing: agents form conclusions based on internal knowledge and use retrieval mainly to confirm them, resulting in confirma tion bias and inefficient this http URL, we propose Contextual Information Policy Optimization (CIPO), an evidence-oriented reinforcement learning framework that explicitly aligns policy optimization with external evidence use. CIPO assigns dense, turn-level credit to reasoning ac tions influenced by retrieved information, while combining this evidence-use signal with a global outcome reward to pre this http URL,CIPOdiscourages evidence-detached guesses and promotes reasoning trajecto ries in which retrieved facts can guide or revise subsequent reasoning. Importantly, CIPO requires neither human process annotations nor an additional reward model. Extensive exper iments on seven in-domain and out-of-domain benchmarks show that CIPO reduces the prevalence of prior-driven rea soning and achieves excellent performance on most tasks.
[AI-26] Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
链接: https://arxiv.org/abs/2608.06122
作者: Omar Coser,Antonio Orvieto,Paolo Soda,Loredana Zollo
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 21 pages, 7 figures,4 tables
Abstract:Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series. Our objective is to assess the impact of SPT on the performance and scalability of transformer-based models across diverse medical applications, particularly under limited data conditions. We evaluate transformer architectures on three representative medical time-series tasks: rehabilitation robotics (Camargo dataset), stress detection (Non-EEG Stress), and Parkinson’s disease detection (Gait Parkinson’s Disease). Models are trained either from scratch or through SPT using four masking-based objectives designed to promote temporal and cross-modal representation learning, and we systematically vary model depth to examine how capacity interacts with pre-training benefits. Across datasets and configurations, SPT consistently improves classification accuracy by 0-6 percentage points depending on masking strategy, dataset and architecture, with gains observed not only in multivariate settings but also when models are restricted to simple univariate inputs. The improvements increase for deeper models that can better exploit the enriched temporal representations learned during pre-training. These findings indicate that SPT is a simple and general strategy that enhances transformer performance on medical time-series tasks without requiring task-specific architectural changes, supporting its potential to improve robustness and accuracy in data-limited clinical settings.
[AI-27] Mind the Gaps: Mixture-of-Minds for Human Simulation
链接: https://arxiv.org/abs/2608.06115
作者: Pranav Dahiya
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population’s central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.
[AI-28] Evaluating Investment Logic in Large Language Models : A Real-World Benchmark Towards Personalzied Financial Agents
链接: https://arxiv.org/abs/2608.06108
作者: Yuanhong Jiang,Jingjie Zou,Zhenghong Lin,Xusheng Yu,Qiqi Huang,Shuai Jia,Shijie Dai
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \textscInvestLogicBench, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \textbfP \rightarrow E \rightarrow R \rightarrow D \rightarrow O trace: investor \textitProfile, observable market \textitEvents, investment \textitReasoning, executable \textitDecision, and delayed \textitOutcome. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8–2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P \rightarrow E \rightarrow R \rightarrow D \rightarrow O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.06108 [cs.AI] (or arXiv:2608.06108v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.06108 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-29] Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference
链接: https://arxiv.org/abs/2608.06085
作者: Yifan Lyu,Xinran Li,Jiaqi Qiao,Xiujuan Xu
类目: Artificial Intelligence (cs.AI)
备注: 7 pages, 2 figures
Abstract:Survey-country metadata can improve an LLM’s forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label’s uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.
[AI-30] When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
链接: https://arxiv.org/abs/2608.06057
作者: Xiaoqing Wu,Xingyu Fan,Feifei Li,Wenhui Que
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semantically plausible after they cease to be authoritative for the current request. We show that such history can hijack a policy the model already possesses: on Qwen3-1.7B, pollution flips 32.1% of decisions that are correct under the original trajectory and frequently induces reuse of corrupted entities or interface conventions. We introduce bench, a paired benchmark with synchronized Original, Polluted, and Oracle State views that preserve the system policy, current tools, latest request, and gold next action. Eleven gold-preserving interventions isolate failures in decision state, entity binding, and interface execution across complete calls and non-call decisions. We further propose ours, which transfers an Oracle-conditioned teacher policy to a student observing only polluted history through soft supervision on student-generated prefixes. On Qwen3-1.7B, ours achieves 87.0% Balanced Tool-Use Accuracy, outperforming Gold-SFT (66.3%), Oracle sequence distillation (82.3%), and off-policy token distillation (85.0%). The method scales consistently: an 8B teacher raises the same compact 1.7B student to 91.9%, while an 8B student reaches 93.0%. The resulting policies further transfer to clean histories, unseen functions, independently regenerated evaluation contexts, external tool-use benchmarks, and noisy multi-hop question answering. These results establish history reliability as a distinct tool-use bottleneck and demonstrate reliable-state policy transfer as an effective and scalable solution.
[AI-31] From Economic Agents to Agent ic Economies: A Systems Blueprint for Economic World Models
链接: https://arxiv.org/abs/2608.06020
作者: Jiale Han,Xiang Li,Jing Qian,Wenyuan Gu,Pin Gao,Ye Luo,Hongyuan Zha,Dacheng Tao,Benyou Wang,Lin William Cong
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Project page: this https URL
Abstract:Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. This paper develops an implementation roadmap for building economic world models as generative engines in which heterogeneous agents act, interact, adapt, and co-evolve with markets and institutions, thereby producing economic dynamics from the inside. We organize EWM systems into a six-level capability ladder, from fixed rule-based agent worlds to adaptive and LLM-based agent worlds, self-evolving agents, evolving institutional worlds, and sim-to-real economic twins aligned with real observations. A systematic literature survey across these levels reveals that existing work remains concentrated in lower-level agent and simulation environments, while systems with self-evolving agents, endogenous institutions, persistent empirical alignment, and validated economic mechanisms remain rare. By translating the EWM agenda into an implementation blueprint, this paper aims to accelerate the development of the next generation of economic simulation environments that can serve as high-fidelity sandboxes for human decision-makers and as training, planning, evaluation, and safety substrates for AI agents. We release a curated paper list and related resources to support future research.
[AI-32] ProDVI: Programmatic Dynamics Priors for Value Network Initialization
链接: https://arxiv.org/abs/2608.06015
作者: Xinwei Liu,Junyuan Liang,Jianting Zhang,Wuhui Chen
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-relevant knowledge through online interaction. Existing approaches obtain informative initializations through pre-collected datasets, high-fidelity simulators, or meta-learning over related tasks, but these prerequisites may be difficult to access or even unavailable. In this paper, we propose Programmatic Dynamics Priors for Value Network Initialization (ProDVI), a framework that leverages the commonsense and domain knowledge encoded in large language models to initialize RL agents without relying on these resources. Specifically, ProDVI prompts a code-generating language model to produce executable Python functions that encode coarse hypotheses about environment dynamics. These functions are then used to generate synthetic transitions. Based on these transitions, we construct an auxiliary dynamics prediction objective to pretrain the state-action encoder of the value network in an actor-critic framework. The learned representation provides dynamics-aware inductive biases before online RL begins. Importantly, the generated programs are used only for representation pretraining and are not required to faithfully simulate the target environment. While the generated programs may be inaccurate, their induced initialization can be corrected through online learning from real transitions and rewards. Experiments on OpenAI Gym and DeepMind Control Suite tasks show that ProDVI can effectively improve the sample efficiency of model-free RL algorithms.
[AI-33] HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
链接: https://arxiv.org/abs/2608.06012
作者: Zhuowen Liu,Bohan Cui,YinShang Guo,Yuting Wang,Hao Li
类目: Artificial Intelligence (cs.AI)
备注: 9 pages, 3 figures, and 4 tables
Abstract:Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, R_0 rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete 2^3 ablation identifies targeted strengthening of L —citing a corpus passage absent from the retrieved evidence—as the observed inclusion-minimal repair: R[L] has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, R[L] meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural L is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.
[AI-34] Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems
链接: https://arxiv.org/abs/2608.06001
作者: Muhammad Riaz Hasib Hossain,Rafiqul Islam,Shawn R. McGrath,Md Zahidul Islam,David W. Lamb
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid machine learning framework for herd level cattle weight forecasting using automated sensing observations collected between 2022 and 2024 in southeastern Australia. Weekly live weight observations, demographic variables, and lagged environmental predictors were integrated into structured forecasting datasets. Herd level forecasting trajectories were generated through temporal aggregation of animal level predictions. Four hybrid architecture families were evaluated, including residual, stacked, cascade, and ensemble assisted frameworks. ARIMA, LSTM, and GRU models were used as comparative baselines. Independent testing demonstrated strong predictive agreement across multiple forecasting horizons. The cascade GB to RF to NN architecture achieved the best performance, with a test R^2 of 0.889, RMSE of 21.319 kg, and MAE of 15.462 kg. Hybrid architectures maintained greater robustness than recurrent sequential models under sparse observation conditions. Forecasting error increased progressively across extended prediction horizons. Feature importance analysis identified animal age, rainfall, and temperature as dominant predictors influencing herd level growth forecasting. The proposed framework may support feed allocation, grazing management, and livestock marketing decisions under heterogeneous sensing environments.
[AI-35] OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
链接: https://arxiv.org/abs/2608.05990
作者: Ning Xu,Xiang Zheng,Fuqiang Zhong,Huadong Wang,Xiaolong Wu,Zhiyuan Liu,Hui Ning
类目: Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Applied Physics (physics.app-ph); Optics (physics.optics)
备注: 34 pages, 15 figures
Abstract:Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents experimental actions as optical operators and evaluates their outcomes using physically interpretable residuals. Operators specify executable changes to measurement, control or reconstruction, while residuals report departures from specified physical conditions. The agent uses both to select, combine or generate operators, and physical performance is evaluated independently against a withheld reference. Across three optical tasks, score-only feedback produced score increases without physical improvement in 23.6–39.0% of decisions, compared with 0.9–1.9% for operator-residual feedback. Operator-residual feedback increased the probability of reaching and maintaining task targets and reduced experimental budgets. Protocols selected in digital twins were transferred to three optical instruments, and repeated experiments showed a lower projection budget in structured-light reconstruction. Together, operators and residuals guide autonomous decisions using measurable physical evidence.
[AI-36] Agent OPSD: Recursive Self-Distillation for Agent ic Reinforcement Learning
链接: https://arxiv.org/abs/2608.05987
作者: Zi-Han Wang,Zhengxi Lu,Zhiyuan Yao,Jinyang Wu,Jie Wu,Zhengzhou Cai,Yueqing Sun,Ziang Ye,Linji Hao,Qi Gu,Xunliang Cai,Yongliang Shen,Yujiu Yang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Code: this https URL
Abstract:Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
[AI-37] mporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
链接: https://arxiv.org/abs/2608.05981
作者: Yichen Zhang,Yixiong Xiao,Congxi Xiao,Jingbo Zhou
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. However, the acquisition of such high-resolution climate information is often prohibitively costly, necessitating the development of data-driven meteorological prediction models. These models aim to generate fine-grained climate data from low-resolution inputs, a process termed climate data super-resolution (SR). Nevertheless, recent advancements in deep learning for climate data SR have primarily focused on leveraging single-frame spatial information, largely neglecting the temporal correlations between different time frames that could enhance SR outcomes. Furthermore, climate data are inherently stochastic and noisy, rendering widely used temporal alignment methods, such as optical flow models, ineffective in this context. Consequently, the development of a framework tailored for climate data SR that effectively captures implicit temporal correlations remains an unresolved challenge. To this end, we propose a novel Temporal-Enhanced framework with bidirectional temporal alignment. In essence, our framework establishes a temporal bridge to enhance spatial resolution in climate data SR through bidirectional alignment, leading to improved SR performance. Within this framework, Paired Latent Mapping achieves spatial alignment and noise reduction by unifying latent spaces. Then a Bidirectional Temporal Alignment captures temporal correlations by training forward and backward networks on consecutive latent frames. Temporal Enhanced Super-resolution then optimizes the entire framework for climate data SR. Experiments on large-scale real-world datasets demonstrated the superior performance of our framework.
[AI-38] RACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions
链接: https://arxiv.org/abs/2608.05975
作者: Taehyeon Kong,Woojin Kim,Jemin Hwangbo
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注: 8 pages, 7 figures. Submitted to IEEE Robotics and Automation Letters (RA-L)
Abstract:In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged robots under unreliable contact conditions. The proposed estimator directly predicts relative displacement, relative rotation, and body-frame velocity from a recent history of onboard inertial and joint measurements. To improve robustness under unreliable contact conditions, we introduce a foot-aware cross-attention module that adaptively weights IMU and leg-wise kinematic tokens without relying on manually defined contact or slip thresholds. The estimator is trained with direct supervision and two physics-inspired auxiliary losses that promote kinematic consistency and reliable use of leg information. To reduce policy-specific overfitting and consequently improve sim-to-real transfer, simulation training incorporates policy randomization, followed by partial real-world fine-tuning of the temporal encoder and prediction head. Experiments across diverse indoor and outdoor terrains demonstrate consistent reductions in position drift compared with classical filtering-based, hybrid, and purely learning-based baselines. Ablation studies further validate the contributions of the proposed training objectives, policy randomization, and real-world fine-tuning, particularly under unreliable contacts and sim-to-real mismatch.
[AI-39] SkillM emo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
链接: https://arxiv.org/abs/2608.05970
作者: Changyuan Wang,Chubin Zhang,Zhenyu Wu,Runhao Li,Angyuan Ma,Ke Chao,Yinan Liang,Xiuwei Xu,Ziwei Wang,Yansong Tang,Jiwen Lu
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, leading to insufficient compositional generalization in out-of-distribution (OOD) scenarios with limited capability to capture reusable skill structures. To address this limitation, we propose Skill-Based Memory (SkillMemo) framework that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks. Specifically, we first introduce an expert-guided trajectory segmentation module built upon a Mixture-of-Experts (MoE) architecture, which implicitly partitions trajectories into distinct skill primitives represented by learned gating coefficients. We further design a skill-level episodic memory architecture that stores compact skill representations as retrievable key-value pairs. During inference, the memory bank retrieves the most relevant skill primitives which are subsequently fused with the model’s current gating distribution, providing a robust contextual prior to refine action predictions. Extensive experiments on the simulation benchmark and real-world manipulation tasks demonstrate that SkillMemo consistently enhances both DP and VLA backbones, achieving state-of-the-art performance and outperforming \pi_0.5 , while exhibiting strong compositional generalization to unseen task configurations.
[AI-40] Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
链接: https://arxiv.org/abs/2608.05958
作者: Vitaliy Tsyganok,Sergii Kadenko,Oleh Andriichuk
类目: Artificial Intelligence (cs.AI)
备注: 14 pages, 14 figures
Abstract:The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the quality of expert data during estimation and help reduce the number of comparisons that the experts need to perform. In the paper we compare three incomplete ranking-dependent pair-wise comparison patterns which can be used in the Analytic Hierarchy Process - Best-worst method, Best-Second Best (Top 2) method, and the original maximum difference method. The first two comparison patterns (and respective methods) are incomplete, while the third can be a complete one. We determine conditions under which these three methods can be compared in terms of stability to expert errors. We also present the results of a simulation-type experiment, in which the three methods are compared. The research allows us to define the most stable incomplete ranking-dependent pair-wise comparison pattern and reduce the number of comparisons without loss of credibility of expert session results. The research contributes to algorithmic, cognitive, and applied aspects of decision support in uncertain environments.
[AI-41] BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
链接: https://arxiv.org/abs/2608.05926
作者: Guanqiao Qu,Shuo Chen,Qian Chen,Kin K. Leung,Xianhao Chen
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注: 10 pages, 7 figures
Abstract:Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decoding (AD) generates output tokens sequentially, resulting in long latency; Speculative decoding (SD) accelerates inference by using a small language model (SLM) to generate multiple draft tokens for LLM verification, but incurs extra memory costs. Due to this latency-memory tradeoff, neither approach alone can efficiently serve users with heterogeneous demands under limited edge computing resources. To address this challenge, we propose a hybrid autoregressive-speculative inference (BALANCE) framework for edge LLM inference. In BALANCE, an edge server hosts both an SLM and an LLM, assigns each user to AD or SD, and performs the two modes simultaneously. To maximize the number of served users, we formulate a task throughput maximization problem to jointly determine user scheduling and computing resource allocation between AD and SD under user latency requirements and server memory constraints. Since the problem is NP-hard, we develop a polynomial-time algorithm that transforms the original problem into two sub-problems and obtains a sub-optimal solution with a constant approximation guarantee. Experiments demonstrate that BALANCE consistently outperforms conventional AD and SD and significantly improves task throughput.
[AI-42] CourseGraph: Finding overlaps and differences in Computer Science courses across universities
链接: https://arxiv.org/abs/2608.05910
作者: Arthur Nijdam,Paul Stankovski Wagner,Sara Ramezanian
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注:
Abstract:Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural horizons. However, this flexibility also leads to a practical challenge: ensuring that students do not take courses elsewhere that substantially overlap with courses in their home curriculum. In this work, we propose CourseGraph, a methodology that automates the evaluation of external courses based on insights obtained from the process followed by curriculum administrators when assessing courses for inclusion in a degree program. Course- Graph extracts information such as course titles, descriptions, and learning outcomes from the course webpage. Then, this information is represented semantically using a BERT-based language model, after which the pair-wise similarity between courses can be computed. This information is then used by a Random Forest classifier to determine whether a candidate course abroad overlaps with a course already contained in the student’s curriculum. We evaluate CourseGraph using (1) the Computer Science program at Eindhoven University of Technology, which contains information about courses with substantial overlap, and (2) six approved international programs from students enrolled in the Computer Science program at Lund University, including the corresponding decisions made by a curriculum administrator. The experimental results indicate that CourseGraph provides an effective approach for identifying overlapping courses and supporting curriculum alignment across universities.
[AI-43] GSBF: Gaussian Splatting for Environment-Aware Beamforming
链接: https://arxiv.org/abs/2608.05896
作者: Yijie Bian,Wei Guo,Zixin Wang,Shenghui Song,Jun Zhang,Khaled B. Letaief
类目: Artificial Intelligence (cs.AI); Information Theory (cs.IT)
备注:
Abstract:Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity. Recognizing that radio propagation is intrinsically governed by the physical geometry, we develop a 3D Gaussian splatting for environment-aware beamforming (GSBF) pipeline based on multi-modal data, which characterizes the environment through a persistent 3D Gaussian representation. Specifically, GSBF models the environmental scattering response with reciprocity-preserving bidirectional spherical Gaussian (Bi-SG) kernels and performs two-sided electromagnetic rasterization to render an angular propagator map. The rendered map is then aggregated through an over-complete array-manifold dictionary and projected to the constant-modulus beamformers, thereby synthesizing beams directly from the access point (AP) pose and user position without online instantaneous CSI. Simulations demonstrate that GSBF consistently outperforms baselines such as exhaustive beam alignment (EBA) with lower latency.
[AI-44] ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
链接: https://arxiv.org/abs/2608.05893
作者: Akanta Das,Tasinul Islam Ahon,Ahmed Mahir Sultan Rumi,Md Mahbubur Rahman,Tausif Amim Shadly,Tanzima Hashem
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-lead ECG recordings into reliable clinical reports remains challenging. Automating ECG report generation could reduce clinicians’ interpretive workload, improve diagnostic efficiency, and expand access to cardiac assessment in underserved communities. Unlike image-based report-generation tasks, ECG interpretation requires the analysis of subtle temporal morphologies, followed by coherent diagnostic reasoning expressed in dense clinical terminology. Existing systems predominantly focus on classification, while current report-generation methods often produce outputs that remain inadequate for practical clinical use. To address these challenges, we propose ECG-LENS, an end-to-end ECG report-generation framework that jointly integrates multi-lead signal modeling, diagnosis-aware representations, and clinically grounded text generation. ECG-LENS combines lead-wise encoders that preserve localized waveform morphology with a global encoder that captures inter-lead dependencies. To guide report generation, we fuse signal representations with clinically enriched textual prompts that condition a GPT-2 decoder. We further introduce an ECG-specific report-preprocessing strategy that helps the model focus on clinically meaningful findings. Finally, because lexical metrics may under- or overestimate report quality, we propose F1-ECGBERT, a BERT-based, ECG-specific metric that measures agreement between diagnostic labels extracted from generated and reference reports. In-domain experiments on PTB-XL and cross-domain evaluation on MIMIC-IV-ECG show that ECG-LENS consistently outperforms state-of-the-art methods, with absolute gains of 4.0%, 6.3%, and 11.5% in METEOR, ROUGE-L, and F1-ECGBERT, respectively, over the strongest baselines.
[AI-45] CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents
链接: https://arxiv.org/abs/2608.05886
作者: Wuya Chen,Yihao yang,Yang Cao,Yue Lin
类目: oftware Engineering (cs.SE); Artificial Intelligence (cs.AI)
备注:
Abstract:Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the file to patch, rather than patching it. On SWE-Bench Verified, a 30B OpenHands agent averages 23 rounds and 631K tokens per resolved issue, with many calls spent on grep, glob, and view_file during repository exploration. We introduce CodeGrep, a 14B retrieval agent trained end-to-end with GRPO to issue multi-turn parallel grep, glob, and read tool calls and return candidate files to a frozen downstream coding agent. On all 500 SWE-Bench Verified instances, CodeGrep preserves resolve rate while substantially improving efficiency: 27.0% versus 25.8% for the no-retrieval baseline, with 15% fewer rounds and 19% fewer tokens on resolved instances. Across retrievers, downstream utility follows a precision threshold: BM25 with precision 0.375 degrades the agent, Jina with precision 0.445 is neutral, and CodeGrep with precision 0.677 crosses the threshold at which retrieval begins to reduce rollout cost. To enable this study, we mine supervision from 67K open-source agent trajectories using CATM and build a Git-worktree environment for multi-turn agent RL. In our setting, applying the efficiency signal at the advantage layer rather than the reward layer reduces KL drift and translates cleanly into downstream efficiency. We will release the model, training pipeline, RL environment, and evaluation harnesses. Subjects: Software Engineering (cs.SE); Artificial Intelligence (cs.AI) Cite as: arXiv:2608.05886 [cs.SE] (or arXiv:2608.05886v1 [cs.SE] for this version) https://doi.org/10.48550/arXiv.2608.05886 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-46] Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
链接: https://arxiv.org/abs/2608.05880
作者: Benjamin Connor,Anna Jurek-Loughrey,Lu Bai,Muhammad Fahim
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: 6 pages. Accepted in 36th Irish Signals and Systems Conference (ISSC) 2026
Abstract:Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis. Results demonstrate that although each method can successfully recover relevant features, none consistently detects all injected pattern types. These findings high- light a critical gap between existing explainability tools and the requirements of pattern-level cluster interpretation, motivating the development of dedicated pattern detection methodologies.
[AI-47] Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks ISWC2026
链接: https://arxiv.org/abs/2608.05867
作者: Jonathon Dilworth,Pedro Giesteira Cotovio,David Herron,Paul Cripps,Nigel Dewdney,Catia Pesquita,Ernesto Jiménez-Ruiz
类目: Artificial Intelligence (cs.AI)
备注: Paper accepted at the Resource Track, the 25th International Semantic Web Conference (ISWC 2026), 25 - 29 October 2026, Bari, Italy
Abstract:The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontologies have been developed through initiatives led by academia, industry, and government. Achieving interoperability across diverse defence and national security ontologies remains a major challenge due to the domain’s breadth and specialisation. In this work, we analyse and document over 60 publicly available ontologies and introduce a new track for the Ontology Alignment Evaluation Initiative (OAEI). This track comprises eight matching tasks, consensus alignments and manually-curated (silver-standard) mappings. The consensus alignments are derived by aggregating the outputs of several state-of-the-art ontology alignment systems. The silver-standard is obtained from the manual validation of the consensus alignment together with a subset of the unique mappings (i.e., mappings suggested by only one system).
[AI-48] Seeing Is Not Deciding: Can Multimodal LLM s Act as Effective CEOs?
链接: https://arxiv.org/abs/2608.05864
作者: Yuyang Dai,Xueqing Peng,Yuxia Wang,Preslav Nakov,Zhuohan Xie
类目: Artificial Intelligence (cs.AI)
备注: 25 pages
Abstract:Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.
[AI-49] Runtime Observability for Heterogeneous Attention Memory
链接: https://arxiv.org/abs/2608.05863
作者: Fanzhe Wei,Li Liu,Ziyang Wang,Chenyu Wang
类目: Artificial Intelligence (cs.AI)
备注: 29 pages, 5 figures. Code, artifacts, and Lean 4 development: this https URL
Abstract:Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model’s memory in a different form, and each fails differently under compression. We give a runtime observability contract that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger. Contracts carry their error metric as a type – composition is only defined when metrics match, and this check rejected our own first composed chain; the repaired chain crosses metrics through two proved bridges, and whatever no formal system can certify is measured instead, dropping the composed tier to empirical automatically: every claim is certified, partially certified, or empirical, composition inherits the weakest tier, and the tier is decided by the machine. Replayed over 12.4 M entry reads and run under eight-way concurrency with per-request budgets and fail-closed identity attribution, the ledger quantifies the honest trade-off on today’s witness and holds its risk budget with zero violations. A fused always-on probe observes a declared one-layer subset under CUDA graphs inside the serving noise floor. Applied to a served DeepSeek-V4 stack with a packed compressed-KV prototype, the same machinery localizes a silent corruption to a precise structural boundary – exact in the eviction-free, identity-isolated regime, with every observed failure in an eviction or slot-reuse regime – through a machine-adjudicated discrimination campaign whose calculus rejected two of our own confounded inferences along the way. All artifacts, guards, and the Lean development are released at this https URL every number in this paper regenerates from the shipped artifacts by one command.
[AI-50] Evidential Rule Learning for Interpretable Classification with Abstention
链接: https://arxiv.org/abs/2608.05859
作者: Javier Fumanal-Idocin,Javier Andreu-Perez
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the evidence behind their decisions and abstain when they cannot decide reliably. We introduce Fast Evidential Rule Learning (FERL), a method that learns interpretable, accurate fuzzy rule models whose outputs are evidential. Unlike post-hoc calibration, FERL’s belief, plausibility, and abstention capabilities arise directly from the fuzzy memberships in a single deterministic pass, with no auxiliary head, held-out set, or repeated inference. Our theoretical analysis further shows that FERL is Lipschitz stable, which means that its evidential outputs vary smoothly with the input. Against state-of-the-art rule learners, FERL is statistically significantly more accurate across a 30 tabular-dataset benchmark ( +2.6% average accuracy over the second best). Its native set predictions attain the best utility-discounted accuracy among credal classifiers ( u_65/u_80=0.80/0.83 vs.\ 0.79/0.80 for the naive credal classifier), at higher set coverage ( 0.92 vs.\ \le0.82 ). FERL also matches dedicated out-of-distribution detectors on tabular near-OOD detection ( 77.7 vs.\ 77.4 AUROC for the strongest baseline). Under detector-class-disjoint concept-bottleneck evaluation, its it is within 2.3 AUROC points of the strongest dedicated detector on both CUB and AwA2, while attaining the best AwA2 AUPR-Out ( 68.3 ) and novel-class rejection ( 57.2 ), while being able to name which attributes are anomalous.
[AI-51] ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
链接: https://arxiv.org/abs/2608.05833
作者: Jiafan Li,Mengxue Yang,Jiaqi Zhu,Liang Chang,Ying Li,Hongan Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images. Traditional representation learning approaches follow the embedding-based paradigm and may struggle when relation-specific evidence is limited. Meanwhile, LLM-based reasoning methods typically linearize graph structures into textual prompts, which obscures structural topology and neglects vital visual information. While vision-language models (VLMs) excel at multimodal reasoning, they cannot natively interpret structured graph topology, particularly when it comes to knowledge graphs where nodes and edges carry complex semantics. To bridge this gap, we propose ViSR-KGC, a visual subgraph reasoning approach for KGC. It integrates three complementary capabilities to capture semantic correlations: identifying global topology dependencies via representation learning, analyzing local multimodal evidence using VLMs, and providing necessary commonsense knowledge inherent in pre-trained models. Based on learned multimodal embeddings, our framework first extracts a compact and query-aware subgraph from the MMKG. Then, this subgraph is transformed into a visually interpretable image using a layout strategy selected through empirical this http URL, the visualized subgraph, entity images, textual descriptions, and candidate answers are combined into a unified prompt, enabling the VLM to infer the missing entity.
[AI-52] Cautious Context Steering for Language Model Personalization
链接: https://arxiv.org/abs/2608.05813
作者: Gihoon Kim,Jeyoung Lee,Suhan Woo,Sekwon Oh,Minsu Jeon,Hyounsoo Han,Euntai Kim
类目: Artificial Intelligence (cs.AI)
备注: 11 pages, 3 figures
Abstract:Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds. Existing methods typically train a separate adapter for each user or learn a reward model whose scores depend on the user. Despite explicitly optimizing for each user, these methods must learn from limited observations and therefore suffer from data sparsity and poor generalization to unseen users and domains. In-context learning (ICL) and Context Steering (CoS) can instead provide more effective personalization by conditioning the base LM directly on user context and leveraging its pretrained capabilities without per-user training. Yet neither adapts the influence of that context across decoding steps: ICL leaves it uncontrolled, whereas CoS applies a fixed steering coefficient and requires two LM forward passes per step. We propose Cautious Context Steering (CCS), which adds a lightweight adapter to a frozen backbone LM to decide at each token whether and how strongly user context should affect generation. The adapter learns this behavior from an oracle context-conditioned LM and preserves the base LM when the context is not helpful. A single CCS adapter trained on only one dataset improves generation quality both in-domain and across four out-of-distribution personalization benchmarks, demonstrating robust generalization to new users and domains. CCS also avoids per-user fine-tuning and the additional context-conditioned forward pass required by CoS, substantially reducing inference cost.
[AI-53] When Agent ic AI Meets Integrated Sensing and Communication
链接: https://arxiv.org/abs/2608.05792
作者: Kai Li,Conggai Li,Sarah Ali Siddiqui,Syed Sohail Ahmed,Xin Yuan,Shenghong Li,Wei Ni
类目: Artificial Intelligence (cs.AI)
备注: 35 pages, 132 references, 10 tables, 9 figures
Abstract:Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer technology into a goal-driven, closed-loop intelligent system, a paradigm we term AISAC. Existing work on learning-based sensing, resource allocation, reconfigurable intelligent surfaces (RIS), edge intelligence, multi-agent coordination, and resilient networking has developed largely in isolation. This survey unifies the literature within a six-stage closed-loop framework comprising observation, contextualization, reasoning and prediction, planning and orchestration, execution and collaboration, and feedback and resilience. It also introduces five levels of agentic maturity, ranging from physical-layer primitives to fully closed-loop agentic ISAC. We use this framework to review advances in multimodal intelligence, large language models, reinforcement learning, federated learning, RIS-assisted control, Unmanned Aerial Vehicle (UAV) and vehicular networks, and AI-native network management, and analyze privacy, security, resilience, and sustainability as cross-cutting requirements of the full perception-reasoning-action loop. An audit of representative studies against nine agentic-specific evaluation criteria shows that no system reports more than one or two of them, exposing a gap between claimed and demonstrated agentic maturity. We identify open challenges in physical-to-semantic grounding, predictive world models, real-time agent-PHY interaction, safe tool use, heterogeneous multi-agent collaboration, benchmarking, and resource-efficient autonomy.
[AI-54] ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution
链接: https://arxiv.org/abs/2608.05790
作者: Jiacheng Wei,Zhaoxin Fan,Xin Wen,Yuqin Lan,Dongrun Li,Wenjun Wu,Faguo Wu,Xiao Zhang
类目: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注: 8 pages,3 figures
Abstract:General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockchain environments. On-chain execution is stateful, adversarial, and economically irreversible, exposing three fundamental gaps: Reactivity, Irreversibility, and Observability. We propose ChainClaw, a blockchain-native agent framework built on OpenClaw, that addresses all three gaps through a layered architecture comprising an event-driven orchestration layer, a simulation-based safety intelligence layer, and an on-chain monitoring runtime layer, unified by a cross-layer memory subsystem. ChainClaw closes the Reactivity gap via event ingestion and simulation feedback, the Irreversibility gap via a pre-execution safety pipeline with transaction simulation and action guard, and the Observability gap via an on-chain read adapter and transaction monitor. We evaluate ChainClaw on a purpose-built benchmark covering seven tasks across four categories and five dimensions. ChainClaw consistently outperforms representative baselines on both safety and task completion.
[AI-55] Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
链接: https://arxiv.org/abs/2608.05784
作者: Nossa Iyamu
类目: Artificial Intelligence (cs.AI)
备注: 14 pages, 5 figures, 4 tables
Abstract:Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent’s memory today records what the user said, not what the user did. We compile passively captured screen activity into agent memory with a deterministic, zero-model pipeline: it segments a local capture stream into typed activity frames, bounded episodes carrying application, site, timing, input volume, and evidence pointers back to the raw rows, with no model in the loop, so the output is byte-identical, cacheable, and mechanically auditable. On one professional’s single-user corpus of 128,756 frames over 51 active days, the compiler reduces a day of raw capture to a prompt-ready context block 86x smaller in 68 ms, and an agent reading that block answers questions about the day at 98.4% accuracy (Wilson 95% CI 91.7-99.7%) against an independent oracle, versus 66-80% for an LLM summary of the same capture, a mid-tier model reading the block matching a frontier one. The same compiler doubles as a demand-side cost instrument. Read off passive, pre-delegation human activity rather than agent rollouts, it supplies two parameters that agent-cost models assume but, to our knowledge, have not measured: the Routine Overhead Ratio R and the routine recurrence h. We report first values of R, a modeled upper bound, at 60-343x, and a delegable recurrence of 9.0% in-sample and 7.7% out-of-sample, for a realistic all-fleet token ceiling near 8%; a compiled routine replays deterministically with the model out of the loop, demonstrated live at zero model tokens on a guard-matched hit. Schema, compiler, and evaluation harness are open. Comments: 14 pages, 5 figures, 4 tables Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.05784 [cs.AI] (or arXiv:2608.05784v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.05784 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-56] When Do Prompt-Side Agent Playbooks Transfer? Accuracy Cost and Runtime Shift in Agent Deployment
链接: https://arxiv.org/abs/2608.05778
作者: Weihong Lin,Lin Sun,Xiangzheng Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill–validate–transfer protocol. On ALFWorld, transfer is beneficial under controlled greedy decoding and, in one near-budget-matched comparison, distilled guidance outperforms five fixed demonstrations. On TAU2-Bench, a prespecified aggregate contrast supports a modest average matched-domain advantage, but global Holm correction retains only one of 135 route-level effects; the remaining grid provides descriptive evidence of compatibility-sensitive heterogeneity. On XBench-DeepSearch, one artifact–runtime pairing preserves useful first-try heuristics while producing repeated queries, delayed stopping, and substantial cost inflation after a context-runtime shift. Across benchmarks, transferred and target-derived playbooks both require target-side validation of success, termination, protocol compatibility, and cost. Frozen transfer is therefore a conditional cold-start option, not a reuse-by-default strategy or a universally preferable alternative to target-side redistillation.
[AI-57] Multivariate Time Series Forecasting needs Cross Variable Loss
链接: https://arxiv.org/abs/2608.05742
作者: Kuiye Ding,Yifan Hu,Hanchen Wang,Hao Xue
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. Specifically, modern forecasting models largely follow the Direct Forecasting (DF) paradigm, generating multi-step forecasts with point-wise objectives that do not explicitly constrain cross-variable structure. In this work, we show that the DF objective is mismatched in the presence of cross-variable and lagged dependencies, revealing an objective gap. To address this issue, we propose \textbfCross-\textbfVariable \textbfLoss (CvLoss), a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph. CvLoss penalizes inconsistent edge-wise residual differences over forecast patches, encouraging consistency across both synchronous and asynchronous interactions. Our experiments show that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones.
[AI-58] ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge ICDE2026
链接: https://arxiv.org/abs/2608.05737
作者: Incheol Baek,Hyungbin Kim,Yon Dohn Chung
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at IEEE ICDE 2026
Abstract:Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is that existing LDP mechanisms require a predefined data domain, which is often unknown in practice. This lack of prior knowledge creates a critical dilemma for the data collector: if the chosen domain is too narrow, values outside the range are clipped, leading to information loss. Conversely, if the domain is too wide, excessive noise is added during the privatization process, which degrades the quality of collected data. This highlights the need for methods that can dynamically estimate the data domain. In this work, we propose an adaptive LDP framework that addresses this problem. In our method, each user sends two pieces of information: their perturbed numerical data, and a privatized signal indicating if their original value was clipped by the current domain. By aggregating these signals, our proposed method, Adaptive Bounding of Clipping regions (ABC) method, iteratively adjusts the domain to fit the underlying data distribution without prior knowledge. Our theoretical analysis shows that the estimated data domain converges to an appropriate range. In the empirical evaluation, the results demonstrate that our framework significantly improves the quality of numerical data collection across various datasets and underlying LDP mechanisms. We also show that the estimated range successfully converges in practice and our approach is robust to its hyperparameters through comprehensive ablation studies. Comments: Accepted at IEEE ICDE 2026 Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2608.05737 [cs.CR] (or arXiv:2608.05737v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.05737 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-59] Subliminal Learning is Non-Semantic Distillation ICML2026
链接: https://arxiv.org/abs/2608.05734
作者: Ethan Hadley,Eren Gultepe
类目: Artificial Intelligence (cs.AI)
备注: Accepted as spotlight paper for the ICML 2026 Mechanistic Interpretability Workshop
Abstract:Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or behavior from a teacher model to a student by distilling from seemingly unrelated or random synthetic data from the teacher. This presents challenges in ensuring AI systems remain predictable and are trained safely, as standard auditing of the input data would not catch the hidden subliminal signal. Here, we investigate several open questions as to the enabling mechanisms and drivers of SL. First is the nature of the process by which biases are encoded in the data. We find that by adding Gaussian noise to the weights of the teacher and student models, the magnitude of subliminal transfer is increased by a factor of 1.9 in Gemma and 1.3 in Llama, suggesting that non-semantic weight structures play a crucial role. We show that steering vectors can be applied to the teacher to produce subliminal data, in addition to prompting and finetuning as used in previous studies. Analysis of the activations of the student models that have been trained on steered and prompted data demonstrates that students inherit not just the semantic meaning of the teacher’s bias, but also the type of intervention that was used to apply it: steered students imitate steering vectors, prompted students do not. Additionally, the gradients of steered subliminal data show a linear correlation with the teacher’s steering vectors, showing promise for data auditing. More broadly, as synthetic data becomes central to frontier training pipelines, being able to see the latent signals hidden in training data becomes paramount.
[AI-60] BlockPython: A Process-Aware Agent -Supported Platform for the Transition from Block-Based to Python Programming
链接: https://arxiv.org/abs/2608.05716
作者: Jesse Yusuf Chan(Zexi Chen),Haoming Wang,Mingwei Xu,Xianlong Xu
类目: Artificial Intelligence (cs.AI)
备注: AIED 2026 Interactive Event Track
Abstract:The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual expressions, which may create a cognitive gap between understanding computational concepts and expressing them in Python syntax. To support this transition, we designed and implemented BlockPython. The platform centers on bidirectional translation between blocks and Python and guides learners through four stages: Task Decomposition, Block-Based Practice, Code Challenge, and Extended Interaction. Across these stages, learners progressively establish connections among program structure, runtime behavior, and textual code. During learning, the platform continuously collects process evidence, including block artifacts, code versions, run outcomes, use of support, and dialogue. Deterministic diagnosis, program visualization, and the learning assistant use this evidence to identify different difficulties in computational understanding and Python expression. The rule-based system is responsible for program execution, objective evaluation, and stage control, while the learning assistant uses verified evidence to provide explanations, prompts, and guiding questions. This report describes the design rationale, learning workflow, and process-aware support mechanisms of BlockPython and provides a system-design reference for supporting the transition from block-based to text-based programming and for analyzing learning processes.
[AI-61] Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
链接: https://arxiv.org/abs/2608.05715
作者: S. M . Bhagya P. Samarakoon,M. A. Viraj J. Muthugala,W. K. R. Sachinthana,Mohan Rajesh Elara
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI)
备注:
Abstract:Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into executable actions grounded in visual scene understanding. This tight coupling between perception and instruction-following introduces a new attack surface: adversarial text placed within the robot’s visual field can act as an indirect prompt injection into the VLM’s reasoning stack. We present a systematic study of physical prompt injection attacks against VLM-controlled sorting, introducing a four-category taxonomy, indirect signage, task redefinition, authority impersonation, and conflict injection, instantiated as a benchmark of 20 attack prompts evaluated across three physical scene layouts and three command formulations that vary in destination specificity and rule explicitness. Across 5,670 trials on three frontier VLMs (GPT-4o, Gemini 2.5 Flash, Qwen3-VL-32B), attacks succeed at 27.0%, 29.4%, and 5.0% respectively, with authority-impersonating and negation attacks transferring across all three models. Analysis of reasoning traces reveals that successful compromise is nearly always conscious (99.9% acknowledgment rate), and that models defend through structurally different mechanisms, explicit rejection for Gemini, perceptual inattention for GPT-4o. We evaluate three simple mitigations: prompt-based defense (75-100% effective, model-dependent), two-stage verification (85-100%), and pre-processing text masking (100%). Our findings show that VLM-controlled manipulation is meaningfully vulnerable to human-readable physical signage, and that simple defenses substantially reduce risk, though defense choice involves trade-offs. The defenses preserve general task capabilities in our benchmark, but they may impair tasks that require reading in-scene labels.
[AI-62] RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
链接: https://arxiv.org/abs/2608.05714
作者: Shuhao Yan,Changhao He,Xi Peng,Peng Hu
类目: Artificial Intelligence (cs.AI)
备注: 17 pages, 7 figures, 5tables, 8 listings
Abstract:Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes, reducing the expertise and effort required for manual modeling. Existing methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms to optimize the generation process, but they do not necessarily optimize how feedback is interpreted and translated into effective corrective actions throughout the generation process. To bridge this feedback-utilization gap, we present RA-CAD (ReAct Agent for CAD), a state-aware agent that interacts with the CAD environment through a Generate–Execute–Critique–Rewrite loop. At each iteration, RA-CAD executes the current code and observes its outcome. Conditioned on the design instruction, current code, and execution feedback, the agent then generates an explicit post-execution critique as an intermediate policy action. This critique either validates the current result for termination or provides revision-oriented guidance that conditions the next rewrite. CAD Code Bootstrapping (CCB) first establishes fundamental parametric CAD coding capabilities through supervised fine-tuning. Feedback-Driven Agent Optimization (FAO) subsequently applies trajectory-level Group Relative Policy Optimization to both policy-generated code and critique sequences, assigning terminal F1 and Chamfer Distance rewards to the complete interaction trajectory. This formulation makes critique an outcome-aligned, learnable policy decision rather than an unoptimized auxiliary output. Experiments on CADFusion and Text2CAD show that RA-CAD achieves state-of-the-art execution validity and geometric quality compared with existing methods and strong proprietary language models, demonstrating the effectiveness of the proposed state-aware text-to-CAD agent.
[AI-63] Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives
链接: https://arxiv.org/abs/2608.05710
作者: Keziah Naggita
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop accurate beliefs about the AI systems so they can improve and or secure favorable outcomes at minimal cost, (2) encourage improvement and or discourage gaming behaviors, and (3) ensure that the AI system continues to achieve its intended objectives, such as maximizing accuracy. To address these goals, the thesis is organized into three complementary parts that examine and study human-AI interactions from the perspectives of both evaluated individuals and AI systems. Together, the work presented in this thesis advances human-centered machine learning by providing principles and methods for designing AI systems that align with human needs, values, and capabilities. Methodologically, this thesis integrates theoretical analysis, data-driven modeling, human-subject experiments, and empirical evaluations on real-world and semi-synthetic datasets.
[AI-64] Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery
链接: https://arxiv.org/abs/2608.05705
作者: Victor Gialis,Maxime Metz,David Esteve,Abdenour Soualhi
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We propose Spectral Aliasing Pretext (SAP), a self-supervised learning method that pretrains models on unlabeled vibration data by exploiting spectral aliasing. We deliberately undersample signals to create folded spectrum, then train a Transformer to reconstruct the original unfolded spectrum. This pretext task forces the model to learn frequency-domain invariants characteristic of mechanical faults, without potentially destructive augmentations. Experiments on the CWRU dataset show that SAP learns stable and highly discriminative representations. In a linear probing setting, SAP quickly achieves very high classification performance with only a small fraction of labeled data and low variance. In contrast, full fine-tuning, including fully supervised training, does not lead to more stable or better results. Overall, these findings suggest that SAP combined with linear probing can be more effective and reliable than fully supervised training for fault diagnosis with limited labeled data.
[AI-65] Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels a Baseline and a Dual-Head Model
链接: https://arxiv.org/abs/2608.05685
作者: Gospel Bassey,Vincent Fakiyesi
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known. Real production fields rarely offer that. They give you sensor histories with no fault log attached, which is exactly the situation where an anomaly-detection method has to invent its own labels, and where quiet assumptions can slip in unnoticed. We work with the open Volve field data released by Equinor and take two things seriously that such datasets usually skip. First, we build anomaly labels that are not just patterns in the numbers but are checked against what the field’s own engineering documents say can physically go wrong, and we release the reasoning behind every label. Second, we test whether those constructed labels are learnable at all, using both an unsupervised baseline and a small dual-head model that marks when an event happens and what kind it is, an idea we carry over from earlier work on defect detection in metal parts. The results are honest. An unsupervised detector that never sees the labels still lands on the same regions our rules flagged, which tells us the labels are not arbitrary. A compact supervised model recovers event presence and event type well across wells it has never seen, and locates events in time only roughly. We report what worked, what did not, and every assumption in between. The dataset, grounded labels, per-label provenance, baseline scores, trained model, and code are released publicly under CC-BY-NC-SA 4.0.
[AI-66] Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot
链接: https://arxiv.org/abs/2608.05684
作者: Hyoto Yamaguchi,Zenji Yatabe,Seiya Kasai
类目: Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Systems and Control (eess.SY)
备注: 5 pages, 7 figures, The paper has been submitted to IEEE SCIS ISIS 2026 for consideration
Abstract:Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walking robot. To classify ground condition without image sensing and processing, we implemented artificial proprioception by integrating a three-axis accelerometer, eight foot pressure sensors, and reservoir computing (RC). Even when large fluctuations in the sensor outputs are caused by dynamic motions of a four-legged robot in walking, our system can classify the ground condition, flat or rough, with high accuracy. We demonstrate on-site switching of walking gait depending on ground condition in the robot. We also discuss the contribution of each sensor to ground condition classification.
[AI-67] Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying
链接: https://arxiv.org/abs/2608.05642
作者: Shimon Honda,Takuma Miyaguchi,Koji Koizumi,Takanori Sano,Tristan Briard,Hideyoshi Yanagisawa
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design action as its expected reduction of epistemic uncertainty about action–outcome relations. The model addresses one part of the Uncertainty Driven Action (UDA) model’s open question concerning how changes in uncertainty perception determine action selection. We examine two environmental properties: generalizability, or how far knowledge from one trial extends to neighboring candidates, and outcome discriminability, or how clearly differences among outcomes can be distinguished. We tested the model through simulations and human experiments using a graph-shape guessing task that isolates learning about action–outcome relations under a limited trial budget. Epistemic value followed an inverted-U-shaped relationship with generalizability and increased with outcome discriminability in the simulations. In the human experiments, the subjective value of trying and enjoyment followed inverted-U-shaped relationships with generalizability, while choice behavior reflected both properties. The B-EUR model provides a computational account of candidate-action evaluation within uncertainty-driven design activity and offers implications for constructing prototype sets, framing design problems, and organizing feedback to support informative exploration. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.05642 [cs.AI] (or arXiv:2608.05642v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.05642 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-68] SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
链接: https://arxiv.org/abs/2608.05628
作者: Yuru Feng,Yaoqi Chen,Beidi Zhao,Qianxi Zhang,Xinjiang Wang,Jianan Lu,Zhirui Wang,Shusen Xu,Zengzhong Li,Qi Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus require autonomous, on-demand skill evolution at test time, constrained by limited interaction budgets and a lack of training or validation sets. This setting introduces a severe sparse reward challenge, where outcomes conflate multiple latent failure causes. Under such ambiguity, existing methods that greedily refine a single incumbent skill are particularly vulnerable to an exploitation trap, allowing early misdiagnoses to exhaust limited trials along unproductive trajectories. To address this, we introduce SkillHEX, a closed-loop framework coupling hypothesis-driven self-verification with evidence-guided tree search. SkillHEX translates falsifiable failure hypotheses into executable tests, producing diagnostic evidence as dense reward without additional environment attempts. This evidence guides a search over persistent skill-revision branches, dynamically balancing the exploitation of supported edits with the exploration of plausible alternatives. Evaluated on 87 tasks from SkillsBench, SkillHEX outperforms existing self-evolving methods and achieves an average pass rate of 55.9% and 57.9% using GPT-5.3-Codex and Claude Opus 4.7 under a five-iteration budget, respectively.
[AI-69] GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification
链接: https://arxiv.org/abs/2608.05608
作者: Yunping Shi,En Yu,Kairui Guo,Jie Lu
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynamic fusion can mitigate such incompleteness, but existing methods operate at a coarse modality level and thus cannot retain reliable components while suppressing misleading ones within the same recovered modality, compromising prediction reliability. To address this issue, we propose GAUGE, a lightweight counterfactual gating framework for incomplete multimodal classification. GAUGE first imputes missing modalities with a frozen imputer and encodes observed and recovered inputs uniformly as fine-grained evidence units. Rather than intervening on each unit explicitly, GAUGE scores the counterfactual effect of replacing every unit with a reference representation through prediction-aware Taylor evidence scores, all obtained in a single forward-backward pass. These scores are mapped to continuous gates, which are converted into additive attention-logit biases for unit-wise evidence modulation without altering the backbone architecture. Experiments across six benchmarks demonstrate that GAUGE outperforms strong baselines across diverse incomplete-input settings. Furthermore, a Taylor remainder theoretical analysis characterizes the error of the first-order approximation relative to the exact counterfactual effect, establishing GAUGE as a principled and scalable framework for fine-grained evidence control under modality incompleteness.
[AI-70] StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
链接: https://arxiv.org/abs/2608.05587
作者: Linqiang Guo,Wei Liu,Li Gu,Yang Wang,Tse-Hsun(Peter)Chen
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended multimodal reasoning after each action, which is costly and poorly matched to the structured nature of GUI state transitions. We propose StepReflect, which formulates per-step GUI reflection as supervised structured prediction conditioned on explicit transition specifications and paired visual evidence. StepReflect is trained through a staged pipeline combining supervised fine-tuning, teacher-student distillation, and preference- and reward-based refinement. Offline, the resulting 8B model achieves 82.16% transition-level accuracy on AndroidWorld, exceeding zero-shot GPT-5.2 by 11.83 percentage points under the same structured input. Online, across M3A, Agent-SAMA, MAI-UI-8B, and Seed-2.0-Pro, StepReflect achieves higher task success in three of four agent configurations and remains within one successful task of the GPT-5.2 Reflection Agent in the fourth. It also reduces paid API charges relative to GPT-based reflection in all four configurations. These results establish StepReflect as a practical, locally deployable alternative to repeated frontier-model reflection for long-horizon mobile GUI agents.
[AI-71] he Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions
链接: https://arxiv.org/abs/2608.05583
作者: Hadi Hosseini,Samarth Khanna,Leona Pierce
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: Accepted at AIES 2026
Abstract:As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients’ own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-culpable patient. Compared to humans, LLMs also place greater emphasis on access to information, reducing responsibility judgments when health-risk knowledge is unavailable. These findings reveal that LLMs apply a systematically different moral framework than humans when responsibility and resource scarcity intersect, surprisingly often amplifying normative disagreement with humans as reasoning capability increases. Comments: Accepted at AIES 2026 Subjects: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2608.05583 [cs.CY] (or arXiv:2608.05583v1 [cs.CY] for this version) https://doi.org/10.48550/arXiv.2608.05583 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-72] SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agent ic Execution
链接: https://arxiv.org/abs/2608.05573
作者: Zhi Han,Chenxi Zeng,Liuhaichen Yang,Zihan Guo,Ming Zhou,Yang Li
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical. However, existing judge benchmarks often expose final responses or static trajectories, and rarely combine task-time skills with directly inspectable artifacts and environments. We therefore introduce SkillTV-Bench, a 681-case benchmark of real agent trajectories from 50 tasks across eleven domains, designed to evaluate skill-aware trajectory verification for both LLM-as-a-Judge and Agent-as-a-Judge methods. Additionally, we propose SkillTV-Evolve, which externalizes verification knowledge as a reusable JudgeSkill that guides an agent judge to plan targeted inspections and issue evidence-grounded verdicts. On a disjoint development pool, an automated evolution loop further refines the JudgeSkill using misjudged cases. On SkillTV-Bench, the refined skill increases the same agent judge’s accuracy by 14.8 percentage points. In offline rollout-pool selection, it increases selected-trajectory success from 22.9% with one rollout to 45.5% with ten rollouts. The code and data are available at this https URL
[AI-73] When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
链接: https://arxiv.org/abs/2608.05563
作者: Jialuo Chen,Lingqi Jiang,Xinhao Deng,Xiaohu Du,Jianan Ma,Yunhao Feng,Yuqi Qing,Zhihao Yuan,Linkang Du,Jingyi Wang
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注:
Abstract:Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instruction. We introduce PoisonedEvolution, a trajectory-poisoning attack on this promotion process. Our skill-visible black-box attacker can inspect a target skill and contribute bounded evidence, but cannot observe private pools or evolution logic or edit the skill bank. Artifact poisoning requires Inclusion, Evolution Attribution, and Realization. Attribution is the distinctive bottleneck: the target behavior must appear causally useful, recurrent, and generalizable before promotion. We evaluate four representative security-effect families using inert canary specifications. At 10% attacker support, across six mainstream LLM evolvers in SkillClaw, PoisonedEvolution embeds target behaviors in 546/600 trials (91.0% SER). On the structurally different Trace2Skill pipeline at the same ratio, it embeds target behaviors in 369/600 trials (61.5% SER), demonstrating transfer across evolution architectures. In a representative controlled study, three consistent attacker records suffice in a 30-record batch, whereas a single record is much weaker. Ablations identify recurring support, causal framing, and domain-aligned encoding as the main determinants of success. These findings expose evidence promotion as a security boundary for self-evolving agents.
[AI-74] Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
链接: https://arxiv.org/abs/2608.05541
作者: Yu Gu,Zhi Zheng,Yunpeng Ba,Xialiang Tong,Mingxuan Yuan,Zhenkun Wang
类目: Artificial Intelligence (cs.AI)
备注: 19 pages, 4 figures, 14 tables. Code: this https URL
Abstract:Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reasoning. However, directly applying ES to billion-parameter LLMs is highly ineffective. In such high-dimensional parameter spaces, most random perturbations are nearly orthogonal to useful update directions, leading to unstable optimization. We propose Hyper-ES, a subspace-based ES framework that avoids the weakness of ES in full-parameter search while exploiting its strength in low-dimensional optimization. Instead of asking ES to discover useful directions from random perturbations in the LLM parameter space, Hyper-ES first performs a small number of inexpensive gradient-based fine-tuning runs to obtain descent directions. Although each direction may provide only a limited improvement on its own, their span forms a compact adaptation subspace that captures useful reasoning updates. Hyper-ES then applies CMA-ES to optimize layer-wise DARE-TIES merging coefficients within this subspace, allowing ES to search over combinations of meaningful descent directions rather than over arbitrary full-model perturbations. We evaluate Hyper-ES on three Qwen2.5-Instruct and DeepSeek-R1-Distill backbones across six mathematical reasoning datasets. Results show that Hyper-ES consistently outperforms GRPO-LoRA by 1% while requiring 10% fewer space-consuming gradient updates. Code at this https URL.
[AI-75] Innovation-Residual Auditing of Autonomous Analysis Agents : Localization Detection Limits Error Control and Identifiability
链接: https://arxiv.org/abs/2608.05490
作者: Ahmed Hassoon,Mark Dredze
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML)
备注:
Abstract:Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step supervision. When such an analysis turns out to be wrong, someone must determine which operation caused it. A recent approach does this without any labelled mistakes, learning instead from analyses known to be sound and flagging operations that depart from what that model predicts; how reliable such audits are has not been studied. This paper supplies that analysis. The choice of score determines whether an error can be localized at all. If each operation is scored by how surprising it is given the operation immediately preceding it, then operations that merely inherit an earlier error are indistinguishable from correct ones, so one mistake produces one flag; scores computed against a longer reconstruction of the intended analysis instead spread a single mistake across many operations. We quantify how far they spread, and how to choose the comparison length when an error accumulates gradually rather than at once. We then give procedures that control the proportion of falsely flagged operations within a single audited analysis, requiring only that sound analyses be exchangeable rather than that the fitted model be correct, and we quantify how much the guarantees weaken when the model is imperfect or when the analysis was selected for review in a way that depends on its content. Finally we establish a limit on what any such audit can report: errors below a certain magnitude cannot be attributed at all, being indistinguishable from ordinary variation among sound analyses. This limit falls so slowly as more sound analyses are collected that at the representation sizes now in use a hundredfold increase reduces it by under two percent, so the dimension of the representation rather than the volume of training data is the binding constraint.
[AI-76] Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers
链接: https://arxiv.org/abs/2608.05472
作者: Zhen Zhang,Amr Alanwar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set. We study the consequences of this asymmetry for permutation-invariant set targets. We introduce the Transformation Degrees of Freedom (TDOF) of a target operator, a complexity measure counting the input-dependent directions an exact representation requires, and present a depth-separation analysis showing that context-rigid attention needs depth proportional to the target’s TDOF, whereas a single layer with a context-adaptive value family can represent the same target. Building on this analysis, we propose Matrix Zonotopic Attention (MZAttn), which replaces the fixed value projection with a context-adaptive matrix-zonotope family: a centre matrix plus a sum of generator matrices weighted by input-dependent gates. The construction reduces to standard multi-head attention at initialisation, preserves permutation equivariance, and admits a data-driven reachability interpretation. Experiments on a range of set-prediction tasks are consistent with the TDOF prediction that the architectural advantage is selective: it appears on targets that depend on the input set in a high-rank, sparsely combinatorial way, and is small on aggregate-statistic targets where parameter-matched standard attention is already competitive.
[AI-77] Recursive Synthesis for Long-Horizon Terminal Tasks
链接: https://arxiv.org/abs/2608.05466
作者: Zhongzhi Li,Yucheng Shi,Zongxia Li,Ruhan Wang,Anhao Li,Zixun Huang,Junyao Yang,Lei Ke,Ninghao Liu,Haitao Mi,Leowei Liang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct generation with large language models (LLMs) often breaks these dependencies. We present Recursive Synthetic Terminal Tasks (RST), a recursive verified synthesis framework for constructing long-horizon terminal-agent tasks at scale. Starting from verified seed tasks, RST extends the reference solution, realigns the verifier and instruction to the new workflow, validates the result in a fresh sandbox, and reuses accepted tasks as seeds for subsequent rounds. Across fifteen recursive rounds, RST produces 37,484 synthesized terminal-agent tasks at roughly \ 0.05 per task. Task difficulty increases substantially over rounds: the median reference solution grows from 67 to 374 lines, the median number of executed commands grows from 40 to 244, and DeepSeek-V4-Pro pass@4 drops from 90% at R_1 to 2.5% at R_15 . To demonstrate training utility, we collect rejection-sampled Qwen3.5 trajectories on the synthesized tasks and use them for supervised fine-tuning. Fine-tuning on these trajectories improves Qwen3.5-27B and Qwen3.5-122B-A10B by up to 10 points on Terminal-Bench~2, Terminal-Bench Hard, and Long-Horizon Terminal Bench, while agentic PPO lifts Qwen3.5-27B to 49.44%, 32.00%, and 22.07% on the three benchmarks, corresponding to relative gains of 20.0%, 41.2%, and 21.9% over the base model. Moreover, after 15 rounds, the recursion shows no ceiling: synthesis yield and validation rates remain stable as difficulty keeps climbing, indicating that the process can continue well beyond the scale reported here.
[AI-78] Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs
链接: https://arxiv.org/abs/2608.05455
作者: Zonglin Han(1),Yichen Chen(1),Jiawen Jiang(2),Tongan Shi(3),Kristian A. Stevens(1) ((1) Department of Computer Science, University of California, Davis, (2) International Digital Economy College, Minjiang University, (3) School of Computer Science and Artificial Intelligence, Liaoning Normal University)
类目: Artificial Intelligence (cs.AI); Data Structures and Algorithms (cs.DS)
备注: 11 pages, 1 figure, 1 table. Equal contribution among Y. Chen, J. Jiang, and T. Shi (alphabetical order)
Abstract:When a student must learn concepts connected by prerequisite dependencies, when does the order of instruction matter, and what does it cost to find the best one? We study instructional sequencing as a stochastic shortest-path problem in which attempting a concept succeeds with a state-dependent probability and failure leaves the learner state unchanged. We first prove that this stochasticity can be eliminated exactly: the problem collapses to a deterministic shortest-path problem on the lattice of prerequisite order ideals, preserving optimal values and actions. The collapse removes stochastic complexity but not combinatorial complexity: optimal sequencing remains NP-hard – via reduction from feedback arc set in tournaments – even with no prerequisite edges, unit costs, uniform binary nonnegative transfer, and success probabilities at least 1/2 . Hardness is not uniform: when realizable transfer preferences remain jointly acyclic with the prerequisites, any topological order of the residual joint graph is optimal, and fixed prerequisite width yields polynomial-time exact dynamic programming. A computable diagnostic, m\Delta , bounds the value of sequencing before optimization. On 70,893 interactions from an introductory CS course, the diagnostic certifies a doubly easy regime – little value to optimize and little space to search – while constructed transfer instances realize the challenging regime, where myopic sequencing suffers large regret yet exact A* with a consistent heuristic expands only linearly many states on that family.
[AI-79] SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
链接: https://arxiv.org/abs/2608.05439
作者: Yixuan Wang,Licheng Luo,Yu Fu,Kaidi Xu,Yue Dong,Mingyu Cai
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to capture the user’s intent, creating risks in safety-critical applications. Inspired by selective conformal prediction, we propose a selective translation framework that not only generates formal specifications but also determines when they can be trusted. Reliability is scored by two complementary black-box signals, the fidelity of the specification back-translated into natural language and the dispersion of repeated translations under exact semantic equivalence, which fail on different errors and jointly separate incorrect translations more sharply than either alone. Conformal risk control calibrates this score into a decision that accepts a specification or abstains, with a distribution-free bound on the rate at which incorrect specifications are accepted for execution, and a conformal anomaly detector on instruction embeddings screens out-of-distribution inputs before any translation is attempted. The proposed framework is general across formal specification languages, with experiments on Signal Temporal Logic (STL), Linear Temporal Logic (LTL), and geometric Spatio-Temporal Logic (SpaTiaL) demonstrating improved translation reliability, robustness under the evaluated cross-tier shifts, and effective uncertainty-aware abstention. This work establishes a foundation for trustworthy natural language interfaces by enabling AI systems to recognize when generated specifications may not be reliable.
[AI-80] Why the Third Axis Is Freedom
链接: https://arxiv.org/abs/2608.05423
作者: Michael Timothy Bennett
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注: Experiment code available on GitHub, in the papers folder: this https URL
Abstract:In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a model that produces one common answer can outperform a model retaining a broader repertoire. Explorative Modeling (XM) produces K outputs per comparison and updates on the closest, claiming exploration as a “third pretraining axis” associated with generative expressivity. Here I show the third axis is actually freedom, meaning the weakness of the constraint implied by a model’s behaviour. Previous work showed freedom is a property of function rather than form. Parameters, architecture, minimum-description-length (MDL), and data can vary while the behavioural constraint remains unchanged. It was formally proved that weakest models are likeliest to generalise, and freedom selection beat MDL by 110-500% in induction experiments. I prove average XM loss depends on the chance a candidate misses an acceptable region, with exploration raising miss probability to power K . For K1 , match probability rises with freedom. I then demonstrate empirically that XM optimises for freedom. In a Forward XM experiment, larger K increased or saturated measured freedom, and increased freedom at every tested value under context-dependent targets. I trained XM candidate pools and compared validation selection with a freedom selector that read unlabelled parent contexts. Freedom won in 29 of 30 cases. Generative expressivity is a mode-count proxy for freedom, that discards the extension structure that gives freedom its generalisation significance. XM is a means, freedom an end, and selecting for freedom improved XM under distribution shift.
[AI-81] Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples
链接: https://arxiv.org/abs/2608.05419
作者: Nilesh Kumar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpopulations where the correlation does not hold. Existing methods for identifying the affected samples without group annotations rely on signals from early training, which requires locating the epoch at which to intervene, a hyperparameter typically selected using group-labeled validation data. We show that a usable signal is available after convergence, when loss no longer distinguishes the two populations. Samples consistent with the spurious correlation are classified by a shared rule, while the remaining samples are fit through configurations specific to individual inputs and are correspondingly more fragile. Applying a fixed perturbation to a converged model’s inputs flips the predictions of the latter far more often than the former. The resulting procedure requires two forward passes per training sample, no group annotations at any stage, and no early-stopping epoch. Using the detected samples to rebalance training raises worst-group accuracy on Waterbirds from 57.3% to 80.8%, against 85.8% with ground-truth group labels.
[AI-82] Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation
链接: https://arxiv.org/abs/2608.05418
作者: Mackenzie Jorgensen,Jo Reilly,Alex Sutherland,Miri Zilka
类目: Artificial Intelligence (cs.AI)
备注: Accepted to AIES 2026
Abstract:AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixed-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias. We found that participants were broadly open to AI adoption, rejecting only three use cases outright, most notably recidivism risk assessment, where objections targeted the premise rather than the implementation. Our analysis reveals that foregrounding racial equity did not narrow the deliberation. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning, reminiscent of the curb-cut effect in inclusive design, highlights the benefit of incorporating the racial bias lens into the risk-benefit analysis of AI use cases from the outset.
[AI-83] Evaluating and Improving Pedagogical Fit in LLM -Based AI Tutors with the Pedagogical Suitability Index
链接: https://arxiv.org/abs/2608.05411
作者: Benjamin Barlog,Hudson Craig,Zedong Peng
类目: Artificial Intelligence (cs.AI)
备注: This paper has already been accepted and presented at the IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026) from July 31 to August 2, 2026, in Seattle, WA, USA
Abstract:Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In classroom learning, effective help depends not only on correctness, but also on whether a response matches the learner’s current foundation, the course sequence, and the timing of concept introduction. Existing evaluations focus mainly on answer quality, leaving this instructional fit under-measured. We present the Pedagogical Suitability Index (PSI), a composite metric of six theory-informed sub-scores that evaluates how well LLM-generated tutoring responses align with learner readiness and curricular progression, and we further use PSI as a structured feedback signal for response improvement. We evaluate four LLM tutors (ChatGPT, Gemini, Gemma4, and Qwen3) across 240 scenario-based evaluations using paired standard and defective prompts, then apply a PSI-guided regeneration protocol to 62 weak-performing cases. Baseline differences across the four tested models were modest overall (PSI range: 0.557 to 0.638), and open-weight and closed models did not exhibit a clear separation in pedagogical fit. Under the tested prompt perturbations, overall PSI remained largely stable (Delta = -0.002), though sub-score trade-offs emerged. More importantly, PSI-guided feedback substantially improved weak-performing cases: 51 of 62 cases improved (82.3%). Focused manual evaluation of the 62 PSI-selected weak cases provides initial evidence that the identified weaknesses are instructionally meaningful and that many PSI-guided regenerations correspond to human-judged improvement. These results suggest that learner- and curriculum-aware alignment may matter more for effective tutoring than model category alone, and that such alignment is both measurable and improvable.
[AI-84] C3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
链接: https://arxiv.org/abs/2608.05381
作者: Swapnanil Mukherjee,Agyeya Negi,Tanuja Ganu,Ponnurangam Kumaraguru
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C ^3 PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C ^3 PO’s paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C ^3 PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities’ structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature
[AI-85] Counterfactual Analysis via Large Language Models
链接: https://arxiv.org/abs/2608.05367
作者: Zonghao Yang
类目: Artificial Intelligence (cs.AI); General Finance (q-fin.GN)
备注:
Abstract:Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. This paper investigates the application of large language models (LLMs), specifically the GPT-3.5 model, for counterfactual analysis. We focus on the online lending context, where the counterfactual return on investment (ROI) is crucial for evaluating different interest rate schemes. We begin by assessing the predictive performance of GPT and comparing it with advanced machine learning algorithms. The results show that prompt engineering can significantly enhance GPT’s predictions, with the R-squared increasing from 1.97% to 2.84%, closely approaching the 3.48% achieved by gradient-boosted regression. Subsequently, we utilize GPT to generate counterfactual ROIs under a set of alternative interest rates. GPT exhibits logical coherence and causal reasoning in its responses. The findings underscore the potential of LLMs as effective tools for counterfactual analysis in online lending, suggesting broader applications for LLMs in various predictive and decision-making contexts.
[AI-86] CASCADE: An Agent ic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction
链接: https://arxiv.org/abs/2608.05359
作者: Jose A. Bird
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP. Prior work validates such tools by checking whether predicted genes are known cancer genes (membership); we instead test whether the predicted direction of change matches reality, using focal-gene copy-number amplification as a dosage-based proxy for the inverse of knockdown against real TCGA patient tumor data. For MYC, CASCADE’s predicted knockdown targets show strong concordance with real amplified-vs-non-amplified tumor expression across three cancer types (BRCA: 90.0%, COAD: 72.0%, STAD: 85.7%; all p0.0013), well above permutation baselines, surviving a PAM50 subtype control and replicating in an independent cohort (METABRIC, 87.2%). Compared against curated MSigDB gene-set baselines via Fisher’s exact test, CASCADE’s accuracy is not shown to exceed existing public knowledge of MYC- or E2F-driven biology, though its gene-specific direction-calling clearly outperforms a naive uniform guess. Extending to fifteen additional genes, validation proves gene-specific rather than universal: proliferation-machinery regulators mostly replicate, while lineage-identity transcription factors and one cyclin-D paralog (CCND2) consistently fail, a pattern we discuss as a hedged, post-hoc hypothesis. We separately benchmark whether an LLM-based agent correctly grounds natural-language requests into CASCADE’s real MCP tool calls. Across 35 queries, a documented local model reaches 71.4% exact match (85.7% for a larger model); schema and gene-alias failures are resolved by scale or server-side correction, but both models confidently default to the wrong perturbation type on ambiguous queries, a failure a targeted fix could not resolve because its trigger condition never occurs. Subjects: Artificial Intelligence (cs.AI) Cite as: arXiv:2608.05359 [cs.AI] (or arXiv:2608.05359v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.05359 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Jose Bird PhD [view email] [v1] Wed, 5 Aug 2026 19:30:44 UTC (67 KB) Full-text links: Access Paper: View a PDF of the paper titled CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction, by Jose A. BirdView PDFHTML (experimental)TeX Source view license Current browse context: cs.AI prev | next new | recent | 2026-08 Change to browse by: cs References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[AI-87] Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application
链接: https://arxiv.org/abs/2608.05346
作者: Marcos Carvalho,Fatih Temiz,Shavbo Salehi,Melike Erol-Kantarci,Daniel F. Macedo
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:
Abstract:Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). However, existing TSN scheduling solutions predominantly rely on static optimization techniques or centralized learning models that are based on fixed traffic patterns, limiting their effectiveness in dynamic environments. In practice, MEC environments often host multiple co-located XR traffic flows whose characteristics evolve over time, creating complex inter-queue dependencies that current schedulers fail to capture. Addressing these challenges requires adaptive, decentralized scheduling mechanisms capable of coordinating multiple TSN queues under varying traffic conditions. To this end, this paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent. The Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues. The simulation results demonstrate that the proposed approach reduces average frame waiting times by up to 26.8% and worst-case delays by approximately 16.8%, highlighting its effectiveness in dynamic XR-driven MEC scenarios.
[AI-88] Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks
链接: https://arxiv.org/abs/2608.05340
作者: Marcos Carvalho,Fatih Temiz,Shavbo Salehi,Melike Erol-Kantarci,Daniel F. Macedo
类目: Networking and Internet Architecture (cs.NI); Artificial Intelligence (cs.AI)
备注:
Abstract:Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (XR). However, the widespread adoption of XR introduces significant challenges due to co-located services in MEC environments, leading to contention for shared network resources. Moreover, XR traffic types have distinct characteristics and criticality in terms of timing requirements, further increasing the complexity and dynamics of such environments. Although reinforcement learning has shown promise for TSN scheduling optimization in dynamic network scenarios, existing approaches rely on centralized or high-level multi-agent designs and are typically tailored to periodic and predictable industrial traffic, limiting their applicability to XR workloads. As a result, these approaches suffer from (i) limited ability to capture inter-queue dependencies due to coarse-grained control, and (ii) poor adaptability to highly dynamic and heterogeneous XR traffic. To address these gaps, we propose a multi-agent reinforcement learning approach for queue-level XR traffic scheduling. We adopt the multi-agent transformer (MAT) to model inter-queue dependencies via attention over agents’ observations and actions, enabling implicit coordination across heterogeneous co-located XR applications. Our simulation results show that the proposed method outperforms baselines, achieving up to 71.42% latency reduction and up to 83.2% reduction in failure rate, while consistently achieving high reliability across all queues.
[AI-89] Hierarchical Server Architecture for Agent ic Science
链接: https://arxiv.org/abs/2608.05332
作者: Vanessa Sochat,Daniel Milroy
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI)
备注: 10 pages, 9 figures, 3 tables
Abstract:Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions. If assessing workload needs against environments is required for scheduling, automated discovery of resources is an essential step. In this paper, we present a hierarchical, dynamic architecture and software to discover resources across diverse cloud, edge, and HPC systems. The design enables concurrent, asynchronous negotiation, selection, and dispatch of requests for work using secretary agents. The agents probe and discover 51 real and simulated providers across 7 categories. We perform 19,973 negotiation and 6,952 selection simulations to assess reliability of decisions, demonstrating high (87.71%) negotiation accuracy and selection costs comparable to more traditional strategies. Designed for extensibility and currently supporting the Genesis Mission, this architecture exemplifies the importance of careful coordination between agents, discovery tools, and infrastructure for agentic science.
[AI-90] Agent ic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks
链接: https://arxiv.org/abs/2608.05266
作者: Nathan S Johnson,Ian Abshire
类目: Artificial Intelligence (cs.AI); Materials Science (cond-mat.mtrl-sci); Machine Learning (cs.LG)
备注: 20 pages, 8 figures
Abstract:Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and synchrotron beamlines. Research into agentic control of physical infrastructure is nascent and there are few well-established paradigms for how to engineer an agentic system. There are many choices to make when designing a microscopy agent, including the choice of LLM, the number of agents to use, agent responsibilities and delegation rules, retrieval-augmented generation parameters, and more. When designing and optimizing an agentic microscope controller, researchers not only want to ensure that the agent can correctly perform known tasks but also that the agent can generalize to new tasks that it has not encountered before. In this study, we develop a benchmark and trace-logging framework that reveals a) how different choices of agent architecture impact performance at microscopy tasks and b) the limitations of benchmarks for predicting if a particular agent will perform well on unseen microscopy tasks. The framework was used to evaluate one-, two-, and three-agent graph topologies, five LLMs, RAG and context parameters, and operational constraints across 53 microscopy benchmark tests. In total, 105 agent configurations, 1,949 individual test runs, and 49,109 RAG retrievals were recorded. Direct comparisons showed clear differences in latency, token use, cost, and failure mode between configurations. However, surrogate models trained on agent architecture and test results did not reliably predict an agent’s performance on new, unseen tasks. These results show that these benchmarks are useful for qualification, regression testing, diagnosis, and direct comparison, but the current heterogeneous test suite does not support a task-independent global configuration model.
[AI-91] OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes Recovery and Decomposition Quality
链接: https://arxiv.org/abs/2608.05263
作者: Yidian Chen,Yingzi Gu,Natan Vidra,Spurthi Setty,Sharon Zheng
类目: Artificial Intelligence (cs.AI)
备注: 8 pages, 4 figures
Abstract:Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown. OrchestraBench evaluates failure, recovery, and decomposition through a controlled, seed-reproducible failure-injection harness over templated enterprise workflows. It introduces cascade radius and per-failure-mode recovery as primary metrics and compares routing policies with bootstrap confidence intervals and paired tests. On a 26-case gold-labelled diagnostic, a keyword/flag router scored 0% on adversarial cases with misleading or missing surface flags, whereas an intent-reasoning model router scored 100%, matching the oracle. Controlled mechanism probes with a real Claude agent over a verifiable arithmetic dependency chain revealed three failure-handling tiers across five MAST modes: tool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0). This ordering persisted when the computation was reframed as a loan-approval workflow and across Sonnet, Opus, and Haiku, although absolute rates shifted with context. Blind retry reproduced latent faults and increased time to detection, indicating that detection and attribution are necessary for containment. Cascade radius increased with pipeline depth (mean 0.9 to 4.7 across depths 3-7). A trusted-state repair ablation showed that apparent containment gains primarily came from the trusted-state signal rather than autonomous detection. These results are controlled-chain mechanism probes, not domain-workload claims.
[AI-92] IMMENSE: Inductive Multi-perspective User Classification in Social Networks
链接: https://arxiv.org/abs/2608.05259
作者: Francesco Benedetti,Antonio Pellicani,Gianvito Pio,Michelangelo Ceci
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI)
备注:
Abstract:Online social networks increasingly expose people to users who propagate discriminatory, hateful, and violent content. Young users, in particular, are vulnerable to exposure to such content, which can have harmful psychological and social repercussions. Given the massive scale of today’s social networks, in terms of both published content and number of users, there is an urgent need for effective systems to aid Law Enforcement Agencies (LEAs) in identifying and addressing users that disseminate malicious content. In this work we introduce IMMENSE, a machine learning-based method for detecting malicious social network users. Our approach adopts a hybrid classification strategy that integrates three perspectives: the semantics of the users’ published content, their social relationships and their spatial information. Such contextual perspectives potentially enhance classification performance beyond text-only analysis. Importantly, IMMENSE employs an inductive learning approach, enabling it to classify previously unseen users or entire new networks without the need for costly and time-consuming model retraining procedures. Experiments carried out on a real-world Twitter/X dataset showed the superiority of IMMENSE against five state of the art competitors, confirming the benefits of its hybrid approach for effective deployment in social network monitoring systems.
[AI-93] Posture and Sustainment Optimization Under Adversarial Uncertainty
链接: https://arxiv.org/abs/2608.05256
作者: Amelie Norris,Alyssa Lee,Natan Vidra,Spurthi Setty
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational planning. Current practice relies on greedy heuristics that maximize value and ignore geographic coverage and are structurally vulnerable to adversaries that target high-strategic value locations. This paper presents a scenario-weighted adversarially robust posture optimization engine for the Posture and sustainability allocation (PSA) problem, modeled as a finite-horizon Markov Decision Process over assets, theater locations, and time steps. We introduce the Composite Expected Value (CEV) optimizer, which places assets by maximizing scenario-weighted expected posture efficiency over a distribution of threat scenarios, and the RobustCEV extension, which iterates against a Bayesian adversary that updates its targeting distribution in response to observed placement. Across three experiments in an Indo-Pacific basing environment with 20 assets and 5 theater locations, we demonstrate that: (1) the greedy baseline incurs a permanent 25.1% posture efficiency penalty due to geographic under-coverage and a 57.3% scenario-weighted readiness collapse under value-correlated adversarial threat; (2) the CEV optimizer recovers up to 19.8% efficiency over greedy when the threat distribution carries a geographic signal, with a curated set of 5 to 20 scenarios sufficient to capture the majority of this gain; and (3) the RobustCEV extension recovers up to 158% efficiency relative to a naive optimizer when an adaptive adversary employs a deceptive threat prior. All findings are validated using paired t-tests with Bonferroni correction and two-level variance decomposition, confirming that the performance gaps reported are structural properties of placement strategies rather than sampling artifacts.
[AI-94] An Emerging Retail Portfolio Management Application: Personalized Tax-Aware Reinforcement Learning with Natural Language Goals
链接: https://arxiv.org/abs/2608.05255
作者: Ramin Pishehvar
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
备注:
Abstract:Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted – existing robo-advisors use static, rule-based allocation, and institutional-grade systems require account minimums and technology stacks unavailable to individual investors. We present a fully built, integration-tested application that closes this gap: a FastAPI backend and web dashboard that let a user describe an investment goal in plain language (e.g. “I want steady growth but need to sell some shares next month for a down payment”), routes that goal to one of six investment mandates, and produces a live, broker-integrated portfolio recommendation from athree-phase reinforcement learning system – a self-supervised cross-asset encoder, a Mixture-of-Experts (MoE) allocation policy with a learned intent router, and a lightweight LoRA adapter that personalizes recommendations from an individual’s revealed brokerage behavior without retraining the shared model. The system is functionally complete and integration-tested end-to-end against a live brokerage API (Alpaca, paper-trading mode), including multi-user authentication, a trust first preview-before-apply confirmation flow, daily email digests, and an auditable action-integrity chain, but has not yet been opened to real end-users; we report this honestly as an emerging, pre-deployment application with a concrete path to full deployment, alongside 14-day walk-forward backtests (bootstrapped confidence intervals included) as preliminary, pre-deployment validation rather than production performance. We also report several practical engineering lessons – silently-inactive integration paths, hanging third-party API calls, and the value of end-to-end empirical verification over trusting checkpoint metadata – that we believe generalize to other applied RL systems built on external, live data sources.
[AI-95] PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
链接: https://arxiv.org/abs/2608.05249
作者: Xiaomin He,Dongling Xiao,Jiahao Xie,Ruiqi Lu,Qianle Wang,Zhongbin Guo,Wanxuan Sun
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
备注:
Abstract:Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still reduce instruction following to answering one self-contained question. We study this gap through \textbfrubric comprehension, which casts the model not as a generator measured against rubrics but as an \textbfexecutor that follows them: given an image and a typed, prioritized rubric, the model must verify each rule before producing an overall judgment. To support this setting, we propose \textbfPRISM, a four-stage data synthesis framework that produces persona–task pairs, prefix-guided rule sets, quality-filtered rubrics, and structured verification traces. We further introduce \textbfPRISM-Eval, whose Loose and Strict metrics use deterministic matching against fixed labels and therefore require no inference-time judge model. With only 10K synthesized samples, PRISM lifts Qwen3-VL-4B from 9.5% to 30.1% Strict accuracy on PRISM-Eval while preserving average performance on general benchmarks, and the gains transfer to four additional open-source MLLMs across dense and MoE architectures, suggesting that structured rubric supervision is a scalable path toward multi-rule, priority-aware multimodal instruction following.
[AI-96] LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
链接: https://arxiv.org/abs/2608.05246
作者: Jiahao Zhang,Yongzhi Tong,Zelin Fu,Pengde Zhao,Yanmei Jiang,Jiang Feng,Min Yang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.
[AI-97] Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
链接: https://arxiv.org/abs/2608.05245
作者: Muyang Ye,Tian Lan,Feihu Jiang,Yongshi Ye,Wuyunsiqin,Bin Zhu,Qianghuai Jia,Zhao Xu,Weihua Luo,Ye Wang,Jinyang Zhang,Longyue Wang,Lingfeng Bao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path toward self-evolution in expert domains. Existing self-evolving skill methods construct skills internally from the model’s parametric knowledge or trajectories, and are therefore bounded by what the model already knows. However, the domain conventions and standard procedures underlying professional skills often lie beyond this boundary and are hard to elicit from the agent alone. To address this issue, we therefore propose a novel framework, Search2Skill, that automatically identifies the agent’s capability gaps, searches external sources to address them, and distills the retrieved evidence into structured, reusable skills. Specifically, Search2Skill is optimized by a rubric-based reinforcement learning scheme that jointly improves when to search, how to search, and how to generate skills. Experiments on eight expert-level domains from three benchmarks show that Search2Skill consistently outperforms both search-augmented and trajectory-based skill-learning baselines under both streaming and held-out evaluation protocols. Further analyses show that the gains arise from skill abstraction rather than raw retrieved evidence, and that the acquired skills transfer across model scales.
[AI-98] Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
链接: https://arxiv.org/abs/2608.05243
作者: Duong Bach,Hai Nguyen Hong,Cuong Do
类目: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
备注:
Abstract:Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information. We show that this interpretation is incorrect. Matching only the marginal distribution places no constraint on the class-conditional distributions, allowing the latent style to remain highly predictive of the label despite appearing perfectly Gaussian in aggregate. We derive an exact decomposition showing that this mismatch is one of four conditions required for factorized sampling, and demonstrate that eliminating it is necessary but not sufficient to obtain the intended factorization. Empirically, our case-study model and four representative latent baselines achieve near-zero global MMD while still allowing a linear probe to recover class labels with 74%–100% accuracy (10% chance level). Our model reaches 99.15% clustering accuracy, whereas externally evaluated class-conditional generation succeeds only 16% of the time. This leakage remains under six independent perturbations involving model capacity, curriculum, prior geometry, and supervision across two datasets. Four mitigation strategies reduce probe accuracy to 21%–46%, although they leave within-class dependence largely unchanged. A post-hoc conditional prior improves externally evaluated class generation to 0.97 on MNIST without retraining but reaches only 0.41 on CIFAR-10, while an empirical style bank achieves 0.88 on CIFAR-10. These results demonstrate that no divergence computed solely on the marginal distribution of the style latent can certify independence from class labels, and that reporting marginal statistics alone does not verify the property commonly claimed in factorized generative models. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML) Cite as: arXiv:2608.05243 [cs.LG] (or arXiv:2608.05243v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.05243 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[AI-99] Project2Task: Graph-Guided Project-Level Planning for Autonomous Research
链接: https://arxiv.org/abs/2608.05225
作者: Huirui Xu,Runtao Xu,Shuo Ren,Jiajun Zhang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single topic. However, a research project is not merely a larger task: it is a long-horizon agenda that must be advanced through multiple bounded tasks with distinct but related objectives, parallel alternatives, and dependency-aware sequences. Existing single-task systems often treat the project as one oversized task, produce a flat set of vague or overlapping tasks, or leave task boundaries and execution order to manual coordination. We introduce Project2Task, a graph-guided project-level planning layer for autonomous research. Given a project brief, it represents candidate contributions as innovation atoms and organizes them in a directed lineage graph. A lightweight Bernoulli block-model objective selects among horizontal, vertical, and hybrid portfolio decompositions. Project2Task then generates bounded tasks with explicit contribution ownership, repairs overlaps and missing execution fields, and emits dependency-aware task contracts that specify objectives, inputs, expected artifacts, evaluation requirements, boundary constraints, dependencies, and execution order. The contracts are independent of any particular downstream research executor and support integration of task outputs into a coherent project-level result. On a benchmark of ten project briefs yielding roughly 30 tasks, manuscript-based portfolio evaluation gives Project2Task an average quality score of 7.15, compared with 4.58 for the Brief Baseline and 5.31 for the Topic-only Setting. Integrating its contracts with AutoResearchClaw increases average downstream task accuracy from 0.536 to 0.759. These results demonstrate the value of explicit project-to-task planning for producing coherent, non-redundant, and executable research-task portfolios.
[AI-100] Small Foundation Models of Human Cognition and Behaviour
链接: https://arxiv.org/abs/2608.05224
作者: Nick Oh,Fernand Gobet
类目: Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
备注: published as a conference paper at COLM’26
Abstract:Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels – task instructions, experimental stimuli, outcome feedback, and choice history – across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.
[AI-101] When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
链接: https://arxiv.org/abs/2608.05219
作者: Junzhuo Liu,Weiwei Li,Jun Ling,Peng Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student’s response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student’s preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state–reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student’s current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student’s current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from 0.746 to 0.865 on ALFWorld and from 0.574 to 0.693 on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at this https URL.
[AI-102] PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads ACM-MM2026
链接: https://arxiv.org/abs/2608.05218
作者: Ao Fu,Yi Zhou
类目: Artificial Intelligence (cs.AI); Sound (cs.SD)
备注: Accepted to ACM MM 2026
Abstract:3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth motion is often over-smoothed and may violate hard articulatory constraints such as bilabial closures, producing the notorious ``leaky mouth’’ artifact. A key difficulty is that brief, discrete articulatory events are inferred from a continuous acoustic embedding under a regression objective, which biases predictions toward averaged mouth configurations. While modern self-supervised speech encoders provide rich prosodic and phonetic cues, they do not provide an explicit, frame-aligned linguistic target that reliably disambiguates closure-level events. We propose \textbfPhoneme-Driven Gaussian Splatting (PD-GS), which augments a 3DGS talker with time-aligned phoneme tokens obtained from an automatic ASR and forced-alignment pipeline. Our core component, the \textbfLinguistic Fusion Module (LFM), adaptively fuses continuous audio context with discrete phoneme embeddings through a learned gate, allowing the model to preserve smooth audio-driven dynamics while strengthening phoneme guidance on articulation-critical segments. PD-GS is trained purely from monocular video using image reconstruction and lip landmark supervision. On HDTF, PD-GS achieves the best lip geometry among the compared baselines (LMD 2.66) and qualitatively reduces closure violations in challenging phoneme sequences, yielding more linguistically faithful neural avatars.
[AI-103] SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
链接: https://arxiv.org/abs/2608.05212
作者: Zhixiang Liang,Yifei Liu,Yidan Huang,Haozhe Zhao,Beichen Huang,Jiaqi Wang,Nan Duan,Qiong Cao
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers. Diagnosing such failures is difficult, requiring the manual inspection of extremely long execution traces, which could be beyond human capacity. We therefore introduce SearchAuditBench, a benchmark that evaluates whether LLM auditors can localize, attribute, and repair these failures, thereby reducing the human burden. SearchAuditBench comprises 1,243 failed trajectories, averaging 73.1 messages and 65.1K tokens, collected from eight open-weight models on five deep-search benchmarks, each expert-annotated with the critical error step, a search-specific root cause, and a reference repair with grading rubrics. We further propose SearchAuditor, a multi-perspective auditing framework that effectively localizes, attributes, and repairs search-agent failures through evidence-grounded adjudication. Experimental results show that even the strongest baseline, when powered by a frontier model like GPT-5.5, attains only a 26.6% end-to-end pass rate. In contrast, our SearchAuditor consistently outperforms all baselines across different frontier models, achieving an end-to-end pass rate of 32.3%, and resuming failed runs with its repairs enables agents to better recover from errors.
[AI-104] Otter: A Time-Aware History-Conditioned Human Chess AI
链接: https://arxiv.org/abs/2608.05206
作者: Tarun Kumar S
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather than treating each position in isolation. It combines two conditioning signals: (1) a move history encoder that conditions predictions on the last 20 moves, capturing opening preferences, positional drift, and intra-game behavioral tendencies; and (2) a time control module that modulates predictions based on clock pressure. Otter is trained on 6.1 billion positions from 117 million Lichess rapid games over 30 days on a single T4 GPU. Otter achieves 55.23% top-1 and 90.95% top-5 move-prediction accuracy, surpassing the prior state-of-the-art human chess model, Maia 2, with far fewer parameters and less training data. Across 11 Elo brackets (1100 to =2000), accuracy peaks at 57.38% in the 1900-1999 bracket. These results show that modeling chess as a time-aware, sequential activity yields more human-accurate move prediction than position-only approaches, using a smaller model. Code, trained models, and complete training logs are publicly released. Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG) Cite as: arXiv:2608.05206 [cs.AI] (or arXiv:2608.05206v1 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2608.05206 Focus to learn more arXiv-issued DOI via DataCite
[AI-105] Abstract Event Causal Rules: Induction and Application
链接: https://arxiv.org/abs/2608.05205
作者: Ziwei Zheng,Peiqiong Chen,Bang Wang
类目: Artificial Intelligence (cs.AI)
备注:
Abstract:Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making support and narrative comprehension. Nevertheless, existing instance-level causal pairs suffer severe generalization deficits on low-frequency long-tail and unseen event combinations. To address this limitation, this work proposes Abstract Event Causal Rule (AECR), a novel relation-level causal abstraction paradigm that transforms concrete cause-effect pairs into generalized abstract causal logic while retaining their intrinsic causal relationships. We design a multi-agent Concrete-to-Abstract Causal Induction (CACI) system coupled with similarity-constrained clustering to distill trustworthy AECRs from noisy raw causal data, based on which two complete AECR knowledge bases are built. To validate the practical utility of abstract causal knowledge, we propose an Abstract Rule-Guided Causal Attention Encoder (AR-GCAE), which injects the retrieved AECRs into the causality Graph Event Prediction (CGEP) benchmark task via rule-guided attention layers and gated representation fusion. Quantitative experimental results reveal that applying AECRs substantially strengthens the generalization capacity of event causal reasoning and brings consistent performance improvements to event prediction, with the most prominent gains observed on rare and unseen event samples.
[AI-106] SkillTrace: Multi-Trace Provenance Auditing for LLM -Agent Skill Reuse
链接: https://arxiv.org/abs/2608.05204
作者: Jialuo Chen,Minghe Wang,Lingqi Jiang,Jianan Ma,Xinhao Deng,Xiaohu Du,Ruixiao Lin,Yunhao Feng,Linkang Du,Jingyi Wang
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注:
Abstract:LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code, tools, references, and operational workflows. As skills become marketplace artifacts, auditing their reuse is no longer the same problem as ordinary code clone detection. Existing detectors target single-modality source code or whole-package similarity, yet skill reuse evidence is distributed across authored text, implementation fragments, and operational structure. As a result, they can miss reuse that preserves only one part of a skill. We present SKILLTRACE, a multi-trace provenance auditing framework for LLM-agent skill reuse. SKILLTRACE extracts three provenance traces: Expression, Implementation, and Operational. It represents the Operational Trace as a Skill Operational Graph (SOG) that captures activation, procedure, and resource-flow structure. An LLM assists only the Operational-trace extraction, once at ingestion; at audit time SKILLTRACE compares cached traces deterministically, calibrates each trace against same-function strict negatives, and reports which trace supports a reuse decision. On SKILLTRACE-BENCH, with 820 transformed reuse positives over 100 marketplace anchors and 751 negative controls, SKILLTRACE achieves AUROC 0.938 and F1 0.898. A 36,446-skill wild audit further shows that trace-attributed evidence surfaces actionable reuse review queues beyond repository-level baselines.
[AI-107] From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
链接: https://arxiv.org/abs/2608.05203
作者: Esra Zihni,Katryna Cisek,Hamzah Ziadeh,Hendrik Knoche,Robert Mikulik,John D. Kelleher
类目: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 9 pages, 2 figures
Abstract:Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians’ reasoning. Motivated by a clinician user study calling for clinical guideline-aligned cut-offs, we ask whether continuous predictors can be replaced by clinically informed categorical encodings without sacrificing performance. On a multi-centre European registry stratified into three treatment cohorts, we compare standard and fully categorised gradient-boosted models, the latter using stroke guideline-aligned, treatment-specific thresholds. The fully categorised models are statistically indistinguishable from their continuous counterparts in two of the treatment cohorts, with a significant drop in predictive accuracy in one cohort. Global feature importance rankings remain consistent, suggesting that discretising continuous predictors into guideline-based categories preserves the core hierarchy of prognostic factors across all treatment groups. Guideline-based categorisation is thus a viable design choice for stroke-outcome models.
[AI-108] ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design Evaluation and an OpenClaw Case Study
链接: https://arxiv.org/abs/2608.05201
作者: Siyuan Li,Peng Shu,Churan Yu,Peilong Wang,Ruidong Zhang,Bowen Guo,Xinliang Li,Ruiyu Yan,Arif Hassan Zidan,Yi Pan,Wei Ruan,Lifeng Chen,Junhao Chen,Zhaojun Ding,Yiwei Li,Zhengliang Liu,Haixing Dai,Lin Zhao,Yu Bao,Xiang Li,Wei Zhang,Tianming Liu
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 40 pages, 4 figures, 6 tables. Introduces and empirically evaluates the ASTELD six-axis classification framework across eight autonomous AI agent platforms, with OpenClaw as an in-depth case study
Abstract:Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet the field lacks a common classification scheme for comparing these design choices. We propose ASTELD, an operational six-axis classification framework for autonomous AI agents: Architecture pattern, Security posture, Tool integration model, Execution paradigm, Level of autonomy and human control, and Deployment topology. ASTELD is constructed by synthesizing prior agent taxonomies with observable platform properties and explicit category-assignment rules. We evaluate its discriminative and explanatory utility by mapping eight representative frameworks and by using OpenClaw as an in-depth case study. The resulting profiles separate all eight platforms under their dominant configurations and reveal three cross-platform patterns: a security-accessibility diagonal, strong execution-architecture coupling, and capability convergence with persistent architectural differentiation. We further classify 50+ OpenClaw derivatives and find that innovation concentrates on the Security, Execution, and Deployment axes, indicating that ASTELD can explain where ecosystem fragmentation occurs. The OpenClaw case study also supplies a six-category vulnerability taxonomy, evidence from five institutional assessments, and adoption and governance analyses that connect platform coordinates to observed risks. These results position ASTELD as a reproducible method for comparing agent platforms, identifying unoccupied design regions, guiding framework selection, and organizing future empirical research. The analysis also exposes a consequential empty region: none of the evaluated systems combines local-first deployment with enterprise-grade security.
[AI-109] Post-Hoc Trajectory-Risk Certification for Modular LLM -Based Security Agents
链接: https://arxiv.org/abs/2608.05199
作者: Zhenpeng Li
类目: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
备注: 18 pages, 11 tables; preprint
Abstract:Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Split conformal prediction gives each stage finite-sample coverage, but deployment requires a trajectory-level guarantee across the full chain. These guarantees do not compose automatically when stages are independently trained and calibrated. Bonferroni allocation is distribution-free but conservative under correlated errors. We show that a natural pairwise-correlation extension to three or more stages is invalid because it gives a lower rather than an upper bound, and derive a valid spanning-tree alternative. We distinguish whether stages are dependent from whether an audit sample is large enough to certify that dependence, and give matching upper and information-theoretic lower sample-complexity bounds. We also show that coarse-to-fine label selection can create near-perfect measured correlation without learned dependence. On a two-stage intrusion-detection pipeline across 6 open LLMs and 2 datasets, removing this artifact reduces measured correlation from near 1 to 0-0.78. A direct audit of trajectory failure becomes 13.7% tighter than Bonferroni once the audit reaches the required sample size, but is worse when undersized. A modular certificate using per-stage certificates and a pairwise overlap bound yields a positive average gain of 0.6%, quantifying the cost of lacking joint access. Same-model, cross-model, and permuted-pairing tests show that residual dependence reflects shared sample difficulty, not shared model representations. Average trajectory coverage across 12 configurations is 92.7% +/- 2.4% at alpha = 0.10. Under cross-dataset deployment, single-step miscoverage reaches 100% even when accuracy remains 78%, showing that distribution shift destroys calibrated confidence before raw accuracy. Comments: 18 pages, 11 tables; preprint Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) MSC classes: 68M25, 68T05, 62G15 ACMclasses: C.2.0; I.2.6; G.3 Cite as: arXiv:2608.05199 [cs.CR] (or arXiv:2608.05199v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.05199 Focus to learn more arXiv-issued DOI via DataCite
[AI-110] Automatic Detection of Deaths from Social Networking Sites
链接: https://arxiv.org/abs/2608.05183
作者: Nuhu Ibrahim,Riza Batista-Navarro
类目: ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI)
备注:
Abstract:This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media content, and reported machine learning (ML) classifiers that achieved high performance in automatically detecting deaths of social networking site users from posts associated with their profiles. A new dataset was developed using Wikidata and Twitter. ML models, both traditional (RF, KNN, LR, and SVM) and deep learning (BiLSTM, CNN, and the state-of-the-art BERT), were trained on features extracted using TF-IDF and pre-trained embeddings (Glove, Word2Vec, and FastText) to classify post-mortem content from its pre-mortem counterpart. The results showed that RF outperformed all other traditional ML models; BiLSTM outperformed CNN; TF-IDF consistently outperformed pre-trained word embeddings for the traditional models; Word2Vec consistently outperformed Glove and FastText for the deep learning models; and BERT outperformed all other models. It was found that although pre-mortem and post-mortem tweets express similar levels of positive sentiment, post-mortem tweets exhibit higher negative sentiment, whereas pre-mortem tweets exhibit higher neutral sentiment. Feelings suggesting negativity (sad, angry, surprise, and fear) are more dominant in post-mortem tweets, while happy is more dominant in pre-mortem tweets. It was also found that words, personal pronouns, verbs, family words, religious words, death words, and swear words occur more frequently in post-mortem tweets, whereas impersonal pronouns and informal words occur more frequently in pre-mortem tweets. Additionally, analytical thinking is expressed more in post-mortem than pre-mortem conversations. This experiment’s significant contribution is the successful development of an exceptionally high-performing technique for automatically detecting user deaths on social networking sites.
[AI-111] Autonomous Research Agents : A Survey of AI Scientists and the Verification Gap
链接: https://arxiv.org/abs/2608.05179
作者: Tianyu Ding,Aditya Nannapaneni,Bingfan Liu,Ling Zhang
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review. End-to-end AI scientist systems can now produce paper-like manuscripts, but their claims are often harder to verify than their code is to run. This survey studies that gap in computational AI/ML research, where code, benchmarks, experiments, and write-ups are most visible. We screen 125 candidate works and include 35, with full-text coding of 26 entries: 24 runnable systems and two study or position papers. We code seven audit dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human-in-the-loop points, novelty verification, and result-selection disclosure. The main pattern is that code release is now common, but reproducibility-grade and claim-verification artifacts remain much less common. In the 24 runnable systems, 83 percent release code, while 38 percent release seeds or execution traces and 38 percent report any novelty-verification method. Among nine closed-loop L4 systems, seven are mechanical reruns and one is author-claimed without an external check; no LLM-era system in the corpus demonstrates an externally validated in-loop oracle under our coding rule. We contribute a coded corpus, a lifecycle-by-autonomy map, an auditability-gap analysis, and a reviewer-facing reporting checklist. The survey argues that the field’s central bottleneck is no longer only whether agents can complete research tasks, but whether reviewers can verify the claims those agents produce.
[AI-112] Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios
链接: https://arxiv.org/abs/2608.05178
作者: Nouar AlDahoul,Hezerul Abdul Karim,Myles Joshua Toledo Tan
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from the Global South due to pro-equity bias that results from equity-focused safety alignment, open-weight and small models frequently flip this preference to favor the Global North, reflecting the global region bias and unaligned geographic distribution of their baseline pre-training data. Our findings highlight how normative assumptions embedded in model behavior can shape gatekeeping decisions, underscoring the importance of auditing AI systems for fairness and value alignment.
[AI-113] Challenges for Musical Education in the Age of AI and Digital Transformation
链接: https://arxiv.org/abs/2608.05176
作者: Jean-Pierre Briot
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 14 pages
Abstract:Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what they teach, how they teach it, and why. We now stand at what may be the most consequential of such turning points. Three deeply intertwined transformations have been converging simultaneously: 1. The very nature of music has changed: how it is made, distributed, consumed, and valued; 2. The public for music has changed: listening habits are now shaped by streaming algorithms and the boundary between consumer and creator has blurred; 3. Music-making itself has changed: digital audio workstations (DAWs) have for two decades been reshaping compositional practice. In addition, generative AI has now irrupted, capable of producing complete, stylistically coherent musical pieces from a short text prompt. These changes are not independent of one another, and they all bear directly on musical education - both the content that must be taught, and the pedagogical tools and methods available to teach it. This paper attempts to map these challenges and to consider how education might adapt. Section 2 surveys the changes in various aspects of music (nature, production, public, economics). Section 3 examines the implications for education before Section 4 concludes.
[AI-114] he Closing Window: How Governments Could Lose Their Ability to Restrain Advanced AI
链接: https://arxiv.org/abs/2608.05173
作者: Peter Barnett
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may eventually conclude that these risks warrant restraining AI development. This motivates the question: will governments still be able to restrain AI development in the future, should they want to do so? In this paper we analyze which world events and changes to the state of AI development would make future governance more difficult or even effectively impossible. Our analysis surfaces likely pathways that would lead to these difficulties, including hardware proliferation, continued algorithmic progress, and the release of catastrophically dangerous AI models. Due to the field’s lack of understanding of AI development, it may be difficult or impossible to know when we will hit a “point of no return”, and we therefore recommend a conservative approach. The window may be closing, but governments currently have an opportunity to preserve their optionality if they act soon. Our policy recommendations would enable governments to restrain AI development in the future, while imposing relatively small costs today.
[AI-115] Estimating time spent on work tasks
链接: https://arxiv.org/abs/2608.05172
作者: Stephane Hatgis-Kessell,Tomás Aguirre,Alexander Wan,Rishi Bommasani
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate to occupation-level effects. We study how tasks should be weighted in this aggregation. Prior work has relied on idiosyncratic or ill-justified choices for task weights. While recent work suggests weighting tasks by time spent, existing time shares are either based on coarse ONET data not intended for this purpose or estimated via black-box language models. We address this gap by proposing a principled method for estimating time shares for nearly 18,000 tasks that constitute nearly all U.S. jobs. Our estimates factor a task’s time into (i) the expected frequency of the task, derived from ONET, and (ii) the time to complete a single instance of it. To estimate the latter, we solve a constraint satisfaction problem based on pairwise comparisons elicited from language models about which tasks are longer per instance. We validate our estimates by characterizing the solution space of the constraint satisfaction problem and collecting data from workers for multiple occupations. We apply our time shares to analyze how AI exposes U.S. occupations and find that some prior results are sensitive to time weights. Accounting for the share of working time exposed to AI, rather than the share of tasks like prior work, widens the gap between the least and most exposed jobs: it lowers measured exposure for most occupations but raises it for the most exposed. Re-weighting by time also reshuffles 11 of the 25 occupations widely reported as most exposed to AI, shifting the top of the list away from clerical work and toward analytical roles. Time shares can serve as a general primitive for research and policy on the labor economy and the economics of technology.
[AI-116] Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving
链接: https://arxiv.org/abs/2608.05171
作者: Jiaxin Zou,Xiaoming Zhai,Chunlei Gao
类目: Computers and Society (cs.CY); Artificial Intelligence (cs.AI)
备注:
Abstract:Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in two conditions: with and without GAI. Chi-square analysis showed significant differences in CPS behavior distributions between groups. Compared with the control group, the GAI-supported group demonstrated more social behaviors, particularly engagement and conflict management, but less frequent cognitive behaviors such as task planning and solution reasoning. Lag sequential analysis further revealed distinct interaction patterns: while the control group followed a more conventional transition from listening to planning, the GAI group showed a robust pathway from task planning to conflict management to solution reasoning. Thematic analysis of AI interaction logs suggested that students used GAI as an epistemic arbiter, drawing on AI-generated facts and visualizations to resolve disagreements constructively. These findings suggest that GAI reshapes CPS by mediating the transition from social conflict to collaborative reasoning.
[AI-117] Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling
链接: https://arxiv.org/abs/2607.27698
作者: Yuan Tian,Yi Mei,Mengjie Zhang
类目: Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
备注:
Abstract:In dynamic multi-mode project scheduling, activities have alternative execution modes and uncertain durations, while precedence relations and limited resources constrain their execution. Heuristic priority rules support fast online decisions, but their design requires substantial domain expertise. Genetic programming (GP) hyper-heuristics can automatically evolve such rules. Large language models (LLMs), meanwhile, provide a flexible interface for interpreting scheduling information and explaining decisions. However, zero-shot LLM decisions may lack domain knowledge, consume many tokens, and vary across repeated queries. GP-evolved rules therefore provide a potential source of scheduling knowledge for guiding LLM decisions. Unlike existing LLM–GP hybrids that use LLMs to support heuristic evolution, we transfer knowledge in the reverse direction, using knowledge extracted from high-quality GP rules to guide an online LLM decision maker. We extract knowledge from high-quality GP rules and inject it through Feature Selection, Feature Hint, Rule Reference, and Rule Follow. These mechanisms are evaluated in terms of scheduling performance, token consumption, decision stability, and the feature focus expressed in generated rationales. GP-derived guidance generally improves the unguided LLM, but its representation matters. Simplifying the decision context or supplying explicit decision logic is more effective than highlighting important features. Feature Selection offers the best token efficiency, whereas Rule Follow achieves strong performance at greater token cost. Guidance also improves decision stability and changes the features expressed in generated rationales.
[AI-118] he ethics of artificial intelligence in the life sciences: Universality cultural diversity and an architecture of care
链接: https://arxiv.org/abs/2608.05436
作者: Jean-Pierre Changeux,Gustavo Deco,Morten L. Kringelbach
类目: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI)
备注:
Abstract:The life sciences and health research have started to benefit from artificial intelligence, which raises ethical concerns that are real but, we argue, not special. Any science should be governed by values that rest on how the human brain is built and socialised rather than anything distinct to artificial intelligence. Importantly, the human brain has a different, much less costly computational architecture than these machines. This is achieved through the orchestration of a global neuronal workspace, and through reward best described not as a quantity to be maximised but as a continuous cycle of wanting, liking and satiety. As such, this creates the deep tension running through the ethics of the human person, between the universality of ethical judgement and the diversity of morals. The brain networks of the global workspace and emotion are universally shared, but the diversity of content is shaped by epigenetic appropriation of the particulars of the physical, social and cultural world, which makes every person unique. Still, if we were to build machines on these principles rather than the present unaffordable reward maximisers, the question of their governance would change from restraint to upbringing. We set out the institutions such a future would require, together with the questions that remain open.
[AI-119] One Qubit Can Beat One Bit: Quantum Advantage for Post-Training Quantization
链接: https://arxiv.org/abs/2608.05240
作者: Yuma Ichikawa,Moeto Mishima
类目: Quantum Physics (quant-ph); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
备注: 40 pages, 6 figures
Abstract:One-bit post-training quantization represents each weight using only its sign, requiring all deployment contexts to share the same binary weight matrix even when their activation statistics favor different sign patterns. We study this shared-sign constraint and introduce Quantum Random Access Quantization (QRAQ). This framework encodes context-dependent signs in a quantum random-access code and retrieves them via context-matched Pauli measurements. Under an explicit fresh-copy logical readout model, QRAQ produces an unbiased, context-specific binary surrogate with a tractable shot-noise penalty. We prove a row-wise separation from shared-sign one-bit PTQ with signed per-row scales. When the optimal context-wise signs are incompatible, QRAQ achieves a strictly lower ideal reconstruction risk. We also derive finite-shot and calibrated-noise conditions under which this separation is retained. Fixed-readout quantum schemes are classically simulable, so the relevant resource in this model is measurement incompatibility rather than quantization alone. Finally, we characterize the role of scale granularity, provide finite-sample certificates, and evaluate the predicted ideal, finite-shot, noisy, and multi-context regimes in simulator experiments.
[AI-120] Quality Diversity for Reliable Data Driven Time-Use Optimization PPSN2026
链接: https://arxiv.org/abs/2608.05230
作者: Aneta Neumann,Ty Stanford,Dorothea Dumuid,Frank Neumann
类目: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
备注: To appear in the Proceedings of Parallel Problem Solving from Nature (PPSN 2026)
Abstract:The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. While predictive models can estimate the relationship between time-use compositions and health outcomes such as body mass index, life satisfaction, and cognition, most optimization approaches focus only on maximizing expected benefit and do not consider the uncertainty inherent in data-driven prediction. Ignoring uncertainty in health-related decisions can lead to unrealistic time-use recommendations. To address this gap, we introduce an uncertainty quantification Quality Diversity (QD) framework for a more reliable time-use recommendation. Objective functions are derived using compositional data analysis using a large child cohort dataset n 1000, to capture the relationship between daily activity compositions and multiple health indicators. We develop a new approach that incorporates predictive uncertainty into QD processes and produces more reliable recommendations that balance the expected health benefits with the confidence of the model. We explore the solution space through variable-based and objective-based behavioral representations, revealing diverse high-quality time-use composition and explicit relationships between health outcomes under uncertainty. By embedding uncertainty directly into optimization, our framework shifts the time-use recommendations toward regions of lower uncertainty while preserving high-quality structures for more reliable decision-making in behavioral health.
机器学习
[LG-0] On-Policy Self-Distillation without Any Supervision
链接: https://arxiv.org/abs/2608.06296
作者: Yijiang Li,Bingyang Wang,Yijun Liang,Yunjie Tian,Di Fu,Nuno Vasconcelos
类目: Machine Learning (cs.LG)
*备注:
Abstract:On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine “self”-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model’s own generations via internal consistency. We propose Unsupervised On-Policy Self-Distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseudo-solution by majority vote under a self-consistency threshold. It then conditions a teacher distribution on the shortest pseudo-solution and distills it into prefixes of the model’s longest incorrect completion, allowing the model to correct itself precisely where it is confidently wrong. Across diverse benchmarks, base models, and training settings, U-OPSD consistently improves over the base models and matches or surpasses supervised methods with ground truth (GT), such as OPSD and GRPO. On AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD improves over the base model by 8.5% and 10.7% on Qwen3 non-thinking mode at the 4B and 8B scales, respectively, and outperforms OPSD by an average of 3.2% and 2.3%. In thinking mode, U-OPSD remains on par with OPSD, outperforming it by 0.9% at 4B and matching it at 8B, while surpassing GRPO by 0.7% and 1.1%, respectively.
[LG-1] Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data
链接: https://arxiv.org/abs/2608.06288
作者: Lev V. Utkin,Stanislav K. Kogan,Andrei V. Konstantinov
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:This work presents a novel attention-based framework for estimating the Individual Probability of Treatment Benefit (IPTB) in survival analysis contexts. The proposed model, called Surv-IPTB, directly quantifies the probability that a specific patient will experience extended survival time under treatment versus control. We reformulate IPTB estimation as a binary classification problem, leveraging pairwise patient comparisons across treatment and control cohorts. The framework incorporates a principled handling of right-censored observations through imprecise probability representations, where uncertain treatment effects are characterized by interval-valued probabilities. An attention mechanism with learnable query-key transformations enables flexible, data-driven aggregation of pairwise comparisons, while simultaneously learning soft class probabilities for censored cases. Through extensive experiments on synthetic datasets with complex nonlinear structures, including spiral, bell-shaped, and circular feature spaces, we demonstrate that our approach maintains robust performance across varying censoring rates and treatment effect strengths. The model consistently outperforms meta-learner baselines (T-learner and S-learner) equipped with random survival forests, Cox proportional hazards, and Beran estimators, particularly in challenging nonlinear scenarios where conventional methods exhibit significant degradation. The results establish the proposed attention-based framework as a scalable and statistically principled solution for personalized treatment benefit assessment in survival settings. The code implementing the model is publicly available.
[LG-2] he Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
链接: https://arxiv.org/abs/2608.06283
作者: Iosif Lytras,Nikolaos Makras,Sotirios Sabanis
类目: Machine Learning (cs.LG); Optimization and Control (math.OC); Probability (math.PR); Machine Learning (stat.ML)
*备注: 53 pages
Abstract:We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.
[LG-3] Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction
链接: https://arxiv.org/abs/2608.06262
作者: Zonghuan Xu
类目: Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注: 18 pages
Abstract:Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space \mathcalX with |\mathcalX|=N . We first ask which pairs of distribution classes can be reliably distinguished. We then ask how many additional queries are required to match an adaptive tester when all queried events must be fixed in advance. We show that learnability holds if and only if the two classes have positive separation in their pairwise conditional probabilities. When this separation is zero, the optimal worst-case error is exactly 1/2 at every finite query budget. For any T -query adaptive policy and any \rho \in (0,1) , we construct a randomized non-adaptive procedure using O(N^2(T + \log(1/\rho))) pair queries chosen before any response is observed. Its simulated transcript is within \rho in total variation of the adaptive transcript, uniformly over all distributions in the model. We also construct a matching family with constant adaptive query complexity and \Omega_\varepsilon(N^2) non-adaptive query complexity. Consequently, the worst-case fixed-error adaptivity gap is \Theta_\varepsilon(N^2) . Thus interaction can reduce the required number of tests by a quadratic factor, but the apparent exponential branching of an interactive evaluation does not yield an exponential query advantage.
[LG-4] RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction
链接: https://arxiv.org/abs/2608.06259
作者: Yiting Zheng,Cheng Fang,Anthony Donofrio,Haote Li
类目: Machine Learning (cs.LG)
*备注: 8 pages, 6 figures
Abstract:Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture chemical transformations, making accurate prediction difficult for reactions with complex substrates. We propose reaction contrastive learning foundation (RxnCLF), a self-supervised contrastive framework for reaction representation learning. RxnCLF is built on a condensed reaction graph (CRG) that unifies reactant and product information into a single graph, enabling the model to learn explicit and enriched transformation structure rather than disconnected graphs. Pretrained on 1.7 million Pistachio reactions, RxnCLF learns a compact and continuous latent space that captures both reaction-center features and broader side chain contexts, making it transformation-aware and chemically interpretable. Fine-tuned on multiple yield prediction benchmarks, including Buchwald-Hartwig, Pd-catalyzed BH coupling, and proprietary HTE C-N coupling and amide formation datasets, RxnCLF consistently outperforms graph and sequence-based baselines, improving R2 and achieving the best performance overall. Our results highlight the promise of CRG-based RxnCLF as a scalable reaction foundation model, with the potential to generalize across broader reaction spaces and support diverse downstream reaction informatics tasks, including regioselectivity prediction, enantioselectivity prediction, and reaction condition optimization.
[LG-5] MetaboLLM : a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction
链接: https://arxiv.org/abs/2608.06253
作者: Dohyun Ku,Min Gu Kwak,Francisco J. Pasquel,Jing Li
类目: Machine Learning (cs.LG)
*备注: 60 pages, 3 figures, 16 tables; includes Supplementary Information
Abstract:Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledge, relational, and description tasks, and transferred to an external public benchmark. MetaboLLM-GIN achieved the highest AUC for stress hyperglycemia prediction after coronary artery bypass grafting (0.8616) and postmenopausal hormone-regimen classification (0.8123), outperforming conventional models, alternative graph constructions, and graphs generated from unadapted or non-retrieval LLM configurations. Model interpretation further produced biologically meaningful findings in both applications. These results show that domain-specialized language models can organize heterogeneous biochemical knowledge into predictive and interpretable metabolite graph representations.
[LG-6] A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance
链接: https://arxiv.org/abs/2608.06246
作者: Fardin Afdideh,Fernando Seoane,Farhad Abtahi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-efficient adaptation, alignment, retrieval augmentation, model editing, unlearning, calibration, and Multimodal Instruction Tuning. However, the literature remains fragmented across technique families, model classes, and deployment contexts, making it difficult to compare methods or describe how a trained model has been modified. This survey synthesizes the post-training adaptation literature and introduces a six-dimensional taxonomy organized by mechanism, goal, data requirement, persistence, structural scope, and model type. The taxonomy distinguishes commonly conflated terms such as fine-tuning, retrieval augmentation, and prompting, and shows how adaptation strategies evolve from traditional machine learning through deep learning, foundation models, large language models, and multimodal large language models. It also maps relationships among techniques, including inheritance, supersession, hybridization, and layered deployment stacks. The resulting vocabulary can support technical documentation, model-change tracking, and governance analysis. The survey concludes by identifying open challenges in evaluation, reproducibility, persistent inference-time adaptation, unlearning, multimodal adaptation, and governance-aware post-training workflows.
[LG-7] mestep-Conditioned Transformers for Global Weather Forecasting
链接: https://arxiv.org/abs/2608.06241
作者: Sam Levang,Fran Bartolic,Ty Dickinson,Chase Dwelle,Paulius Rauba,Viktor Cikojevic
类目: Machine Learning (cs.LG); Operating Systems (cs.OS)
*备注:
Abstract:Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorter timesteps (e.g. 1 to 6 hours) finely resolve atmospheric dynamics within the diurnal cycle but increase error accumulation for a given forecast horizon, while longer timesteps (e.g. 24 hours) reduce error accumulation but limit the usability of short-range forecasts where sub-daily predictability is high. In this work, we present GEM-3, a probabilistic global weather model that addresses this trade-off through explicit multi-timestep inference. With a single set of trained weights, the model timestep can be configured at inference time to balance predictability and usability across a broad forecast horizon. Additionally, we find that mixed-timestep training consistently improves rollout stability relative to timestep-specialist models. Under the hood, GEM-3 is a lightweight neighborhood-attention transformer with ~134M parameters on an equirectangular grid with a number of architectural advancements beyond its predecessor GEM-2. The result is a practical forecasting system that couples near-SOTA medium-range probabilistic skill, stable extended-range rollouts, efficient training and inference, and decision-relevant diagnostics.
[LG-8] SAGA: Score-Weighted Adaptive Generation Alignment for Low-Resource Nordic Language Models
链接: https://arxiv.org/abs/2608.06179
作者: Hoda Fakharzadehjahromy,Emil Wiman,Andreas Bueff,Hafsteinn Einarsson,Fredrik Heintz
类目: Machine Learning (cs.LG)
*备注: 18 pages, 7 figures
Abstract:Preference optimisation has proven effective for improving large language models but typically relies on costly human preference annotations. Extending these methods to morphologically rich, low-resource languages remains challenging because such annotations are scarce. We present SAGA (Score-weighted Adaptive Generation Alignment), a parser-guided preference optimisation framework that replaces human labels with dependency-parser supervision. SAGA converts parser judgements into preference pairs for delta-DPO, combines parser quality with lexical diversity in a composite reward, filters low-information pairs using a reward-gap criterion, and monitors reward hacking to maintain reliable supervision. Across Danish, Icelandic, and Norwegian Bokmål using GPT-SW3-1.3B, SAGA consistently improves grammatical quality without requiring human preference labels. Danish parse success increases from 69.0% to 93.8%, Icelandic achieves a +4.5 percentage-point improvement on an independent Stanza evaluation (three-run mean +3.3 percentage points) while native speakers prefer SAGA outputs in 80% of pairwise comparisons, and Norwegian Bokmål improves by +28 percentage points. These results demonstrate that parser-derived supervision is a practical alternative to human preference annotation for grammatical alignment in low-resource languages where high-quality dependency parsers are available.
[LG-9] hreshold-Based Early Stopping of Accumulations in Neural Networks with Binary Activation
链接: https://arxiv.org/abs/2608.06177
作者: Quentin Luquet de Saint-Germain,Massil Ait Abdeslam,Jean Pierre David
类目: Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注: 9 pages, 3 figures, 1 table
Abstract:Binary neural networks are very attractive for constrained deployment, enabling small footprint and low-power inference. For binary activations, the dot products become sign-controlled additions or subtractions, but the number of operations is unchanged. Indeed, every neuron or output channel still accumulates all of its input, even though only the sign will be retained, which is often wasteful. As the accumulation progresses, the running partial sum frequently drifts so far from zero that its final sign becomes highly predictable long before the last term is reached; every contribution evaluated after that point changes the value of the sum but not the final output activation. This paper turns this observation into a post-training early-stopping mechanism. We characterize the behavior of the running accumulations on the training dataset and use this information to predict the final sign as soon as possible. No model parameter is retrained. We count the number of operations under an idealized ordering of weights. On VGG11 applied to the CIFAR-10 dataset, the method removes 86.6% of the accumulation terms of the deepest convolution for a 0.37 -point accuracy drop, and 25% of the full-network arithmetic when used on the three deepest convolutions simultaneously, for a 1.36 -point drop.
[LG-10] SkillTFM: Gated Skill Evolution for Training-Free Adaptation of Tabular Foundation Models
链接: https://arxiv.org/abs/2608.06137
作者: Yi He,Zhengkang Guan,Anpeng Wu,Peng Cui,Fei Wu,Kun Kuang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Tabular data are ubiquitous in real-world applications and are crucial for data-driven prediction and decision-making across science, industry, finance, healthcare, and public services. Tabular foundation models (TFMs) have emerged as a promising paradigm for general-purpose tabular learning, offering reusable predictors across diverse datasets and substantially reducing the need for task-specific training, tuning, and model development. However, their practical deployment remains constrained by distribution shifts, heterogeneous feature semantics, and task-specific patterns that are difficult to capture without costly fine-tuning or additional labeled data. To this end, we propose SkillTFM, a training-free system that shifts TFM adaptation from parameter updates to the gated evolution of agentic skills. The core of SkillTFM is a verifiable and extensible skill bank that couples boundary evidence identification with gated skill evolution: the former characterizes task structure and base-model failure patterns, whereas the latter retrieves and extends reusable skills subject to explicit validation. Across simulated boundary settings and real-world electricity-price forecasting, SkillTFM improves AUC by 0.128–0.142, raises nonlinear-boundary AUC from 0.699 to 0.898. Furthermore, experiments across TFM backbones demonstrate the effectiveness and generality of SkillTFM. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2608.06137 [cs.LG] (or arXiv:2608.06137v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.06137 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-11] LLM Inference Under Bursty Workload Distribution: Modifying the WAIT Algorithm
链接: https://arxiv.org/abs/2608.06135
作者: Anjali Gangadhar Katageria,Shobha Rani,Raghu Nandan Sengupta
类目: Machine Learning (cs.LG)
*备注:
Abstract:Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving. Recent work has focused on improving scheduling algorithms to boost throughput while maintaining low latency. However, these approaches often assume Poisson request arrivals with constant rates - an assumption that fails to reflect the inherently bursty and dynamic nature of real-world traffic. We propose a lightweight extension to the state-of-the-art WAIT algorithm [1], which adapts to time-varying arrival rates without prior traffic knowledge. The proposed algorithm performs online estimation of request intensity based on observed interarrival times. Using Markov Modulated Poisson Process (MMPP)-based synthetic workloads with diverse request types, we conduct a simulation-based evaluation demonstrating that the proposed method achieves higher throughput than Sarathi-Serve [2], ORCA [3], and vLLM [4] in the evaluated low arrival-rate shift scenarios while maintaining comparable latency.
[LG-12] Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations
链接: https://arxiv.org/abs/2608.06107
作者: Guillaume Couairon,Alexis Jacq,Yu-Han Wu,Renu Singh,Yana Hasson,Quentin Berthet,Romuald Elie
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 34 pages, 32 figures
Abstract:Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models. However, standard auto-regressive ML emulators often suffer from error accumulation over long horizons and struggle to capture the stochasticity of complex physical systems. In this paper, we propose Kastor, a comprehensive methodology to adapt a deterministic physics foundation model into a highly efficient and accurate generative surrogate. First, we introduce a two-stage inference scheme that combines a large-stride causal auto-regressive model with a non-causal temporal super-resolution network, significantly reducing error accumulation while minimizing computational cost. Second, we present Mean prediction regularization (MPR), a novel training objective that constrains the generative model to predict the deterministic distribution mean under null noise conditioning. This regularization dramatically improves the performance and stability of both Functional Generative Networks (FGN) and diffusion-based emulators. Finally, we demonstrate that incorporating spatial gradient matching improves the accuracy and physical fidelity of the simulations as measured by power spectrum density. Extensive evaluations on diverse simulation datasets of the benchmark The Well show that with these components, our model outperforms competing methods in forecasting accuracy, spectral consistency, and computational efficiency. Our model achieves a 42.9% average reduction in forecasting compared to our reference based on the Walrus finetuning methodology, and outperforms Walrus for 8 out of 10 datasets on variance-normalized RMSE (VRMSE).
[LG-13] ML-for-ML
链接: https://arxiv.org/abs/2608.06046
作者: Yutong Zhao,Noga H. Rotman,Gianni Antichi,Ran Ben Basat
类目: Networking and Internet Architecture (cs.NI); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 8 pages, 3 figures
Abstract:AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster. Comments: 8 pages, 3 figures Subjects: Networking and Internet Architecture (cs.NI); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) Cite as: arXiv:2608.06046 [cs.NI] (or arXiv:2608.06046v1 [cs.NI] for this version) https://doi.org/10.48550/arXiv.2608.06046 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-14] Dynamic Graph Prompting via Topology-Routed Mixed-Curvature Experts
链接: https://arxiv.org/abs/2608.06031
作者: Quanxin Wang,Xuanting Xie,Bingheng Li,Xingtong Yu,Shuo Wang,Ruiyi Fang,Zhao Kang
类目: Machine Learning (cs.LG)
*备注: Suggestions and comments are welcomed
Abstract:Dynamic graph prompting freezes a pre-trained temporal backbone and adapts it to label-scarce downstream tasks using lightweight prompts. However, existing methods operate within a single, fixed embedding space. In this work, we reveal that temporal shifts in local clustering and degree heterogeneity actively reorganize the edge curvature spectrum—indicating that the optimal representation geometry dynamically evolves with local topology over time. We formalize this unaddressed mismatch as geometry under-adaptation. To overcome this limitation, we propose CurvPrompt, a topology-routed geometry prompting framework for dynamic graphs. Instead of relying on a single space, CurvPrompt maintains a bank of curvature-diverse Riemannian experts, each paired with a learnable prompt. A topology-aware gate dynamically routes each node–time instance to a sparse subset of experts, constructing a personalized mixed-curvature representation. To ensure parameter efficiency and training stability under extreme label scarcity, CurvPrompt employs soft routing during pre-training to build a continuous topology–geometry mapping, and transitions to hard Top-K routing with uniform weights during downstream adaptation. Extensive experiments across four benchmark datasets show that CurvPrompt significantly advances few-shot link prediction while delivering strong, consistent performance on node classification tasks, validating the necessity of geometry-adaptive prompting.
[LG-15] BioKD: Selective Physiology-to-Video Knowledge Distillation via Reliability Gate for Emotion Recognition
链接: https://arxiv.org/abs/2608.06023
作者: Bojing Hou,Ruohao Li,Yitong Zhu,Hongjun Liu,Luwen Yu,Yuyang Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:To address the limitations of video-based emotion recognition under ambiguous or socially masked behavioral cues, as well as the poor deployability of physiological signals, this paper proposes a reliability-aware physiology-to-video knowledge distillation framework, termed BioKD. The proposed framework leverages physiological signals as privileged information during training to guide a video-based student model in learning deep affective representations, while relying solely on non-intrusive video inputs at inference time. To cope with the high noise and instability of physiological teacher supervision caused by inter-subject variability, signal artifacts, and temporal inconsistency, BioKD incorporates a sample-wise reliability-aware gating mechanism together with a progressive distillation strategy. By adaptively regulating the strength of knowledge transfer, the framework suppresses negative transfer induced by unreliable physiological supervision and enables more stable cross-modal distillation. Experiments on DEAP and AMIGOS show that BioKD consistently outperforms representative baselines under both trial-wise and subject-wise evaluation protocols for valence and arousal recognition. For example, BioKD achieves 68.01% on DEAP (trial-wise arousal) and 65.29% under the more challenging subject-wise setting, demonstrating improved performance under a subject-independent evaluation setting. Further analyses show that BioKD effectively mitigates overconfident teacher errors and outperforms an entropy-only weighting strategy, confirming the importance of explicitly modeling supervision reliability. In addition, BioKD introduces no additional inference-time overhead relative to the same video student architecture and removes the need for physiological sensing and multimodal synchronization.
[LG-16] Do Tabular Foundation Models Agree with Themselves?
链接: https://arxiv.org/abs/2608.06004
作者: Christian Klötergens,Vijaya Krishna Yalavarthi,Lars Schmidt-Thieme,Tom Hanika
类目: Machine Learning (cs.LG)
*备注:
Abstract:Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems. They are constructed as transformers that approximate the Bayesian posterior predictive distribution based on a pre-training prior. These univariate predictors can be converted into multivariate ones autoregressively by sampling one target and adding it to the features. However, the faithfulness of the resulting joint has not been investigated. Furthermore, TFMs cannot be evaluated against the posterior itself, at least not on real-world datasets, because the ground-truth distribution is unknown. We therefore propose asking a different question: could a model’s predictions result from any joint distribution? To answer this question, we pose two requirements that any such model must satisfy. The first is marginalization consistency, which demands that marginalized conditionals are equal to directly predicted marginals. The second is factorization consistency, which demands that different factorization orders result in equal joint distributions. Every TFM that we evaluate violates both of these requirements for both classification and regression across all datasets. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2608.06004 [cs.LG] (or arXiv:2608.06004v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.06004 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-17] A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies
链接: https://arxiv.org/abs/2608.05995
作者: Frieder Wizgall,Georg Tirpitz,Moritz Seiler,Kerstin Ritter,Bálint Mucsányi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. This often requires disentangling epistemic uncertainty from aleatoric uncertainty, yet these uncertainty types are not defined consistently across the literature, making it difficult to assess whether a method produces accurate uncertainty estimates. Evaluation is further complicated by the fact that ground-truth epistemic uncertainty is typically unavailable. Existing benchmarks therefore mostly rely on proxy tasks such as out-of-distribution detection, which do not provide complete ground-truth uncertainty targets and offer limited insight into the structure and quality of uncertainty estimates. We propose a unified definition of uncertainty as pointwise posterior risk, the expected loss of a predictor under the distribution of plausible ground-truth functions given the data. This view combines Bayesian uncertainty over functions with estimator-dependent deviations from the posterior mean, capturing effects such as misspecification and optimization error. This formulation constitutes the foundation of a theory-backed benchmark that enables direct computation of oracle epistemic and aleatoric uncertainty using semi-synthetic datasets with real covariates and known generative processes. By avoiding proxy evaluations, the benchmark enables fine-grained analysis of uncertainty estimates. Empirically, we find that accurate prediction does not guarantee reliable uncertainty disentanglement. The benchmark reveals practically useful differences between methods, identifying approaches with meaningful alignment to oracle uncertainty targets while exposing sensitivity to datasets and modeling choices.
[LG-18] Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
链接: https://arxiv.org/abs/2608.05989
作者: Xinwei Liu,Junyuan Liang,Jianting Zhang,Wuhui Chen
类目: Machine Learning (cs.LG); Robotics (cs.RO)
*备注:
Abstract:Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary prediction performed either in latent space (self-prediction) or observation space (observation prediction). However, state-of-the-art methods from both categories still struggle on challenging visual control tasks when training data is limited. We posit that relying on either predictive objective alone may be insufficient. In contrast, observation prediction grounds learned representations in observation-level dynamics, but does not directly regularize the temporal predictability of latent representations over extended horizons. In this paper, we propose Observation-Grounded Self-Predictive Representations (OG-SPR), a model-free visual RL algorithm for continuous control that learns representations that are both temporally predictive in latent space and grounded in observation-level dynamics. OG-SPR incorporates two core auxiliary objectives: multi-step latent self-prediction and next-observation prediction. We empirically show that directly imposing latent self-prediction on the shared representation may over-constrain it and does not necessarily improve performance. To address this issue, OG-SPR introduces two lightweight adapters for latent self-prediction, allowing the shared representation to benefit from temporally predictive signals without being forced to directly satisfy the self-prediction objective. Experiments on 28 visual control tasks from the DeepMind Control Suite show that OG-SPR improves aggregate performance over state-of-the-art self-predictive and observation-predictive RL methods, with particularly pronounced gains in challenging domains such as dog and humanoid.
[LG-19] HBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction
链接: https://arxiv.org/abs/2608.05982
作者: Pui Chung Siu,Claudia Cabrera,Mani Mudaliar,Arkaitz Zubiaga
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注: 19 pages, 11 figures, 8 tables
Abstract:Inadequate target–disease linkage accounts for 40–50% of Phase~II efficacy failures, so anticipating which programmes will advance would let sponsors back the hypotheses most likely to reach patients. What a programme can be judged on is the evidence that supported its linkage \emphwhen it entered the clinic. No existing biomedical knowledge graph allows that evidence profile to be assembled as of a past date. We present the Temporal Heterogeneous Biomedical Knowledge Graph (THBKG), which describes and predicts therapeutic target–disease links through time: 110,396 entities and 11.1M edges across nineteen relation types, each edge carrying the year its evidence changed, so a pair’s profile can be recovered as it stood when its own decision fell due. On this graph we define a decision-aligned benchmark that predicts, for a target–disease pair entering Phase~II, whether it advances to Phase~III on evidence datable before that decision. Graph propagation over the THBKG outranks every direct-evidence reference scored under the same decision-aligned protocol, reaching a relative success of 4.3–4.5 at the top ten pairs per therapeutic area. The gain concentrates on the 72.8% of pairs with no direct target–disease evidence at their decision point, where a direct-edge model has nothing to read: the encoders still rank five- to sixfold above chance, recovering the signal by propagating over the intervening biology. Adapting a path-based explainer to the decision-time subgraph decomposes each prediction into the evidence landscape behind the hypothesis for explainable prediction. We release the THBKG as a continually updated substrate for studying therapeutic target hypotheses by retrospective validation.
[LG-20] How Far Do Simple Transformations Translate Across Text Embedding Models?
链接: https://arxiv.org/abs/2608.05980
作者: Sid Ali Hamideche,Louis Adrien Dufrene,Quentin Lampin,Guillaume Larue(Orange Research)
类目: Machine Learning (cs.LG)
*备注:
Abstract:We investigate whether simple transformations can translate representations across heterogeneous text embedding models. Understanding how independently trained models organize semantic information is an enabler for AI-to-AI latent communication without decoding into human-readable text. Focusing on lightweight translators such as linear mappings, we test the literature hypothesis of latent universality in a realistic text setting beyond simplified benchmarks. Across nine embedding models differing in architecture, pooling strategy, and training objective, we evaluate compatibility using CKA, downstream transfer, fidelity, and retrieval. Simple translators recover meaningful shared structure and support transfer for some compatible pairs, but fail sharply for others. Compatibility depends jointly on architecture, training objective, pooling, and data distribution. Overall, the results show that heterogeneous embedding spaces are not universally related by simple mappings as often suggested in some literature.
[LG-21] Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage Negative Results and Operational Hardening
链接: https://arxiv.org/abs/2608.05944
作者: Seon Ho Kim,Ui Jeong Jeon,Su Hyeon Kim,Min Tae Hwang
类目: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 13 pages, 5 figures. Experience report
Abstract:We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) – among the first published field accounts on this accelerator. We claim no new algorithm. The individual mechanisms we use are established practice; our contribution is the integrated field experience and a set of calibrated measurements on new hardware. Concretely we offer four practitioner artifacts. (1) A B300-calibrated power-draw triage table that distinguishes compute / communication / data-starvation / checkpoint-or-deadlock / idle by board wattage (utilization% reads 100% during an NCCL hang). (2) A set of honest negative results that dispel common optimization folklore at this scale: a controlled A/B in which per-step NFS reading matches a pretokenized local cache (~53k tok/s) because the corpus fits in page cache and the job is compute-bound; and a reconstruction of an earlier “throughput collapse” as NFS/CPU contention rather than a storage-medium limit. (3) Calibrated 4/8/16-GPU strong-scaling and GPU-hour numbers on B300 (near-linear, as expected in this regime; we report absolute values as reference data). (4) A worked failure case – an epoch-end NCCL deadlock from per-rank token-packing imbalance – together with a 2.7-second pre-run invariant gate and an external watcher that turn multi-hour silent failures into instant rejections. This deadlock and its remedy correspond to PyTorch’s documented Join / equalize-to-minimum practice; we position our instantiation against that prior art and report the GPU-hours the failure cost and the gate saves. The transferable takeaway is operational, not algorithmic: for data-dependent data-parallel jobs, watch power rather than utilization, and verify invariants before launch – a passing smoke test is not evidence of a safe full run.
[LG-22] BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
链接: https://arxiv.org/abs/2608.05928
作者: Yuhao Wang,Zelin Zang,Yuxuan Liu,Zhen Lei,Stan Z. Li
类目: Machine Learning (cs.LG)
*备注: 34 pages, 6 figures, and 13 supplementary tables (Tables S1-S13); includes Supplementary Information with detailed training and evaluation protocols. Numerical source data for all figures are provided as ancillary files; training code and the BioM-JEPA checkpoint will be released via GitHub
Abstract:Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
[LG-23] CohortHijack: Robustness of Single Cell Annotation to Companion Cell Removal
链接: https://arxiv.org/abs/2608.05900
作者: Arash Vashagh,Yasmin Vashagh
类目: Machine Learning (cs.LG)
*备注:
Abstract:Many single-cell annotation tools refine an initial cell label using nearby cells or cluster-level voting. We study whether this refinement can be manipulated without changing the target cell. We introduce CohortHijack, a robustness audit that removes selected non-target cells from the query cohort while preserving the target expression profile, base prediction, and trained model. We evaluate random and structured removal methods, together with greedy, multi-start, and beam search, on PBMC3K and Paul15 using logistic regression and calibrated linear SVM classifiers. Structured removal was consistently stronger than random removal on Paul15. Multi-start search changed 24.33% of linear-SVM targets and 19.67% of logistic-regression targets while removing a small fraction of the cohort and keeping mean collateral changes below 0.4%. Ablations confirmed that the effect disappeared when neighborhood refinement was disabled. We also evaluated CellTypist majority voting, where independent predictions remained unchanged across all evaluations, but refined labels changed after small companion-cell removals. These findings identify query cohort composition as a target-preserving attack surface in single-cell annotation.
[LG-24] Alternating Levenberg-Marquardt Training of Physics-Informed Neural Networks with Fourier-Enhanced Features
链接: https://arxiv.org/abs/2608.05892
作者: Yulun Wu,Matthieu Barreau,Miguel Aguiar,Karl H. Johansson
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 53 pages, 18 figures, 6 tables
Abstract:Physics-informed neural networks (PINNs) often fail to accurately resolve partial differential equations (PDEs) with high-frequency or multi-scale solutions, as well as strongly nonlinear problems. Two factors underlie this difficulty: spectral bias, the tendency of neural networks to underfit high-frequency features; and representation-coefficient coupling, the entanglement of representation learning and coefficient fitting within a single nonconvex optimization objective. In this work, we propose the Fourier-enhanced alternating Levenberg–Marquardt PINN (FALM-PINN), an optimization framework that decouples representation learning from coefficient fitting. The upper-level problem learns a Fourier-enhanced basis that enriches the latent space with high-frequency components, while the lower-level problem resolves the coupling by fitting the projection coefficients on this basis, solving a nonlinear least-squares problem with the Levenberg–Marquardt algorithm. The framework applies to general nonlinear and coupled PDE systems, and reduces to a single-step convex optimization problem for linear PDEs. We prove global convergence of the alternating training scheme in both cases. Numerical examples on multiple challenging high-frequency and nonlinear PDEs show that FALM-PINN achieves relative L^2 errors up to two orders of magnitude lower than state-of-the-art baselines.
[LG-25] A neural operator view on U-Nets for inverse imaging problems
链接: https://arxiv.org/abs/2608.05839
作者: Alexander Auras,Martin Burger,Samira Kabri,Michael Moeller,Michael Schopf-Kuester
类目: Numerical Analysis (math.NA); Machine Learning (cs.LG)
*备注:
Abstract:Deep neural networks have shown great empirical success in the solution of a wide variety of ill-posed inverse problems in imaging. Yet, very few works have studied their behavior in the limit that turns the discretized ill-conditioned problems into truly ill-posed ones, i.e., for an increasing resolution of the discretization. In this work, we review common approaches to neural operator learning in architectures that resemble a U-Net, one of the most common classical architectures for inverse imaging problems. We discuss advantages and drawbacks of the respective approaches, consider a 1D toy example for improved interpretability, and present extensive numerical experiments on how different types of neural operator U-Nets can improve a first (crude) limited angle CT-reconstruction. In particular, we study how well networks trained for a certain resolution of the discretization generalize to other resolutions. Our finding is that while U-shaped neural operator architectures are by design resolution-invariant, the classical U-Net architecture seems to be more robust with respect to resolution changes than expected.
[LG-26] Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation
链接: https://arxiv.org/abs/2608.05819
作者: Alfred M. Pastor,Maribel Castillo,Jose M. Badia
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF); Quantum Physics (quant-ph)
*备注:
Abstract:Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting circuit structure, although its efficiency depends strongly on the chosen contraction plan. On GPUs, plans with similar theoretical complexity may perform very differently because execution also depends on parallelism, reduction structure, memory traffic, and contraction geometry. We present a learning-to-rank framework for selecting efficient contraction plans before executing them. Each plan is represented by structural features derived directly from its sequence of pairwise contractions, and gradient-boosted rankers are trained from GPU measurements using listwise and pairwise objectives. We evaluate the resulting models on diverse circuit families, using separate in-distribution and circuit-family-shift test sets, and compare them with random and MinFill-based baselines. The learned rankers generally identify better plans, with the listwise model providing the strongest overall decision quality. We also study backend shift by comparing empirical plan orderings on two GPU architectures and evaluating the source-trained models on the second device without retraining. The rankings remain substantially, though not perfectly, stable across GPUs, and the models retain useful decision quality. These results support Learning to Rank as a practical way to reduce contraction-plan search, while showing that performance remains partly backend dependent.
[LG-27] Neuro-Symbolic Closed-Loop Control of Laser Powder Bed Fusion with an In-Loop Ontology
链接: https://arxiv.org/abs/2608.05773
作者: Gisuk Hong,Jaebong Cho,Hyunbo Cho
类目: Machine Learning (cs.LG)
*备注: 23 pages, 8 figures, submitted to journal(Journal of Intelligent Manufacturing) and under review
Abstract:A geometry-conditioned, neuro-symbolic closed-loop architecture is proposed for laser powder bed fusion, in which a standards-aligned ontology operates inside the control loop and couples symbolic reasoning with statistical learning to set the targets of a constraint-aware predictive controller. The ontology links the process objectives and constraints to the signals a controller can observe, and a description-logic reasoner converts them into the references and bounds enforced on each scan. The demonstrated case is overhang dross, a quality limit on the melt pool depth, which governs quality yet cannot be measured during the build, is mapped through a geometry- and power-dependent depth-to-width ratio onto a bound on the observable width, with the ratio and its calibrated uncertainty supplied by a Gaussian process. The reasoner classifies each upcoming feature and selects the active constraints-adding a lack-of-fusion floor at overhangs, a monotone guard beyond the calibrated range, and an energy-density cap where a process window is declared while running only on changes of geometric context and otherwise leaving a single small quadratic program on the per-scan path. In an Eagar-Tsai surrogate calibrated to the NIST AM-Bench benchmark for IN625, the architecture eliminates the dross produced by a geometry-blind controller, holds dross at zero with only a small residual lack-of-fusion under dual scoring, degrades gracefully under deliberate plant mismatch, and retargets to new alloys and constraints by editing ontology data rather than code. The results establish architectural feasibility, experimental calibration of the ratio is the principal next step.
[LG-28] Accelerating nanodrug development in continuous flow systems using informed prediction models based on low-cost surrogate nanoparticles
链接: https://arxiv.org/abs/2608.05761
作者: Kai Dahms,Eilien Heinrich,Jochen Schmid,Michael Bortz,Iryna Savych,Regina Bleul
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注:
Abstract:The development of nanotherapeutics often involves extensive empirical optimization due to the sensitivity of nanoparticle properties, such as size and polydispersity index (PDI), to minor changes in process parameters. Factors like formulation concentration, flow rates, and mixing ratios can significantly influence clinical efficacy and therapeutic outcomes. The absence of predictive mathematical frameworks has made iterative experimental screening necessary, increasing both costs and development time. This study introduces and validates a predictive modeling approach based on shape constraints, aiming to enhance the estimation of nanoparticle characteristics across various process conditions. Using controlled microfluidic methods, liposomes and lipid nanoparticles were systematically prepared under varying lipid concentrations, flow rates, and aqueous-to-organic mixing ratios. The shape-constrained model, informed by both experimental data and expert knowledge, was subsequently validated for a pharmaceutical application using minimal empirical data. Results reveal that shape-constrained modeling facilitates accurate prediction of nanoparticle size and dispersity, reducing the need for extensive experimental workflows. This framework supports rational and efficient process development for manufacturing nanomedicine systems. Subjects: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE) Cite as: arXiv:2608.05761 [cs.LG] (or arXiv:2608.05761v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.05761 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-29] Equipment-centric workpiece localization in near real-time using deep learning-based vision and event-driven finite state machines
链接: https://arxiv.org/abs/2608.05744
作者: Dohyeon Kong,Jaebong Cho,Hyunbo Cho
类目: Machine Learning (cs.LG)
*备注: 21 pages, 14 figures, 9 tables. Published in The International Journal of Advanced Manufacturing Technology
Abstract:Continuous workpiece localization is essential for traceability and process coordination in hot forging, but direct tracking is unreliable because of extreme temperatures, surface degradation, and irregular routing. This study presents an equipment-centric framework that infers workpiece locations from handling equipment observed by multiple static 2D cameras. The framework estimates floorplan-space 3D equipment coordinates and recognizes grasp and release activities. Event-driven finite state machines validate these activities as discrete handling events and continuously update workpiece states and locations. A keypoint-guided attention mechanism integrated into a 3D convolutional neural network improves activity recognition by focusing on functionally relevant equipment regions. Evaluation in an operational hot forging factory achieved 100% event detection accuracy within a 33-second tolerance window, a mean localization error of 317.8 mm, and a mean system latency of 21 seconds. The framework connects vision-based perception with interpretable event-driven reasoning and supports visualization of workpiece transfers and quantitative analysis of equipment operations.
[LG-30] CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
链接: https://arxiv.org/abs/2608.05732
作者: Mehrshad Saadatinia,Parsa Razmara,Ardalan Aryashad,Ali Abbasi,Seyedarmin Azizi
类目: Machine Learning (cs.LG)
*备注:
Abstract:Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment. Existing steering methods, such as Contrastive Activation Addition (CAA), typically rely on fixed single-layer interventions derived from aggregate activation differences. These methods impose a single intervention across semantically diverse inputs and often fail to sustain consistent behavioral changes across layers, limiting the effectiveness of the steering. In this work, we introduce CircuitSteer, a novel framework that leverages Sparse Autoencoders (SAEs) to identify and manipulate coherent semantic circuits distributed across multiple layers. By constructing a feature flow circuit based on feature co-activation and the geometric alignment of decoder directions, we isolate the specific multi-layer subcircuits responsible for a target behavior. We then synthesize dense steering vectors from these sparse features and apply multi-point interventions to guide the model’s internal semantic trajectory. We evaluate CircuitSteer using contrastive examples across a diverse set of tasks, including toxicity, emotion-intensity, sycophancy, and refusal, spanning two model families. Across all models and datasets, CircuitSteer is the only method to consistently produce fluency-preserving interventions; competing methods either sacrifice text quality or lack coverage, failing entirely on complex behaviors like sycophancy and refusal. These results demonstrate that multi-layer circuit steering, enabled by enforcing geometric alignment among selected features, yields strictly more robust and effective behavioral control than static single-point interventions. Code is available at this https URL.
[LG-31] LILAC: An Idempotent Neural Speech Codec
链接: https://arxiv.org/abs/2608.05727
作者: June Young Yi,Dongwook Lee,Jiheum Yeom,Sungroh Yoon
类目: ound (cs.SD); Machine Learning (cs.LG)
*备注: 22 pages, 4 figures
Abstract:Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper’s twelve baseline systems, every configuration tested rewrites, on average, at least 15% of its tokens in a single decode-re-encode pass. This poses a problem for utilizing Neural Audio Codecs as token interfaces in pipelines where re-encoding decoded outputs can occur. We present LILAC, a fully convolutional 24 kHz speech codec at 9.375 Hz and 0.75 kbit/s that is codec idempotent by construction; re-encoding the decoded audio of any valid token stream returns the identical stream. LILAC achieves idempotency while maintaining competitive quality, reaching UTMOS 4.14 and 4.24 on LibriSpeech and LibriTTS-R test sets, comparable to SOTA sub-1 kbit/s Neural Audio Codecs.
[LG-32] SEAM: Global consistency beyond local accuracy in scientific machine learning
链接: https://arxiv.org/abs/2608.05702
作者: Gnankan Landry Regis N’guessan,Bum Jun Kim
类目: Machine Learning (cs.LG); Computational Engineering, Finance, and Science (cs.CE)
*备注: 43 pages, 9 figures
Abstract:Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. Yet such local checks cannot establish whether the resulting explanations can be assembled into one globally admissible explanation. We introduce Scientific Explanation-Admissibility Machines (SEAM), a generator-agnostic framework that makes this local-to-global consistency question computable across regions, sensors, regimes, and model components. The finite explanation-sheaf instantiation SEAM- \Omega represents each region by a structured explanation with state, closure, and observation channels together with optional contract metadata; compares neighboring explanations on their overlaps; and converts disagreement into a channel-resolved obstruction. This obstruction locates inconsistency and tests competing declared accounts by restricting each repair to the revisions that one account permits. Exact feasibility refutes or retains an account; when exact repair is unavailable, residual-aware regularized records provide a separately labeled empirical attribution. The framework also separates inconsistency from non-identifiability and monitors learned generators under distribution shift. We establish theorems for minimum-cost intervention and conservation-contract detectability, together with companion results for identifiability and closure recoverability. Across nineteen experiments involving synthetic partial differential equation systems and out-of-distribution Fourier neural operator (FNO) monitoring, SEAM detects incompatible explanations even when local predictions are accurate, and attributes failures to specific channels and overlaps. SEAM adds a global explanation-consistency audit to existing solvers and learning models, testing whether their local explanations form a coherent scientific account.
[LG-33] Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading
链接: https://arxiv.org/abs/2608.05675
作者: Rasul Khanbayov,Hasan Kurban
类目: Machine Learning (cs.LG)
*备注:
Abstract:Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader’s answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at. We act on the complementary relation, equivariance: edit a figure’s data and the correct answer must change by a computable amount. Two matched edits are provably complete for affine reading errors; no suite of swap edits is complete for label permutations, and cyclic relabeling closes most of that gap. We instantiate the theory as the Equivariance-Consistency Score, a label-free, training-free detector, and release REND-EQUIV, pairing matched invariance and equivariance sets over identical data. The predicted ordering holds across three models and a hand-labeled population immune to the one circularity in how it is selected; a second invariance-family method confirms the blind spot belongs to the relation, not to any implementation; and cyclic relabeling delivers its predicted gain on a matched real sample. The same characterization explains a reported inversion of this ordering in the classifier metamorphic-testing literature: detectability is a joint property of the relation and the fault class, never of the relation alone.
[LG-34] When Does Consensus Mean Correctness? Measuring the Agreement-Accuracy Coupling with Semantics-Preserving Re-Rendering
链接: https://arxiv.org/abs/2608.05670
作者: Rasul Khanbayov,Hasan Kurban
类目: Machine Learning (cs.LG)
*备注:
Abstract:A model’s agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness. That coupling is rarely measured directly: natural-image perturbations preserve meaning only by assumption, and no exact answer key localizes errors. Scientific figures remove both obstacles, a figure is drawn from data by a program, so redrawing it yields images that are semantically equivalent by construction and share a programmatically exact answer. We build RENDEQ, a generator of such render-equivalence sets, and measure the coupling on three open-weight VLMs, checking every finding across three independent instantiations. Re-rendering beats resampling on both accuracy and reliability. Agreement beats an evidence-carrying baseline, mean token log-probability, on two of three models and ties on the third, reversing an intermediate, buggy replication traced to a rendering-pipeline failure. The dispersion behind this is concentrated in one style factor, the plotting library, more than double the next-largest factor and an order of magnitude above the noise floor. Fine-tuning on the model’s own cross-render consensus inverts: accuracy falls in every one of five replication runs, the opposite sign to published results on natural images. Agreement certifies correctness only above a threshold set by how diffuse a model’s errors are, and an objective that rewards agreement destroys exactly that diffuseness.
[LG-35] Potential Matching Optimal Transport: Continuous Normalizing Flows for Exact p-Wasserstein Dynamics
链接: https://arxiv.org/abs/2608.05666
作者: Lishuo Zhang(1),Ruizhi Huang(1),Yang Yu(1),Lei Li(1 and 2) ((1) School of Mathematical Sciences, Shanghai Jiao Tong University, (2) Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University)
类目: Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:
Abstract:We introduce Potential Matching Optimal Transport (PMOT), a potential-flow framework for general p -cost optimal transport with c_p(x,y)=|x-y|^p . PMOT parameterizes the CNF velocity field with a scalar potential in the generalized Benamou–Brenier form for the chosen exponent p . It trains the potential gradient with a self-induced matching loss along straight bridges determined by the model’s own endpoints, while allowing flexible terminal distribution matching. Our main result establishes zero-loss exactness: under the stated regularity, exact terminal matching, and uniqueness assumptions, any zero-loss solution satisfies the generalized Benamou–Brenier optimality system and recovers the corresponding p -optimal transport map and dynamics. On synthetic benchmarks, PMOT learns p -specific maps that agree with the corresponding p -matched OT references. It also remains competitive as a likelihood-based density model on high-dimensional tabular data, and an MMD-based color transformation experiment demonstrates flexible sample-based terminal matching.
[LG-36] RASP-QAOA: Resource-Aware Per-Instance Selection for Exact QAOA Simulation
链接: https://arxiv.org/abs/2608.05646
作者: Chih-Chung Hsu
类目: Emerging Technologies (cs.ET); Machine Learning (cs.LG)
*备注: 9 pages, 5 figures. Code and reproducibility package: this https URL
Abstract:Exact QAOA simulation spans several computational representations whose useful regions differ sharply across graph structure, circuit depth, precision, and available memory. Choosing only a backend name hides these differences: an executable choice also fixes the representation, adapter, precision mode, and memory policy. We introduce RASP-QAOA, a per-instance selector over ten such actions. It first removes actions that cannot implement the requested QAOA semantics or execution requirements, then orders the remaining actions using instance features; actions outside learned support are handled by analytical work estimates. On a content-disjoint 60-request H200 evaluation, RASP-QAOA succeeds on all 31 requests for which at least one admissible action completes and validates. Within this set it reaches 27/31 top-1 and 31/31 top-2 selection, with 1.051 geometric-mean regret. Its failure-penalized PAR10 score is 0.0396 times that of development-selected CUAOA (95% interval: 0.0085-0.1644). A separate 30-request crossover shows that graph structure changes 16 decisions and improves the paired penalized score, while a depth-1 stump matches gradient boosting. The evidence supports resource-aware representation selection at n = 35, p = 5, with gains driven by representation features rather than classifier complexity.
[LG-37] LC-Implicit-QAOA: Active-Workspace-Capped Exact Objective-and-Gradient Evaluation for Training over Bounded QUBO Light Cones
链接: https://arxiv.org/abs/2608.05610
作者: Chih-Chung Hsu
类目: Emerging Technologies (cs.ET); Machine Learning (cs.LG)
*备注: 8 pages, 2 figures, 8 tables. Code and technical supplement: this https URL
Abstract:QAOA training repeatedly queries an objective and all shared gradients, making exact evaluation a feasibility bottleneck even when QUBO terms have bounded causal cones. Building on established causal-cone restriction and adjoint differentiation, LC-Implicit-QAOA profiles cone structure and induced-edge counts before local-amplitude and named-workspace allocation, then jointly selects equal-size microbatches and checkpoint schedules under a named active-evaluator workspace budget. “Implicit” means omitting both global state and global cost table, not implicit differentiation; infeasible requests are rejected before those allocations. An independently implemented complex128/float64 dense adjoint agrees with LC over 1,800 graph-angle comparisons, with a worst relative gradient error of 1.56 x 10^-13. LC completes all 104 target requests in a p=2 bounded-cone grid; under a prespecified n = 24 validation cap, the matched state-plus-cost reference is executed for 28 requests and deliberately not run on 76. Across 80 budgeted requests, measured allocated evaluator memory stays within budget, reaching at most 0.797 of it. On 3-regular n=512, p=2, the adjoint reaches the same finite-budget endpoint in 101 objective-equivalent calls and 189 s, versus 909 calls and 1,565 s for central differences. LC targets fixed-depth one- and two-local diagonal QUBO costs with a transverse-field mixer; it provides neither global states, sampling, nor a hardware-independent fastest-backend rule.
[LG-38] Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security Baselining
链接: https://arxiv.org/abs/2608.05605
作者: Mohammad Arafath Uddin Shariff,Byrav Ramamurthy
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG); Networking and Internet Architecture (cs.NI)
*备注:
Abstract:Research and Education Networks (RENs) serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their normal traffic patterns which are characterized by massive, bursty “elephant flows” are statistically indistinguishable from volumetric attacks such as DDoS to conventional monitoring systems. This similarity leads to high false-positive rates in anomaly detection, blinding security operators to genuine threats. In this paper, we propose and evaluate a high-fidelity traffic forecasting framework designed to establish dynamic security baselines for RENs. Leveraging an exclusive 57-day Internet2 dataset spanning ten backbone routers (13.7 billion packets), we perform the first large-scale benchmark of anomaly-aware forecasting models in this domain. We systematically evaluate six model families, from SARIMA to state-of-the-art long-sequence architectures (TiDE, PatchTST), across 960 experimental configurations. Our results demonstrate that these advanced architectures, particularly TiDE, reduce baseline prediction error by 30-42% compared to traditional methods ( p 0.001 ), significantly improving the distinction between legitimate scientific bursts and potential anomalies. Furthermore, we introduce a novel anomaly-integration strategy that improves model robustness by 3.3% in the presence of noise. This work provides the first statistically validated framework for distinguishing scientific workflows from network attacks, enabling more autonomous and resilient network security operations.
[LG-39] Behavioral Residualization for Unsupervised Intrusion Detection in Automotive CAN Networks
链接: https://arxiv.org/abs/2608.05548
作者: Chandan Hegde,Mukundh R Reddy
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注: 9 pages, 5 figures, 8 tables
Abstract:Modern vehicles rely on the Controller Area Network (CAN) bus, whose design prioritizes low cost and real-time performance but provides no message authentication or encryption. An attacker with physical or remote access can therefore inject arbitrary frames, making intrusion detection an important defense-in-depth mechanism. Most published CAN intrusion detection systems rely on presence-based features, such as novel arbitration IDs, frozen payload bytes, or anomalous DLC values. These features perform well on public datasets containing easily separable attacks but fail when attackers reuse legitimate arbitration IDs. We present per-ID behavioral residualization, a CAN-specific representation that extracts fourteen temporal, protocol, and payload features from sliding windows and residualizes them against each arbitration ID’s normal baseline. Our central claim is that this representation, rather than any individual detector, drives the performance gains. Across six unsupervised detectors and two datasets, residualization improves mean F1 in the majority of evaluations (21/24 on HCRL and 30/36 on ROAD across five seeds). On the more realistic ROAD dataset, where attacks reuse legitimate IDs, the representation achieves recall = 0.99 with high ROC-AUC on targeted signal-manipulation attacks. Two limitations are explicitly quantified: novel-ID flooding (HCRL DoS, F1 = 0.02) and cross-ID fuzzing (ROAD, F1 = 0.27), defining the measured coverage boundary of the proposed representation. Comments: 9 pages, 5 figures, 8 tables Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG) Cite as: arXiv:2608.05548 [cs.CR] (or arXiv:2608.05548v1 [cs.CR] for this version) https://doi.org/10.48550/arXiv.2608.05548 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-40] KV-Skill: Forging Expertise in the Models Native Language
链接: https://arxiv.org/abs/2608.05475
作者: Zhaowei Han,Xiang Zhang,Bing Han,Kai Liu,Danqi Hu,Jie Liu
类目: Machine Learning (cs.LG)
*备注: 17 pages, 4 figures, 18 tables. Zhaowei Han and Xiang Zhang contributed equally to this work. Code: this https URL
Abstract:Task knowledge is commonly stored either as text in the prompt or as an update to model weights. Text is modular but must be interpreted on every use, while weight adaptation makes the resulting capability difficult to load, remove, or share independently. We introduce KV-Skill, a design space of external factorized operators that a frozen language model reads through a lightweight interface. KV-Skill supports two complementary paths. Registration converts an authored text skill into a text-derived operator and trains a shared per-backbone interface. Reward learning develops a compact latent operator directly from task outcomes, with or without an authored skill. Neither path adds positions to the prompt. Across ten benchmarks and four backbones from three model families, converting text to a KV-Skill consistently makes the same procedural knowledge more effective. On Qwen3.5-4B LiveMath, registration reaches 77.2 accuracy, compared with 23.4 for the source text skill, 52.0 for SkillOpt, and 64.5 for SoftSkill. Under matched reward training and parameter budgets, KV-Skill gives the best result in seven of eight matched settings against soft prefixes, prefix tuning, and LoRA. A post-hoc rank analysis further shows that text-derived operators retain nearly all of their benefit with one task-aligned direction per injection layer, while matched random directions fail. Finally, one shared interface retains three independently loadable KV-Skills without measurable forgetting. These results show that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone. Code is available at: this https URL
[LG-41] Hybrid Probabilistic Zonotopes for Identifiable and Refinable Predictive Uncertainty
链接: https://arxiv.org/abs/2608.05454
作者: Zhen Zhang,Amr Alanwar
类目: Machine Learning (cs.LG); Machine Learning (stat.ML)
*备注:
Abstract:Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region. Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete choice among modes, bounded systematic drift within the chosen mode, and irreducible stochastic noise. We introduce the Hybrid Probabilistic Zonotope (HProbZ), an output head that represents these three sources as binary, bounded, and stochastic generators of a zonotope, and admits a closed-form likelihood by convolution. Sharing the bounded generator across prediction steps couples future predictions algebraically, so observing one step refines the predictive distribution at every remaining step in a single forward pass. We establish that the three generators are identifiable from the likelihood up to permutation, and that an HProbZ density is representationally distinct from any finite Gaussian mixture. The same shared structure provides analytic per-mode risk and distribution-free multi-modal conformal sets at inference time. Empirical analysis on representative prediction benchmarks supports the effectiveness of the design relative to same-encoder mixture baselines, while offering structural properties that mixture or convex-conformal predictors do not jointly provide.
[LG-42] Discrete energy as an exact label-free training objective for finite-element surrogates
链接: https://arxiv.org/abs/2608.05437
作者: Ruifeng Cao(The University of Manchester),Xidan Song(Wuhan University)
类目: Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注: 9 pages. Theory note of the FE-JEPA programme
Abstract:Supervised training of finite-element (FE) surrogate models requires reference solutions, and each reference solution is obtained by solving the system that the surrogate is intended to replace. The assembled discrete potential energy provides a training signal that requires no reference solution. This note records, with proofs, the identities that make this signal exact for linear elastostatics: the difference between the energy of a prediction and the energy of the reference solution equals one half of the squared stiffness-norm error, and the gradient of the energy equals the stiffness-weighted error. Label-free discrete-energy minimisation and supervised regression in the stiffness norm therefore have the same unique minimiser and identical gradients at every point. Around this central result, the note states a conditioning lemma that bounds the displacement error by the energy gap, a modewise contraction identity that explains why the Euclidean displacement error is an unsuitable primary metric, the Chebyshev bound that governs conjugate-gradient post-processing of surrogate predictions, and a conditional latent-separation proposition for joint-embedding predictive architecture (JEPA) pretraining on a shared stiffness operator, with an explicit numerical counterexample that delimits its scope. Every claim with numeric content is implemented as an executable falsification check; the checks were executed twice, on synthetic test problems and on a probe set of 16 instances from the validation split of a pre-registered experimental run, and every inequality holds, with the measured tightness reported. A closing section explains why the construction does not extend to elastodynamics through direct minimisation of the action functional, and which time-discrete formulation restores exactness.
[LG-43] Robust Context-Aware Detection of Malicious Instructions in Text
链接: https://arxiv.org/abs/2608.05430
作者: Buzhao Liu,Xinhang Ma,Yevgeniy Vorobeychik
类目: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
*备注:
Abstract:The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks. Therein, however, also lies their vulnerability to attacks which embed malicious instructions in text, common variants of which are known as indirect prompt injection (IPI). A fundamental task in addressing this vulnerability is successful segmentation of a given text into benign and malicious sentences (if any). While a number of approaches for this task have been proposed, no detector combines query-relative detection at the segment level, and none are hardened against adaptive evasion attacks realizable in agentic executions. We address the former limitation by developing an approach for malicious sentence classification that is both context- and query-aware. Next, to harden the resulting classifier against evasion, we present two adversarial training methods. The first is directly adapted feature-space adversarial training (AT) in which evasions are approximated using projected-gradient-based optimization in the embedding space. The second simulates realizable evasion attacks in the AT loop through LLM-based paraphrasing. Crucially, we parametrize both AT variants to facilitate a smooth tradeoff between utility and attack robustness. In extensive experiments using indirect prompt injection benchmarks we show that the proposed approach outperforms state-of-the-art IPI defense baselines under static attacks, while in the case of adaptive attacks, our AT variants provide significantly higher utility, lower attack success rate, and often both. Finally, we show that the best AT parameters can depend intimately on the particular application domain. Consequently, domain-dependent tuning of malicious text detectors is likely necessary in practice. Our code is publicly available at this https URL.
[LG-44] Can Open-Weight LLM s Produce Kernel-Verified Coq Proofs? A Pilot Study
链接: https://arxiv.org/abs/2608.05420
作者: Ahmed Ryan,Md Erfan,Akond Ashfaque Ur Rahman,Md Rayhanur Rahman
类目: Logic in Computer Science (cs.LO); Machine Learning (cs.LG)
*备注:
Abstract:Large language models (LLMs) can generate text that resembles a mathematical proof, but resemblance does not establish correctness. A formal proof checker verifies whether each proof step follows established logical rules. Coq bases its rules on the Calculus of Inductive Constructions, a logical framework that defines which proof steps the system may accept. This pilot study evaluated six open-weight LLMs on the same 100 theorems from CoqStoq, a benchmark derived from real Coq projects. Each LLM received one attempt per theorem with the temperature set to 0, and Coq checked every proposed proof in the theorem’s original project environment. We counted a proof as successful only if the Coq kernel accepted it. Gemma 4 verified 12 of 100 theorems, Llama 3.3 verified 8, and DeepSeek Coder V2 Lite verified 1. Qwen 3.5, Mistral Small 3.1, and GPT-OSS verified none. The 21 successful model-theorem results covered 15 distinct theorems, 11 of which were not solved by a baseline of standard Coq tactics. All verified theorems had short or medium human-written reference proofs; no model verified a theorem with a long reference proof. Because the proof-length analysis was exploratory, this pattern does not establish that proof length caused the difference. For the three models with at least one success, the total generation cost per verified proof ranged from 741 to 36,193 output tokens, 14.9 to 178.0 seconds, and 0.0167 to 0.2000 aggregate GPU hours. We could not calculate these ratios for models with no verified proofs. Across 600 attempts, the models produced 21 kernel-verified proofs, giving an overall success rate of 3.5%. The study reports descriptive differences among the models but does not statistically test whether one model outperforms another. Therefore, the results do not establish a universal ranking of the six models. Subjects: Logic in Computer Science (cs.LO); Machine Learning (cs.LG) Cite as: arXiv:2608.05420 [cs.LO] (or arXiv:2608.05420v1 [cs.LO] for this version) https://doi.org/10.48550/arXiv.2608.05420 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Ahmed Ryan [view email] [v1] Wed, 5 Aug 2026 21:29:12 UTC (1,270 KB)
[LG-45] Spectral Distillation: From Nonlinear Dynamics to Linear State-Space Models
链接: https://arxiv.org/abs/2608.05416
作者: Liane Galanti,Devan Shah,Shlomo Fortgang,Elad Hazan
类目: Machine Learning (cs.LG)
*备注:
Abstract:Can nonlinear dynamical systems be learned through a compact linear state-space representation, without directly solving a non-convex system-identification problem? We give a provable pipeline for doing so. Starting from observations of an unknown nonlinear dynamical system, we first learn an implicit spectral predictor using Observation Spectral Filtering (OSF), a convex method that competes with the best linear observer for the system. We then apply spectral-to-LDS distillation to convert this predictor into an explicit recurrent linear dynamical system. Our main theorem shows that the average prediction error of the distilled LDS decomposes into an exponentially-small distillation term and the OSF learning term governed by the Luenberger complexity of the best observer. The guarantee is dimension-free: it depends on observer complexity rather than on the latent dimension needed to represent the nonlinear system. To our knowledge, this yields the first end-to-end provable method for extracting a best-in-hindsight LDS representation of nonlinear dynamics through convex learning followed by provable distillation. Experiments on linear LDS benchmarks and MuJoCo behavior cloning show that the train-then-distill pipeline produces compact LDS predictors that match or outperform directly trained baselines.
[LG-46] Quantum-Structured World Models (QSWMs) for Predictive Latent Dynamics
链接: https://arxiv.org/abs/2608.05371
作者: Hailong Jiang,Emran Hossain,Feng Yu,Jianfeng Zhu,Guilin Zhang,Wulan Guo
类目: Machine Learning (cs.LG)
*备注: 19 pages, 5 figures,
Abstract:World models learn latent states that summarize interaction histories, evolve over time, and support prediction, simulation, or planning. Most existing world models represent these states using classical vectors, probability distributions, recurrent hidden states, or transformer activations. In this paper, we introduce Quantum-Structured World Models (QSWMs), a quantum-inspired framework for predictive world modeling with structured latent states, latent transition operators, and measurement-inspired decoding maps. We study whether mathematical structures inspired by quantum theory, such as complex-valued representations and density-matrix-like latents, provide useful inductive biases for world modeling. We establish three foundational properties: classical inclusion, predictive sufficiency, and structured compactness. We then instantiate complex-valued and density-matrix-like QSWM variants and evaluate them on elementary cellular automata against strong classical baselines. Results show promising local predictive potential for complex-valued QSWMs, while also revealing limitations in long-horizon rollout, density-matrix variants
[LG-47] DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting
链接: https://arxiv.org/abs/2608.05358
作者: Rahil Aftab,Vineet Kumar Rakesh,Soumya Mazumdar,Tapas Samanta
类目: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
*备注:
Abstract:Federated learning repeatedly incurs local optimization and model-update transmission. We study DG-FedReuse, a simulator-level mechanism that allows selected clients to contribute age-decayed cached updates when a stochastic head-gradient discrepancy proxy remains below a round-dependent threshold. A hard cache-age limit and minimum fresh-client quota constrain reuse, while fresh updates use an adaptive per-tensor Top-K numerical-field representation. Experiments cover six image-classification datasets, 50 virtual clients, Dirichlet label heterogeneity (\alpha=0.5), and three seeds. At a common 90-round budget, DG-FedReuse yields 83.36-85.42% modeled update-data-field uplink saving, compared with 76.88% for matched Top-K FedAvg; the seed-aligned accuracy differences range from -5.29 to -0.14 percentage points. Best-observed test accuracies obtained under test-controlled checkpointing are retained only as exploratory archival evidence and range from -2.38 to +0.45 percentage points relative to matched FedAvg. A symmetric dense-model-downlink sensitivity reduces the headline saving to 41.68-F42.71% and the incremental gain over Top-K FedAvg to 3.24-4.27 percentage points, demonstrating the dependence of communication conclusions on the accounting boundary. The study characterizes the proposed reuse rule in the implemented simulator; it does not establish unbiased generalization, end-to-end bandwidth reduction, runtime or energy savings, faster convergence, or superiority over existing stale-update and lazy-aggregation methods.
[LG-48] Computationally Efficient Collaborative Communication Via Regularity-Based Coarsening
链接: https://arxiv.org/abs/2608.05327
作者: Mark Bedaywi,Scott Emmons,Nika Haghtalab,Stuart Russell
类目: Computer Science and Game Theory (cs.GT); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG)
*备注:
Abstract:Our results show that the existence of a short high-utility protocol already suffices for efficient communication. In particular, in a game with n possible observations and m actions: (1) For any achievable target utility \alpha , we give an algorithm with \mathrmpoly(n, m, 1/\epsilon) runtime that designs a protocol achieving utility at least \alpha-\epsilon using only 2^\mathcal O(CC_\alpha(G))/\epsilon^2 bits of communication. Here, CC_\alpha(G) is the minimum number of bits used by any protocol, even a computationally inefficient one, to achieve utility \alpha . (2) We prove that this exponential dependence on CC_\alpha(G) is tight up to a constant. That is, unless \mathrm P=\mathrmNP , no polynomial-time algorithm can in general find optimal protocols using fewer than 2^CC_\alpha(G) -2 bits. We note that our results strictly weaken the assumptions required by prior work in the multi-agent information aggregation literature, filling a gap that had remained elusive even for games with constant CC_\alpha(G) . In particular, prior guarantees for agreement-based information aggregation rely on structural assumptions such as informational substitutes or weak learnability. We show that these assumptions already imply CC_\alpha(G) = O(1) and are therefore more restrictive conditions than required by our protocol to succeed. On a technical level, our results involve a novel strengthening of the Frieze-Kannan weak regularity lemma and yield the following powerful polynomial-time transformation tool: for every communication game G , it constructs a game \hat G that is a coarsening of the agents’ observation spaces into constant-size partitions, such that G and \hat G are indistinguishable with respect to every short communication protocol. This coarsening theorem is the engine behind our algorithm and may be of independent interest. Subjects: Computer Science and Game Theory (cs.GT); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG) Cite as: arXiv:2608.05327 [cs.GT] (or arXiv:2608.05327v1 [cs.GT] for this version) https://doi.org/10.48550/arXiv.2608.05327 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-49] Rectifying Geometric Misalignment: Online Source-Free Adaptation for Class-Imbalanced EEG
链接: https://arxiv.org/abs/2608.05315
作者: Shiwen Chu,Shanglin Li,Motoaki Kawanabe,Reinmar Kobler
类目: Machine Learning (cs.LG)
*备注: Accepted at EUSIPCO 2026
Abstract:Electroencephalography (EEG) based Brain-Computer Interfaces (BCIs) often require unsupervised domain adaptation (UDA) to generalize across subjects and sessions. While Riemannian alignment methods like the Riemannian Centering Transformation (RCT) are effective for handling covariate shifts, they implicitly assume balanced class priors. However, in realistic online BCI scenarios, the label distributions vary dynamically (label shift), causing standard alignment techniques to geometrically misalign the target data distributions. In this work, we propose OSPDIM (Online SPD manifold information maximization), a source-free online UDA framework designed to address label shifts on the Riemannian manifold. OSPDIM introduces a manifold-constrained bias parameter into the tangent space mapping, which is optimized via information maximization to correct the geometric skew caused by imbalanced data streams. Unlike offline methods relying on global batch statistics, OSPDIM estimates and corrects geometric bias on-the-fly. Simulations on 2D SPD matrices visually demonstrate that OSPDIM successfully rectifies the misalignment where standard centering fails. Extensive experiments on multiple motor imagery datasets show that OSPDIM significantly outperforms standard Riemannian baselines, particularly in challenging online adaptation scenarios with severe class imbalance, offering a robust solution for practical, plug-and-play BCI systems.
[LG-50] Beyond Rotations: AuroOFT for Expressive Quantized Orthogonal Fine-Tuning
链接: https://arxiv.org/abs/2608.05253
作者: Yue Han,Dianlin Wang
类目: Machine Learning (cs.LG)
*备注:
Abstract:Quantized orthogonal fine-tuning (qoft) enables parameter-efficient adaptation of low-bit language models by learning structured activation rotations before frozen quantized weights. However, its task-specific updates remain constrained to linear orthogonal transformations, limiting input-dependent nonlinear corrections. We introduce AuroOFT, which keeps qoft as a stable quantization-compatible branch while attaching a zero-start gated low-rank nonlinear residual to each adapted linear layer. AuroOFT maps activations into an RMS-normalized compact latent space and uses adaptive nonlinear bases with bounded or token-dependent gating. The zero-initialized up projection makes AuroOFT functionally identical to qoft at initialization, while orthogonality remains a branch-level stability property rather than a property of the combined nonlinear layer. Under matched data, optimization, decoding, and parser protocols, AuroOFT improves Macro-6 over matched qoft by 1.30-2.70% on the 1.5B/3B Qwen2.5 settings, exceeds QLoRA by 6.52-10.62%, and saves 32.3-44.7% trainable parameters relative to QLoRA in representative scales. The small exam-style multiple-choice math set is treated only as a protocol-sensitivity diagnostic. Our code is available at the anonymous repository: this https URL.
[LG-51] Beyond Full-Model Rollback: AuroSFT for Adapter-State Multi-Task Fine-Tuning
链接: https://arxiv.org/abs/2608.05250
作者: Yue Han,Ziniu Liu
类目: Machine Learning (cs.LG)
*备注:
Abstract:Multi-task supervised fine-tuning (SFT) often casts a heterogeneous data mixture as a single optimization problem, even though different tasks may reach their best generalization at different times. msft exposes this mismatch through task-wise roll-out, exclusion, and rollback, but its original formulation materializes the scheduler state as full-model checkpoints, making stage transitions costly to store, restore, and deploy. This paper introduces AuroSFT, a parameter-efficient framework that recasts the carried state of overfitting-aware multi-task SFT as a compact, mergeable adapter state. AuroSFT freezes the pretrained backbone, trains only injected adapters, rolls back adapter checkpoints at task-wise peaks, and continues on the remaining active mixture. At the layer level, each adapter applies an AuroRA-inspired adaptive nonlinear layer to a low-rank weight factor rather than to the sample representation. The resulting update remains linear in the input, rank-bounded, and exactly mergeable into the frozen projection. Under the retained-backbone comparison protocol, AuroSFT achieves 61.36% average accuracy, compared with 59.85% for the corresponding msft reference row, and obtains higher accuracy on all five backbones. Our code is available at the anonymous repository: this https URL.
[LG-52] Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language AAAI2027
链接: https://arxiv.org/abs/2608.05238
作者: Xinran Feng,Yi Xie,Chao Zhang,Ruikun Li,Wanyun Ling,Ziyue Li,Chenxi Liu
类目: Machine Learning (cs.LG)
*备注: 46 pages, 9 figures, including supplementary material. Submitted to AAAI 2027. Xinran Feng and Yi Xie contributed equally
Abstract:Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is supposed to learn. The data can never teach more than the labeler already knows. A second gap makes this worse: most datasets use a single variable, but the patterns that matter (cross-channel correlation, lead-lag structure, co-occurring anomalies) appear only with several variables, right where the labeling LLM’s limits are most exposed. These two problems create a trilemma: existing methods are reliable, realistic, or scalable, but none achieves all three. We resolve this by decoupling perception from description. Deterministic code computes a set of statistics from real, open-source multivariate series; the LLM verbalizes those precomputed facts. Perception, which LLMs do poorly, is handled by computation, while the LLM handles expression. This produces CGTime, our 4B-parameter computation-grounded time-series-language model. CGTime outperforms far larger general-purpose models on multivariate understanding tasks: it attains the best multivariate fact score on our held-out benchmark (0.283 vs. 0.173 for GPT-4o-mini and 0.203 for GPT-5.4-nano), a gap that survives Holm-corrected paired significance tests against every baseline. It also states verifiable numerical facts in generated captions more accurately and covers a broader range of statistical properties.
[LG-53] PPDL: LLM -Based Flows as Probabilistic Programs ICML2026
链接: https://arxiv.org/abs/2608.05234
作者: Louis Mandel,Guillaume Baudart,Mandana Vaziri,Martin Hirzel
类目: Machine Learning (cs.LG); Programming Languages (cs.PL)
*备注: Published at ICML 2026
Abstract:Building reliable applications that leverage large language models (LLMs) remains a significant challenge. While LLMs offer impressive capabilities across diverse tasks, their outputs often lack accuracy and provide no clear measure of confidence. This uncertainty compounds in flows of multiple calls to LLMs and other tools, making it difficult for developers and end-users to trust the results. This paper introduces a probabilistic language for programming LLM-based flows. It enables developers to quantify and propagate uncertainty throughout the application’s flow, and experiment with different inference scaling techniques without adding a single line of code beyond the flow’s logic. We present an experimental study to demonstrate this capability, and a case study building a theorem proving agent for the Rocq theorem prover.
[LG-54] When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters
链接: https://arxiv.org/abs/2608.05207
作者: Fangxin Wang,Ziyi Zhang,Diyi Zhuang,Langzhou He,Shiyu Wang,Baichuan Mo,Philip S. Yu
类目: Machine Learning (cs.LG)
*备注: 23 pages, 14 tables, 4 figures
Abstract:Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster’s residual to drive a lightweight post-hoc corrector. Prior automated feature engineering models the data-generating process; corrective features instead model the model-failure process. We present CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model (LLM) that proposes named feature combinations, binary flags, and short executable code. A single validation-grounded gate accepts or rejects every candidate regardless of its origin, and a validation-selected corrector applies the accepted features or leaves the forecast unchanged. This source-agnostic pipeline also allows prior feature-engineering systems to be evaluated under identical conditions, making CRAFTER an instrument for attributing forecast improvements to the feature source alone. Across six public datasets and six frozen backbones, CRAFTER surpasses every dedicated feature-engineering system at every feature budget, roughly doubling the improvement achieved by the corrector alone and reducing the error of the weakest backbones by up to 27%. These gains are robust across different LLM backends and persist even when applied on top of fine-tuned backbones.
[LG-55] MS-MLB: An Open Machine Learning Benchmark for Blood-Based MS Classification
链接: https://arxiv.org/abs/2608.05196
作者: Adam Simson,Ankush Dutta,Quang Bui
类目: Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注:
Abstract:Multiple sclerosis (MS) is diagnosed through clinical assessment, magnetic resonance imaging, laboratory evidence when appropriate, and exclusion of better explanations. Blood RNA expression data may contain disease associated immune signal, but a blood RNA classifier cannot be treated as a replacement for clinical diagnosis. This paper presents MS-MLB (Multiple Sclerosis Machine Learning Benchmark), a reproducible open benchmark for machine learning based MS research classification from whole blood RNA expression data. MS-MLB uses the public GSE17048 cohort, converts it into an MS versus healthy control task, and evaluates multiple algorithms under a shared, leakage controlled pipeline that a researcher can rerun without reconfiguring the evaluation. The evaluation includes nested cross-validation, an untouched stratified holdout set, bootstrap confidence intervals, ROC and precision recall analysis, calibration measurement, and an exploratory MS Research Score. In the final benchmark summary, Gradient Boosting ranked first by MS Research Score on the holdout set, with an MS Research Score of 93.83, AUC-ROC of 0.989, sensitivity of 0.950, specificity of 0.778, F_1 score of 0.927, and Brier score of 0.050. Prior studies have applied machine learning to MS blood transcriptomic data, including PBMC stage classification and whole blood diagnostic signature modeling. The contribution here is different and narrower. To our knowledge, MS-MLB is the first open benchmark focused on MS versus healthy control classification from GSE17048 whole blood RNA expression data with a documented external model submission pathway built into the framework. The score is intended for research comparison only and has not been clinically validated. The benchmark is accessible here: this https URL.
[LG-56] Scalable estimation of VARMA models
链接: https://arxiv.org/abs/2608.06340
作者: Daniel Paulin,Victor Elvira
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注: 60 pages, 1 figure
Abstract:Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass over the entire series. Yet their moving-average term captures with a few parameters what a pure autoregression matches only with many lags. We introduce an estimation framework that removes this computational barrier: each optimization iteration is independent of the series length T . The framework combines a partial-autocorrelation reparametrization that guarantees stationarity and invertibility by construction, Gaussian priors on the reparametrized coefficients with separate scales for diagonal and off-diagonal entries, and losses that depend on the data only through fixed-size sufficient statistics, evaluated by a Parseval (Fourier) identity at near-linear cost in the truncation length. This yields two point estimators: a regularized least-squares fit and a covariance-marginalized maximum-a-posteriori estimator. We prove that both recover the infinite-autoregressive representation of the true process at a near-parametric rate in fixed dimension, so the truncation introduces no asymptotic bias. The same machinery extends, at the same leading cost, to seasonal dynamics, exogenous regressors (VARMAX), and rolling-window refits. Empirically, the estimators stay close to the oracle forecast error from d=10 to d=40 (where classical conditional MLE returns non-invertible fits whose forecasts diverge) and match or beat VAR, Bayesian-VAR, component-wise ARMA, and sparse-VARMA baselines on retail-demand, meteorological, and air-quality data. This brings likelihood-based VARMA estimation, at a per-iteration cost independent of the series length, to the problem sizes where practitioners have so far relied on VAR models.
[LG-57] Optimal Rates for Learning with Monotone Adversaries
链接: https://arxiv.org/abs/2608.06337
作者: Anay Mehrotra
类目: Machine Learning (stat.ML); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG); Statistics Theory (math.ST)
*备注:
Abstract:A monotone adversary observes an i.i.d. labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. The learner sees a uniform shuffle of the combined sample and is scored on the original distribution. Every example is correctly labeled, but the insertions depend on the clean sample, so the combined sample is not exchangeable. Larsen, Pabbaraju, and Shetty, who introduced this model, showed that empirical risk minimization attains expected error O((d/n)\log(n/d)) for classes of VC dimension d , and that every known optimal learner can be pushed away from the \Theta(d/n) rate, optimal for PAC learning. They asked whether the extra logarithm is an artifact of those particular algorithms or an inherent consequence of the lack of exchangeability. We show that this additional cost is inherent beyond VC dimension one. In the worst case over classes of VC dimension d and over known finite insertion budgets, the minimax expected error is \Theta(1/n) at d=1 and \Theta((d/n)\log(n/d)) for d\geq 2 . The same rates hold with Littlestone dimension d_\mathrm L in place of d , so the clean online-to-batch rate O(d_\mathrm L/n) is unattainable as well. Thus, somewhat counterintuitively, adding correctly labeled examples can make learning harder by a logarithmic factor, even for classes that admit finite mistake bounds in online learning. The dimension-one upper bound is achieved by a simple improper learner whose analysis adapts the leave-one-out argument underlying the one-inclusion graph. All of our lower bounds are elementary and come from a single construction: an explicit class and prior on which two target hypothesis, which differ a point of nonnegligible mass, produce the same sample. Subjects: Machine Learning (stat.ML); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG); Statistics Theory (math.ST) Cite as: arXiv:2608.06337 [stat.ML] (or arXiv:2608.06337v1 [stat.ML] for this version) https://doi.org/10.48550/arXiv.2608.06337 Focus to learn more arXiv-issued DOI via DataCite (pending registration)
[LG-58] Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning
链接: https://arxiv.org/abs/2608.06276
作者: Farzana Nasrin
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Algebraic Topology (math.AT)
*备注: 27 pages, 7 figures, and 5 tables
Abstract:Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure. While substantial progress has been made in the statistical analysis of PDs, existing literature often treats diagrams as static objects and provide limited frameworks for probabilistic modeling and stochastic evolution on PD space. We introduce a reinforcement learning framework for stochastic dynamics on PD space, where diagrams evolve through topology aware local edit operations. The dynamics define controlled Markov processes on spaces of finite PDs with variable cardinality. We establish conditions under which the induced Markov chains are irreducible, aperiodic, and geometrically ergodic, implying the existence of unique stationary probability laws on PD space. To guide the dynamics toward scientifically relevant topological targets, we formulate objectives that encompass distribution matching, task specific topological statistics, and structure-preserving compression. The resulting rewards balance task specific distributional targets, diagram fidelity, and complexity reduction, and yield a framework for adaptive topological simplification and probabilistic modeling. Experiments on synthetic and neuroimaging PDs demonstrate that the proposed framework can preserve dominant topological structure while reducing diagram complexity.
[LG-59] Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification
链接: https://arxiv.org/abs/2608.06250
作者: Alex Buna,Shirley Xiaoqi Liu,Patrick Rebeschini
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. In this work, we show that early stopping can overcome this suboptimality: in a Gaussian mixture model with label-flipping noise, GD stopped at an appropriate oracle time achieves minimax-optimal excess zero-one risk for covariance spectra with fast and continuous decay, including polynomial and exponential spectral decays. Our analysis combines a sharp upper bound for the early-stopped iterate with a matching statistical lower bound over arbitrary classifiers, yielding optimal rates that are validated by experiments. A central technical contribution is a new calibration result that converts excess logistic risk into excess zero-one risk; it handles the model misspecification induced by the label-flipping noise, and removes the square-root rate in standard bounds. We also establish a lower bound for linear interpolators, showing that interpolation can require exponentially more samples than early stopping to achieve the same excess risk.
[LG-60] Muon on the Stiefel Manifold Admits an Exact Closed-Form Update
链接: https://arxiv.org/abs/2608.06218
作者: Mikhail Solonko,Molozhavenko Alexander,Maxim Rakhuba
类目: Optimization and Control (math.OC); Machine Learning (cs.LG); Numerical Analysis (math.NA)
*备注:
Abstract:We study Muon, a recently proposed matrix-aware optimization method, in the context of the Stiefel manifold. This manifold consists of matrices with orthonormal columns and is ubiquitous in machine learning and scientific computing. Existing extensions of Muon to this manifold rely on heuristic, approximate, or iterative updates with varying computational efficiency. We show that the corresponding Stiefel Muon update admits an exact closed-form solution and use this result to develop Skewon, a practical algorithm for orthogonality-constrained optimization with an efficient implementation. We further establish first-order convergence guarantees for Skewon in the smooth non-convex setting.
[LG-61] Beyond Marginal Validity: Finite-Sample Guarantees for Localized Conformal Prediction
链接: https://arxiv.org/abs/2608.06206
作者: Anton Conrad,Rustam Isaev,Denis Belomestny,Eric Moulines,Sergey Samsonov
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: 68 pages, 8 figures, 2 tables
Abstract:Conformal prediction endows arbitrary black-box predictors with finite-sample, distribution-free marginal coverage, yet marginal validity can hide severe covariate-specific miscalibration, while exact distribution-free conditional coverage is finite-sample unattainable. Randomly localized conformal prediction (RLCP) mitigates this gap by calibrating near the test point while preserving marginal coverage. Existing theory, however, lacks finite-sample guarantees for the realized localized set that jointly control conditional validity and oracle efficiency. We provide such guarantees. For any fixed score, under Hölder regularity of the conditional score CDF and standard density and kernel assumptions, we prove high-probability bounds, uniform over a realized localization neighbourhood, for the conditional-coverage gap and the length error relative to the oracle. The bounds decompose into an O(h^\beta) localization bias and a calibration term decreasing with calibration size, clarifying the bandwidth bias-variance tradeoff and when RLCP tracks the oracle. We also analyze data-split learned scores: when the score targets a pivotal score, as in conformalized quantile regression, uniform local guarantees decompose into fixed-score calibration and uniform score-estimation errors, showing that improved learning sharpens localized guarantees.
[LG-62] Handling Missing Data in Probabilistic Regression Trees
链接: https://arxiv.org/abs/2608.06195
作者: Taiane Schaedler Prass,Alisson Silva Neimaier,Guilherme Pumi
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注: Theoretical background for the companion paper, arXiv:2510.03634
Abstract:Probabilistic Regression Trees (PRTrees) are a smooth and consistent alternative to classical regression trees, producing continuous predictions through probabilistic split assignments. This paper extends the PRTree framework to accommodate missing predictor values directly during tree construction, eliminating the need for prior imputation. Three strategies are proposed, each exploiting the available information differently: a uniform-probability approach, a partial-observation approach, and a dimension-reduced smoothing approach. These modifications are defined to preserve the fundamental probabilistic properties of the original methodology, including probability conservation and marginal compatibility, under arbitrary patterns of missing covariate values. The proposed methods are evaluated on several real-world datasets exhibiting different levels of missingness and are compared with classical regression trees. The results show that the effectiveness of probabilistic tree construction depends strongly on the treatment of missing observations. Across the considered datasets, the fill strategy emerged as the dominant modeling component, often exerting a larger influence on predictive performance than either the smoothing distribution or the proxy-selection criterion. In datasets where a substantial proportion of observations contained missing predictor values, the proposed methods frequently outperformed CART, while maintaining the interpretability and flexibility of tree-based models.
[LG-63] On Same-Sample and Independent-Sample Stochastic Extrag radient for Monotone Variational Inequalities
链接: https://arxiv.org/abs/2608.06182
作者: TaeHo Yoon,Nicolas Loizou
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:We study stochastic extragradient (SEG) methods for solving monotone variational inequality problems (VIPs) over a feasible set. Although extragradient is a foundational algorithm for VIPs and its deterministic convergence theory is well developed, its stochastic counterpart remains less understood. Most existing analyses focus on independent-sample SEG (I-SEG) and assume either that the domain is compact or that the variance of the stochastic operator is uniformly bounded. The behavior of same-sample SEG (S-SEG), a natural variant with materially different properties, has received far less attention. In this work, we address these gaps in the literature. We first show that S-SEG is sensitive to samplewise Lipschitz parameters: mean Lipschitzness and bounded variance alone do not ensure convergence, even on a compact set. Then, for possibly unbounded domains, we establish a high-probability restricted-gap convergence for each SEG variant under a relaxed set of assumptions, and show that certain fundamental improvements to these results are impossible in general. Finally, we show that a known asymmetric double step-size selection that guarantees almost sure last-iterate convergence for I-SEG can fail for S-SEG: there exists a stochastic monotone VIP for which S-SEG diverges almost surely even under the modified step-sizes.
[LG-64] Verifiable Regularity Criterion for Conditional Expectation Operators and Conditional Mean Embeddings with Applications to Nonparametric Regression Bayesian Inverse Problems and Koopman Operators
链接: https://arxiv.org/abs/2608.06155
作者: Maximiliano Hertel,Ilja Klebanov,Manuel Schaller,Karl Worthmann
类目: Dynamical Systems (math.DS); Machine Learning (cs.LG); Numerical Analysis (math.NA); Statistics Theory (math.ST); Machine Learning (stat.ML)
*备注: 37 pages, 3 figures
Abstract:Conditional expectation operators (CEOs) and their associated conditional mean embeddings (CMEs) play a central role across applied mathematics and machine learning, appearing in nonparametric regression, Bayesian inverse problems, and Koopman operator theory. A fundamental question is when a CEO maps a function space on \mathcalY into a prescribed function space on \mathcalX , particularly a reproducing kernel Hilbert space (RKHS). We show that such mapping properties are characterized by the regularity of the Radon–Nikodym density of the conditional law, and establish a simple, verifiable sufficient condition under which the CEO is bounded and Hilbert–Schmidt. For RKHSs norm-equivalent to Sobolev spaces, this condition reduces to Sobolev regularity of the conditional density. The result yields a direct route to validate CME representations and error bounds for Galerkin-type and CME-based estimators. We verify the regularity condition in three settings: nonparametric regression, Bayesian inverse problems, and Koopman operator theory for stochastic dynamical systems. We show in each case that classical regularity results on the underlying probabilistic model imply the required mapping properties. The resulting framework offers a unified perspective on conditional expectation operators across probability, operator theory, kernel methods, and stochastic dynamics.
[LG-65] Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data
链接: https://arxiv.org/abs/2608.05930
作者: Nina van Gerwen,Dimitris Rizopoulos,Manon Hillegers,Loes Keijsers,Sten Willemsen
类目: Machine Learning (stat.ML); Machine Learning (cs.LG)
*备注:
Abstract:The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multiple times a day. Our work is motivated by such data collected by the GrowIt! app, which was released to investigate daily emotions among adolescents during the COVID-19 pandemic. Current procedures to analyse ESM data face various challenges. While standard statistical techniques may not scale well to a high-dimensional setting, machine learning procedures can give biased results due to selection bias introduced by missingness. In our motivating dataset, adolescents dropped out due to previous strong feelings of negative emotions. Hence, the implied missing data are of the missing-at-random type that standard machine learning procedures cannot accommodate. We develop a novel neural network architecture that generalises mixed effects models to deep learning to overcome these challenges. It allows semi-parametric and flexible modelling of data’s mean and correlation structure through fixed and random effects. For estimation, we use an adaptation of variational auto-encoders and a Bayesian data augmentation algorithm. Through this approach, the model can accommodate longitudinal outcomes following generic distributions, scale well to high-dimensional settings and provide valid inference when data are missing-at-random. We applied the Deep Generalised Mixed Model to the GrowIt! study and various simulations. The results show potential for the Deep Generalised Mixed Model, yet suboptimal performance due to model instability.
[LG-66] A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring
链接: https://arxiv.org/abs/2608.05697
作者: Mohammad Hosseini,Hamed Khatounabadi,Mohammad Fakharzadeh
类目: Quantitative Methods (q-bio.QM); Machine Learning (cs.LG); Signal Processing (eess.SP)
*备注: Code available at this URL: this https URL
Abstract:Respiration provides a continuously available window into physiological state and behavior. However, monitoring it outside controlled settings remains challenging because a wearable system must capture small body deformations while remaining comfortable, low power, and robust to changes in posture and motion. We present a compact non-invasive respiratory sensing system based on a force-sensitive resistor (FSR) embedded in an abdominal belt and integrated with a custom Bluetooth Low Energy acquisition board. The system combines a simple piezoresistive readout with a mechanical holder designed to transfer abdominal expansion to the sensor without analog amplification. We evaluate the complete sensing pipeline across multiple breathing patterns and body positions. In stationary settings, the recorded signals exhibit consistent amplitude changes and recurring peak-to-peak timing across breathing maneuvers; under light movement, these variations remain visible despite motion-induced baseline shifts. We further design a five-phase stress-induction protocol and collect respiratory recordings from 12 participants. Using interpretable time-domain features and standard classifiers, we examine whether the acquired signals distinguish relaxation from stress-induction phases. In this preliminary experiment, the best-performing model achieves 88.0% test accuracy, indicating that the extracted respiratory features distinguish stress-induced phases from relaxation phases in this dataset. Overall, our results show that the proposed platform enables real-time respiratory monitoring across diverse daily-life scenarios and captures respiratory changes that distinguish stress-induction from relaxation phases, supporting its potential for affective-computing applications.
[LG-67] Provably Efficient Self-Calibrating Quantum Fault Tolerance
链接: https://arxiv.org/abs/2608.05686
作者: Weiyuan Gong,Hong-Ye Hu
类目: Quantum Physics (quant-ph); Machine Learning (cs.LG)
*备注: 64 pages, 12 figures
Abstract:Quantum error correction protects logical information only when every physical operation remains below the fault-tolerance threshold, a condition that must be maintained continuously rather than only at the initial calibration. In practice, however, analog control parameters inevitably drift because of environmental fluctuations. As future fault-tolerant quantum computations are expected to run for days or even months, interrupting computation for repeated recalibration becomes fundamentally impractical. A promising alternative is to integrate calibration directly into computation by repurposing syndrome measurements as a calibration signal (Sivak et al, Nature 2026), but whether such self-calibration can be achieved with provable efficiency remains an open question. Here we establish a theoretical framework for self-calibrating quantum fault tolerance. We prove that, for a broad class of control-induced errors, the detection rate defines a locally strongly convex surrogate objective for analog calibration with high probability. This geometric property enables efficient online optimization using only syndrome measurements collected during normal error correction. We prove convergence to an \varepsilon detection rate within O(1/\varepsilon^2) epochs for time-independent drifts and also establish guarantees for time-dependent drifts. We further show that the convergence rate is independent of the code distance for quantum low-density parity-check (LDPC) codes. Pulse-level simulations of neutral-atom arrays and large-scale circuit-level Clifford simulations confirm these theoretical predictions. Our results establish self-calibrating fault tolerance as a provably efficient paradigm in which the same syndrome measurements simultaneously protect logical information and stabilize the underlying hardware.
[LG-68] How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits
链接: https://arxiv.org/abs/2608.05595
作者: Prabhjot Singh,Adel N. Toosi,Rajkumar Buyya
类目: Quantum Physics (quant-ph); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
*备注: 11 pages, 6 figures. Includes technical appendix
Abstract:Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponential in the number of cuts - the dominant runtime cost in prior work. We ask whether, for machine-learning tasks, this step is necessary, and replace it with late fusion: each subcircuit is trained and measured independently, and a small classical head combines their outputs - a linear-cost, decision-level combination borrowed from multimodal learning. To characterize the trade-off we introduce a quantumness dial Q , a tunable reconstruction budget interpolating from pure fusion to full reconstruction, and a cut-entanglement diagnostic that indicates how much reconstruction a task needs (Spearman \rho=0.59 over 104 runs). Across synthetic and standard datasets, independently trained late fusion matches full reconstruction accuracy within 0.04 at every point of the controlled sweep and on every classical benchmark, at exponentially lower cost; it is also markedly more robust to shot and device noise. Controlled entangled-data experiments locate the boundary where fusion must fail. We do not claim advantage over classical machine learning - consistent with recent benchmarking, quantum offers no accuracy edge on these datasets. Late fusion is thus an efficient, noise-robust, self-characterizing alternative to reconstruction for circuit-cutting QML.
[LG-69] Equation-Free Period-Aware Forecast-Error Contraction for Estimating Negative Largest Lyapunov Exponents from Short Trajectory Ensembles
链接: https://arxiv.org/abs/2608.05522
作者: Andrei Velichko,N’Gbo N’Gbo,Viet-Thanh Pham
类目: Chaotic Dynamics (nlin.CD); Machine Learning (cs.LG); Data Analysis, Statistics and Probability (physics.data-an)
*备注: 13 pages, 3 figures, 1 table, 37 references
Abstract:Estimating positive largest Lyapunov exponents from data is comparatively natural because neighboring trajectories separate, whereas stable dynamics require resolving contraction before measurement noise or finite precision erases the signal. We introduce a period-aware forecast-error contraction procedure for estimating a dominant negative Lyapunov exponent from ensembles of short scalar trajectories without using governing equations or an analytical Jacobian. A k-nearest-neighbor predictor is trained on trajectory histories, the geometric-mean absolute forecast error is evaluated at phase-consistent horizons, and the exponent is obtained from the slope of the logarithmic error profile. Unlike data-driven approaches that reconstruct local evolution matrices or differentiate a learned surrogate, the proposed method extracts the contraction rate directly from out-of-sample forecast errors. Two adaptations are essential: the forecast step is synchronized with the detected orbit period, and candidate slopes are accepted only when they form a stable consensus across several transient lengths. On the logistic map, the method recovers 92 of 112 negative-exponent parameter values with a mean absolute error of 0.0253 and R^2=0.886 . On a two-dimensional map without fixed points, independent scalar pipelines based on the three observables x_n , y_n , and z_n give mean absolute errors of 0.00879–0.01145 and R^2=0.983 – 0.986 . Because the estimation stage uses only observed trajectories, the framework provides a basis for repeated-relaxation experiments in which short sensor responses are available but the governing equations and analytical Jacobian are unknown. Experimental validation remains a subject of future work.
[LG-70] An Inertial Block Proximal Linearized Method with Adaptive Momentum for Nonconvex and Nonsmooth Optimization
链接: https://arxiv.org/abs/2608.05502
作者: Weifeng Yang
类目: Optimization and Control (math.OC); Machine Learning (cs.LG)
*备注:
Abstract:In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning. To solve this class of problems, we propose the inertial block proximal linearized method with two-phase adaptive momentum (IBPL ^+ -TP). Compared to the current methods, our method possesses three main advantages: (1) it introduces a two-phase adaptive momentum strategy to effectively update the extrapolation parameters, (2) it allows using two different extrapolation points to accelerate the convergence, (3) it allows the extrapolation parameters of these two extrapolation points to be independent of and unconstrained by all other parameters. While maintaining the above advantages, we prove that our method ensures the monotonic convergence of the objective function of this class of problems, and we also prove that the sequence generated by our method globally converges to a critical point, as well as establish the convergence rate of our method. To demonstrate the effectiveness of our method, we apply it to solve two nonconvex and nonsmooth machine learning problems, namely sparse nonnegative matrix factorization with \ell_0 -constraints and sparse nonnegative CP decomposition with \ell_0 -constraints. The numerical experimental results on solving these problems show that our method outperforms several state-of-the-art methods.
[LG-71] Effective pruning of task-trained recurrent neural networks using noisy fluctuations and connection rescaling
链接: https://arxiv.org/abs/2608.05464
作者: Sanjith Senthil,Rishidev Chaudhuri
类目: Neurons and Cognition (q-bio.NC); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE)
*备注:
Abstract:The pruning of network connections is key to brain function but, despite its importance, there exist few biologically-plausible pruning rules with demonstrated good performance. In this work we evaluate noise-prune, a recently introduced unsupervised local pruning rule for recurrent networks that uses noisy fluctuations to determine the importance of connections. Noise-prune has previously only been empirically tested on random networks without a specific computational function. We show that noise-prune preserves task-performance in task-trained recurrent neural networks, greatly outperforming a strategy that only uses the magnitude of connections and performing on par with or exceeding a non-local strategy that uses second-order information. Rather than deterministically removing connections that fall below a certain threshold importance, noise-prune samples connections to preserve based on their importance and strengthens retained connections to preserve average synaptic strength. We show that this sampling and rescaling is essential to good performance, but that the optimal empirical degree of rescaling is lower than that predicted by the original theoretical argument. Our work thus validates noise-prune as a biologically-plausible pruning rule for functional recurrent network architectures and characterizes its optimal parameter settings.
[LG-72] Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation with Explainable Attribution
链接: https://arxiv.org/abs/2608.05373
作者: Alex Chen,Maria Hybinette
类目: Trading and Market Microstructure (q-fin.TR); Machine Learning (cs.LG); Statistical Finance (q-fin.ST)
*备注:
Abstract:Intraday market manipulation is hard to detect because its footprint is brief, buried in millions of quotes, and statistically similar to ordinary volatility. Detectors reach high recall only by flagging so many other days that measured precision collapses, producing alerts no regulator can act on. We show that this manipulation leaves a distinctive dynamic signature: a pump-and-crash pattern visible in the velocity of market state, rather than its level. We build a minute-level detection pipeline, strictly partitioned in time, based on smoothed state velocity: option-Delta velocity for index options and price velocity for equities. We explain every alert with SHAP attribution. We hold the test period strictly out-of-sample and fix all thresholds before evaluation. On the locked Indian BANKNIFTY index-options test, the plain autoencoder recovers 10 of 10 regulator-identified manipulation days. Conditioning detection on market regimes inferred by a hidden Markov model yields an instructive negative result. The regimes are descriptively distinct, but using them trades recall for precision. Under the closed-world assumption that unlabeled days are normal, precision remains near 25%. The same dynamic appears in thinly traded U.S. equities (SEC v. Patel). The shape of the signature survives the transfer; its velocity magnitude does not. A pump-reversal shape score ranks the complaint’s alleged manipulation days with AUC 0.91 (ARQQ) and 0.81 (ACY). On the ARQQ worked example, the score peaks inside the complaint’s documented minute window. Finally, exact SHAP attribution over every alert shows that unconfirmed alerts share the regulator-identified days’ attribution profile (cosine similarity 0.99). The precision ceiling is consistent with incomplete enforcement labels rather than detector failure. What transfers across markets and instrument types is the dynamic signature itself. Subjects: Trading and Market Microstructure (q-fin.TR); Machine Learning (cs.LG); Statistical Finance (q-fin.ST) Cite as: arXiv:2608.05373 [q-fin.TR] (or arXiv:2608.05373v1 [q-fin.TR] for this version) https://doi.org/10.48550/arXiv.2608.05373 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Maria Hybinette [view email] [v1] Wed, 5 Aug 2026 19:51:54 UTC (208 KB) Full-text links: Access Paper: View a PDF of the paper titled Velocity- and Regime-Aware Detection of Intraday Options Market Manipulation, with Explainable Attribution, by Alex Chen and Maria HybinetteView PDFHTML (experimental)TeX Source view license Current browse context: q-fin.TR prev | next new | recent | 2026-08 Change to browse by: cs cs.LG q-fin q-fin.ST References Citations NASA ADSGoogle Scholar Semantic Scholar export BibTeX citation Loading… BibTeX formatted citation loading… Data provided by: Bookmark checked="checked"class=“labs-tab-input”> Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs. Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?) mathjaxToggle(); We gratefully acknowledge support from our major funders, member institutions, , and all contributors. About Help Contact Subscribe Copyright Privacy Accessibility Operational Status (opens in new tab) Major funding support from
[LG-73] Physics-Based Molecular Fingerprints from Spectral Graph Theory Provide Efficient Geometry-Aware Measures of Chemical Similarity
链接: https://arxiv.org/abs/2608.05336
作者: Jacob W. Toney,Ayleen Y. Farnood,Samir Darouich,Heather J. Kulik
类目: Chemical Physics (physics.chem-ph); Machine Learning (cs.LG)
*备注:
Abstract:Molecular representations are essential for the evaluation of molecular similarity and the development of structure-property relationships. Despite the known importance of 3D structure to determine chemical and physical properties, the most widely used molecular fingerprints encode only two-dimensional connectivity. Such representations fail to distinguish similar but distinct stereoisomers and conformers. Alternative 3D methods are typically defined pairwise, making their application to large chemical spaces prohibitive, while deep learning embeddings are expressive but uninterpretable and limited by their training data diversity. Here, we introduce novel physics-inspired molecular fingerprints based on principles from spectral graph theory. We represent molecules as a complete graph in 3D space, with edge weights encoding heuristic physical interactions. Eigenvalue decomposition of the resulting graph Laplacian matrix results in a computationally efficient fixed-length chemical fingerprint that encodes 3D structure while obeying necessary physical symmetries of permutation and E(3) invariance. Spectral fingerprints differentiate between unique molecular structures with identical 2D connectivity, overcoming a limitation of 2D descriptors, while maintaining the low computational cost needed for efficient screening of vast chemical spaces. We evaluate our fingerprints with community detection algorithms and observe strong performance against representative baselines across datasets from organic, inorganic, biological, reticular, and reaction chemistry. Nearest-neighbor property estimation and applicability domain analyses reveal the utility of our molecular representation in machine learning and cheminformatics. We anticipate that spectral fingerprints will serve as generalizable, interpretable, and efficient measures of chemical similarity that incorporate 3D information at minimal cost.
[LG-74] A Unified Causal Inference Framework for the Desirability of Outcome Ranking Paradigm in Benefit-Risk Evaluation
链接: https://arxiv.org/abs/2608.05244
作者: Yuan Feng,Shiyu Shu,Yixin Fang,Ionut Bebu,Toshimitsu Hamasaki,Scott Evans,Guoqing Diao
类目: Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME)
*备注:
Abstract:We developed a unified covariate-adjusted causal inference framework for estimating the desirability of outcome ranking (DOOR) probability for benefit-risk evaluation in randomized trials and observational studies. The framework expresses the DOOR probability as a bilinear functional of the marginal ordinal outcome distributions under the two treatment strategies, estimates conditional ordinal distributions through sequential risk-set hazards, and derives the efficient influence function (EIF) of the DOOR probability. The point-estimation simulations compared G-computation, normalized inverse probability weighting (IPW), augmented IPW (AIPW), and targeted maximum likelihood estimation (TMLE), with nuisance functions estimated using generalized linear models or Super Learner (SL). TMLE-SL showed the strongest and most consistent point-estimation performance, with AIPW-SL ranking second. EIF-based inference was then evaluated for AIPW-SL and TMLE-SL, with and without cross-fitting, across settings varying in overlap, treatment-effect heterogeneity, and treatment allocation. CVTMLE-SL showed the strongest overall performance across DOOR-scale bias, recovery of the underlying ordinal distributions, standard-error accuracy, and confidence-interval coverage. We illustrate the methodology using data from the multidrug-resistant organism network of the Antibacterial Resistance Leadership Group.
[LG-75] CLARA: Clarification of Language Ambiguity through Result Analysis for Natural-Language Cancer Genomics Queries
链接: https://arxiv.org/abs/2608.05195
作者: Pratyush Kumar Shukla,Manveer Singh Tib,Siddhant Garg
类目: Genomics (q-bio.GN); Machine Learning (cs.LG); Quantitative Methods (q-bio.QM)
*备注: Preprint of an article submitted for consideration in the Pacific Symposium on Biocomputing (PSB 2027). Copyright 2026 World Scientific Publishing Company. this https URL
Abstract:A natural language interface can be used to make cancer genomics databases easier to use, but even if a question is perfectly fluent, its scientific meaning can be ambiguous. We propose CLARA, a framework that represents a question as a typed scientific query specification, considers a few possible interpretations, executes them, and asks for clarification when the estimates diverge. CLARA was assessed on mutation-prevalence contrasts among eight TCGA PanCancer Atlas cohorts and a 30-gene panel. This benchmark consisted of 330 unique executable contrasts varying in mutation scope, assay denominator, and sample context; 115 contrasts were result-sensitive and 215 were result-stable, per the preregistered definition of relative divergence greater than 0.10 or absolute divergence greater than 5 percentage points. An independently implemented pandas execution engine perfectly replicated all 660 results from the SQLite engine. In a separate 120-question LLM-generated, manually vetted language stress test, CLARA recognized all 60 result-sensitive contrasts and needlessly clarified 13 of 60 stable contrasts (accuracy 89.2%, sensitivity/recall 100%, specificity 78.3%). Standalone machine learning had superior overall accuracy (97.5%) but missed one critical contrast. This demonstrates that downstream execution can distinguish consequential from inconsequential ambiguity and reveal an explicit trade-off between safety and burden.
附件下载


