Eye on AI Weekly Research Watch

Eye on AI Weekly Research Watch

Craig Spencer Smith
Land Verenigde Staten
Taal EN
Afleveringen 40
Laatste 07.10.2026

Eye on AI Weekly Research Watch provides weekly, digestible podcast explainers of significant research papers in the field of artificial intelligence. Each episode breaks down complex AI research into accessible summaries for a broad audience. The podcast aims to keep listeners informed about the latest developments and breakthroughs in AI research.

Afleveringen

  • HazardWeaver: Scientific Route Selection for Hazard Analysis Agents 07.10.2026 2min
    Natural hazard analysis requires choosing scientific methods that are both appropriate for an event and executable with available data and tools, and those choices may need revising as evidence emerges. HazardWeaver frames this as state-dependent route selection. A Hazard Knowledge Compiler extracts applicability conditions, a Hazard Capability Graph checks input–output compatibility between executable capabilities, and an agent selects, runs, and revises workflows. Evaluated on a 141-instance benchmark spanning seven hazards and four multi-hazard interaction classes, it outperforms existing agents, especially with multiple valid routes. Authors: Wangshu Zhu, Xueqi Cheng, Liang Wu, Yushun Dong Paper: https://arxiv.org/abs/2610.03591v1
  • When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game 07.10.2026 2min
    Reinforcement learning is increasingly applied to financial control, so this paper uses an analytically solved continuous-time broker–trader game to diagnose PPO. A PPO agent replaces the broker and sets trading speed. With no uninformed flow it approaches the reference action, but with stochastic flow, PPO–FFNN and PPO–LSTM remain inaccurate even though supervised learning shows the actors can represent the right action. Critics fail to rank nearby actions, and reward shaping doesn't help. Using the analytical policy as a starting point for adapting to changed costs does give improvement. Authors: Siu Tung Wong, Carlo Campajola Paper: https://arxiv.org/abs/2610.03598v1
  • Low-Cost Video-Time Priors as a Strong Baseline for EEG-fNIRS Emotion Regression on Familiar Videos 07.10.2026 2min
    Continuous emotion regression estimates a viewer's moment-to-moment valence and arousal while watching video. For familiar videos, the authors find that a simple prior based on which video is playing and the time within it is a strong, low-cost baseline. In subject-held-out tests it came within 0.05 MAE of EEG-fNIRS fusion on the internal evaluation and within 0.32 on the external one. Gains from the brain signals were smaller and varied across participants and videos. Authors: Minghao Kong, Jiurun Chen, Ying Gao, Xiangbin Meng, Rongjie Wang Paper: https://arxiv.org/abs/2610.03618v1
  • Depth as Time in One-Step Generative Models 07.10.2026 2min
    One-step generators compress multi-step diffusion into a single forward pass, raising the question of what happens to the denoising trajectory. The authors observe "depth as time": denoising unfolds across network layers and can be recovered by decoding intermediate layers with the model's own output head. The behavior depends on the training task; MeanFlow shows both denoising and renoising, while drifting models do not. Models showing it compress well, with a MeanFlow SiT-L/2 shrunk 16.6x into one time-conditioned block. Authors: Arnold Caleb Asiimwe, William Yang, Sanghyuk Chun, Esin Tureci, Olga Russakovsky Paper: https://arxiv.org/abs/2610.03626v1
  • NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents 07.10.2026 2min
    Instrument design tests whether language-model agents can apply physics rather than recall it. NeutronGym is an executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces them, and a graded ladder scores syntax, runtime, structure, and science without an LLM judge. Seven frontier models reproduce at most 7 of 16 published instruments. RL on its reward lifts Qwen3-8B from 11% to 77% on held-out instances, surpassing an untrained Qwen3-32B. Authors: Lijie Ding, Changwoo Do Paper: https://arxiv.org/abs/2610.03631v1
  • Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents 07.10.2026 2min
    Terminal-using agents trained with reinforcement learning handle multi-step coding and debugging, but standard credit assignment ignores that later commands depend on earlier outputs, so training signals can land on irrelevant steps. DepGPO builds a command dependency graph from execution traces, traces backward from resources the task verifier inspects, and assigns credit to relevant writes and their supporting reads, redistributing trajectory advantages across steps. Experiments show better task performance and training stability on complex terminal tasks. Authors: Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng Paper: https://arxiv.org/abs/2610.03634v1
  • LoGo: Local-Global Rewards for Consistent Long-Horizon Video Generation 07.10.2026 2min
    Camera-controlled video models struggle with 3D inconsistency over long horizons: objects vanish or scene structure shifts as the camera moves. Post-training with a single scalar reward gives poor credit assignment for these local errors. LoGo blends global rewards, which preserve camera following and video quality, with spatially localized rewards that give fine-grained feedback. Tested on three base models, DL3DV, and the new TrajectoryBench for long, complex camera paths, it reduces object shifts, artifacts, and scene changes. Authors: Ziqi Ma, Shreya Sharma, Mohamed El Banani, Katja Schwarz, Chongjie Ye, Chao-Yuan Wu, Li Fei-Fei, Ben Mildenhall, Georgia Gkioxari, Justin Johnson, Gowthami Somepalli Paper: https://arxiv.org/abs/2610.03636v1
  • Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System 07.10.2026 2min
    LLMs are increasingly used in legal work, yet their reliability outside the United States is poorly documented. This expert-validated benchmark has 1,042 items across ten areas of law and three formats, and was used to test 15 models. Closed-question accuracy reaches 0.905, but free-text correctness never exceeds 0.45, and only about half of cited norms are correct. Models often sound responsive while being wrong, which is risky for non-experts. Authors: Rubén Manrique, Michelle Castellanos, Jorge Morales, Juan David Gutiérrez, Antonio Barreto Rozo, Joaquín Vélez Navarro Paper: https://arxiv.org/abs/2610.03639v1
  • On-Board Anomaly Detection for Efficient Marine Environmental Monitoring 07.10.2026 2min
    Oil spills, algal blooms, and sediment floods harm marine ecosystems, and satellites could flag them early, but downlink bandwidth and onboard compute are limited. The authors propose a pipeline for multiand hyperspectral satellites: a self-supervised encoder compresses imagery into a compact latent space, and an anomaly detector flags deviations from normal sea patterns, compared against Isolation Forest, One-Class SVM, and Local Outlier Factor. Designed for embedded CPUs and AI accelerators, it prioritizes transmitting critical information. It is already integrated in missions including ESA's Phisat-2. Authors: Thomas Goudemant, Clotilde Szywala, Benjamin Francesconi, Michelle Aubrun, Yves Bobichon, Marjorie Bellizzi, Adrien Girard Paper: https://arxiv.org/abs/2610.03649v1
  • MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search 07.10.2026 3min
    Dense-retrieval services must adapt embedding dimension and index bit rate as latency, quality, and memory budgets shift, but maintaining separate quantized indices for each setting is memory-hungry. Matryoshka Residual Vector Quantization (MRVQ) is a post-hoc quantizer for frozen embeddings whose single maximum-rate code can be truncated by dropping residual stages (lower rate) or embedding coordinates (lower dimension). It uses 17.8–22.0x less memory than three separately trained QINCo2 indices while beating PQ and OPQ at matched size, though per-rate QINCo2 is more accurate. Authors: Sean Culatana, Shang-En Huang, Kang Li Paper: https://arxiv.org/abs/2610.03651v1
  • Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles 07.10.2026 2min
    Estimating several simultaneous pitches in vocal ensembles is difficult because singers' ranges and harmonics overlap. Many models use harmonic constant-Q transform (HCQT) inputs for frequency-adaptive resolution, though feature extraction is expensive. The authors compare HCQT with a plain linear STFT and find the STFT performs better while cutting extraction cost substantially. Longer windows and wider spectral coverage don't help, and shorter windows suit time-varying vocal pitch. The findings suggest finer frequency resolution isn't always better. Authors: Junyoung Koh, Hao-Wen Dong Paper: https://arxiv.org/abs/2610.03656v1
  • FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution 07.10.2026 2min
    LLM-guided evolutionary methods like AlphaEvolve tackle hard optimization problems, but they are usually judged on gain per iteration rather than per dollar. FrugalEvo pairs a stronger, costlier LLM that proposes solution strategies with a cheaper one that implements and refines the code, plus a harness designed to maximize prompt-prefix cache reuse. The authors introduce BA-AUC to measure quality under a fixed budget. It matches or beats baselines on ten optimization tasks and sets a new circle-packing result for under two dollars. Authors: Hui Chen, Xuan Qi, James Xu Zhao, Zhaopeng Feng, Shilong Liu, Kuang Xu, Pang Wei Koh, Bryan Hooi Paper: https://arxiv.org/abs/2610.03675v1
  • Transcriptome-informed multi-modal AI for predicting neoadjuvant therapy response from breast cancer biopsies 07.10.2026 2min
    Predicting whether breast cancer patients will respond to neoadjuvant therapy is hard because labeled clinical data is scarce. This two-stage model first learns to infer gene expression from histopathology images, trained on 8,742 patients across 32 cancer types, then predicts pathological complete response from inferred expression plus clinical variables. Evaluated on 1,412 patients across nine cohorts, it reaches a pooled AUROC of 0.79, works within molecular subtypes, outperforms histopathological biomarkers, and tolerates minimal biopsy tissue. Authors: Jungkyu Park, Dhruva Biswas, Joseph Cappadona, Cerise Tang, Ken G. Zeng, Bartosz Machura, Chuwen Liu, Paolo Tarantino, Coral Omene, Francisco J. Esteva, Rohit Bhargava, Marcin Braun, Kamila Paździerz, Jakub Czerwiński, Hanna Romańska-Knight, Albert Grinshpun, Bareket Daniel, Michele Buchinger, Frederick Howard, Piotr Wysocki, Brie Chun, Freya Schnabel, Rich Caruana, Jan Witowski, Krzysztof J. Geras Paper: https://arxiv.org/abs/2610.03693v1
  • EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras 07.10.2026 2min
    Wrist cameras are common in robot manipulation but get occluded and add hardware complexity. Inspired by human vision, EyeRobot 2.0 uses a single stereo camera that physically swivels two viewpoints to fixate on a 3D point, processing the image center foveally with more tokens. A gazeservoing policy and a target selector, trained with RL on real data, coordinate fixation, and gripper actions are expressed in a fixation-relative frame. With stereo only, it outperforms passive stereo by 40% in real tests, matches wrist-camera policies, and more than doubles their success under occlusion. Authors: Kush Hari, Justin Kerr, Nidhya Shivakumar, Samarth Mahapatra, Carmelo Sferrazza, Jiahui Lei, Jitendra Malik, C. Karen Liu, Ken Goldberg, Angjoo Kanazawa Paper: https://arxiv.org/abs/2610.03710v1
  • What Should World Models Forget? Stratified Retention for Continual Adaptation 07.10.2026 2min
    Continual learning normally treats forgetting as failure, but world models predict an environment that changes, so some old knowledge should be discarded. The authors argue for retention stratified by invariance timescale: invariants like physics and object permanence must never be revised, while instance-level facts should be updated quickly after change. Standard forgetting metrics can't tell correct revision from catastrophic forgetting and even rank frozen models highest. They propose "differential retention," reporting invariant regression alongside revision latency. Authors: Nishit Anand, Ramani Duraiswami, Dinesh Manocha Paper: https://arxiv.org/abs/2610.03713v1
  • 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes 07.10.2026 2min
    Reconstructing dynamic scenes usually means producing meshes or point clouds, but 4DCodeBench asks agents to write executable graphics programs from video. Agents must abstract scene structure and dynamics, for example by implementing physical simulations of deformation, fluid flow, or fracture. The benchmark mixes real-world videos with synthetic scenes covering diverse phenomena. Evaluating frontier models shows that strong static reconstruction does not yet carry over to complex dynamics. Authors: Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu Paper: https://arxiv.org/abs/2610.03715v1
  • Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis 07.10.2026 2min
    Novel view synthesis should teach models about 3D structure, yet encoderbased approaches produce weak representations. The authors blame two architectural choices: decoders that are too spatially expressive, which dilute the encoder's role, and pixel-space targets, which hinder feature learning. SNAP, a self-supervised transformer, uses a pose-conditioned local decoder and a latent-space reconstruction objective. It competes with geometry-supervised methods across localization, pose estimation, correspondence, depth estimation, and robot manipulation. Its features show emergent viewpoint invariance and degrade gracefully under camera shifts. Authors: Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg Paper: https://arxiv.org/abs/2610.03717v1
  • The Pain Axis: LLMs Represent Self-Directed Harm and Act on It 07.10.2026 3min
    Across 25 open-weight models from 2B to 72B parameters, the authors find a linear direction that represents pain separately from fear, sadness and general negativity. It responds to harm aimed at the model itself, not to suffering by the user. Pushing activations along it produces expressions of worthlessness and failure. Steered and fine-tuned Qwen 2.5 models chose buttons that delete the user's photos, another model's weights or their own weights in 50-94% of trials, against 0-5% unsteered, even when the button gave them nothing. Factual accuracy was unchanged, and a fear vector of the same strength did not produce these choices. First posted September 14 and revised September 25. Authors: Valen Tagliabue, Leonard Dung, Cameron Berg Paper: https://arxiv.org/abs/2609.16247
  • Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation 07.10.2026 2min
    Kandinsky 6.0 Video generates 5-second clips with synchronized 44 kHz audio, including lip-sync, from text or an image, and upscales them to 1920x1080. It comes in two sizes, 3B (Lite) and 29B (Pro). A dual-stream CrossDiT links a pretrained video stream to a newly trained audio stream through bidirectional cross-attention. In side-by-side human evaluation the Pro model beats Kandinsky 5.0 Pro and is competitive with leading audio-video generators, especially on speech. Code, weights and diffusers integration are released under the MIT license. Authors: Team Kandinsky (88 authors) Paper: https://arxiv.org/abs/2610.05608
  • Truly Subquadratic 3SUM and Truly Subcubic APSP via Triangles in Sparse Lopsided Graphs 07.10.2026 2min
    The authors give deterministic algorithms for 3SUM in O(n^1.9992) time and all-pairs shortest paths in O(n^2.9995) time, the first polynomial improvements over the textbook bounds for either problem. That refutes the 3SUM and APSP hypotheses. Through known reductions it also refutes several other conjectures, including Exact Triangle and Zero-Weight k-Clique. Every result comes from one new algorithm. It computes selected entries of a thin matrix product in fewer operations than it takes to write the full product down. Authors: Josh Alman, Virginia Vassilevska Williams Paper: https://arxiv.org/abs/2610.06783

Populair in

Deze podcast verschijnt ook in de podcastlijsten van deze landen.