Paper Tracking Β· Archive

Every weekly run is preserved. Click a month to expand the weeks inside. Latest weeks at the top. ← Back to current week

October 2026 15 papers

Week of

Graph Γ— LLM

Vision-Language Models

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

↑ 27 πŸ“š 14588 β˜… 0 Sep 29
  • Problem: Unified multimodal models either quantize images to discrete tokens or use separate objectives for language and vision; a fully continuous shared generative process remains underexplored.
  • Model: "Multimodal Flow (MF-1): fully continuous flow-based generative model using ordered hyperchunks and chunk-causal Flow Matching backbone for joint language-vision modeling in embedding spaces.
  • Code: hustvl/Multimodal-Flow

World Models

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

↑ 77 πŸ“š 3981 β˜… 0 Oct 1
  • Problem: Existing video world models fail to maintain object states and dynamics after they leave the agent's view, limiting persistent world modeling.
  • Model: World Observer: decoupled actor-observer generation using a pretrained video Diffusion Transformer that jointly generates perspective actor and panoramic observer streams sharing a single world state.
  • Code: not released

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

↑ 19 πŸ“š 86589 β˜… 0 Sep 29
  • Problem: Long-horizon video generation loses coherence when early visual details fall outside the model's input window; existing selection methods rely on current content rather than future information needs.
  • Model: FrameMorrow: prospective frame selector predicting compact prospective tokens representing future information needs to guide selection of relevant historical frames across diverse generators.
  • Code: not released

Spatial Single-Cell Study

GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomics

↑ 0 πŸ“š 23935 β˜… 0 Sep 30
  • Problem: Spatial gene expression prediction from histopathology images relies on expensive tests; text-based optimization for incorporating gene semantic information remains underexplored.
  • Model: GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomicsβ€”combines H&E image patches with gene text summaries via cross-attention layers to predict spatial gene expression.
  • Code: not released

Autonomous Labs & Lab-in-the-Loop

Experimental Experience Modeling for Autonomous Research

↑ 0 πŸ“š 21750 β˜… 0 Sep 30
  • Problem: Autonomous research agents lack systematic methods to decide which experiments warrant running when evidence is insufficient.
  • Model: Experimental Experience Modeling (EEM): framework for acquiring, reusing, and accumulating experimental experience to guide experimental decisions through historical retrieval, sufficiency assessment, targeted pilots, and iterative library growth.
  • Code: not released

esQueranto: Differentiable Structured Quantum Light for Automated Scientific Discovery

↑ 0 πŸ“š 1470 β˜… 0 Sep 30
  • Problem: Existing quantum-optical simulators lack a unified framework combining photon-number quantum optics with structured-light propagation for automated experimental design.
  • Model: esQueranto: differentiable JAX-based simulator integrating Fock-state quantum optics, spatial beam geometry, and light propagation with automatic differentiation for optimization.
  • Code: artificial-scientist-lab/esQueranto

A roadmap for polymer informatics super-intelligence

↑ 0 πŸ“š 9 β˜… 0 Sep 28
  • Problem: Polymer informatics lacks integrated systems for inverse design, causal reasoning across chemistry-processing-performance, and autonomous closed-loop experimentation.
  • Model: approach: modular agent-directed architecture with conversational LLM layer coordinating domain-specialized tools for neat polymers, composites, solvents, synthesis, and processing across molecular-to-product design hierarchy
  • Code: not released
September 2026 31 papers

Week of

Vision-Language Models

InternW0-Ξ”: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

↑ 15 β˜… 0 Sep 24
  • Problem: Integrating large-scale pretrained visual, semantic, geometric, and motion priors into a unified framework for robot action generation.
  • Model: InternW0-Ξ”: World Action Model combining a pretrained video expert and action expert via Mixture-of-Transformers with Causal Imprint for action-relevant predictive representations, trained on 20K+ hours of heterogeneous robot and human demonstrations.
  • Code: InternRobotics/InternManip

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

↑ 6 πŸ“š 15073 β˜… 0 Sep 21
  • Problem: Reference-based IQA metrics rely on expensive, noisy human annotations; existing alternatives capture only relative comparisons rather than absolute perceptual distances.
  • Model: "FoMo (Forking Moment): automated data generation pipeline using diffusion model trajectories to assign pointwise perceptual distance labels between image pairs without human annotation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

World Models

Training Object Permanence in World Models

↑ 207 β˜… 0 Sep 22
  • Problem: Video generation models fail at object permanence and solidity, foundational aspects of physical reasoning in world models.
  • Model: PWM-WROP: 16B continuation world model fine-tuned on cognitive-science-inspired object permanence and solidity tasks.
  • Code: not released

The Past Frames the Future: Memory for Autoregressive Video Generation

↑ 49 πŸ“š 11507 β˜… 0 Sep 22
  • Problem: Autoregressive video generation loses critical historical information when context windows exceed bounded limits, breaking temporal consistency in long-horizon generation.
  • Model: approach: systematic survey of memory mechanisms in autoregressive video generation, organizing literature by representational forms, semantic functions, operational lifecycles, learning paradigms, and evaluation methods.
  • Code: HaroldChen19/Awesome-AR-Video-Memory

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

↑ 36 πŸ“š 8523 β˜… 0 Sep 23
  • Problem: Text-to-video generators require enhanced prompts with cinematic planning across shots, but existing prompt enhancers fail to preserve user requirements and align with video-grounded distributions.
  • Model: "WanPE": 397B-parameter prompt enhancement model using video-grounded reverse construction and Semantic-Consistency GRPO to generate shot-level cinematic plans for text-to-video generation.
  • Code: not released

Week of

Graph Γ— LLM

Vision-Language Models

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

↑ 114 πŸ“š 92730 β˜… 0 Sep 15
  • Problem: Evaluating whether multimodal alignment in omni-modal generative models improves physical world reasoning across complementary input modalities.
  • Model: approach: comprehensive evaluation framework for MiniMax-H3 across four physical reasoning dimensions using multimodal inputs (text, images, video, audio).
  • Code: gulucaptain/MiniMax-H3-Reason
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

↑ 33 πŸ“š 49288 β˜… 0 Sep 16
  • Problem: Evaluate whether multimodal language models can close the observe-reason-act-revise loop for embodied manipulation under incomplete observation without privileged geometry.
  • Model: approach: General-purpose MLLMs learning from RGB demonstrations to actively select camera viewpoints, issue metric Cartesian commands, and revise from execution feedback
  • Code: zhangzhongbo2213/VABench

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

↑ 41 πŸ“š 27336 β˜… 0 Sep 15
  • Problem: Action tokenizers for VLAs compress demonstrations lossy, distorting local action relationships needed for precise control despite low pointwise reconstruction error.
  • Model: "ActionPiece": action tokenizer with physical rank preservation and quantization regularization to maintain action relationship fidelity in autoregressive VLA policies.
  • Code: not released

World Models

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

↑ 10 πŸ“š 4313 β˜… 0 Sep 14
  • Problem: Existing controllable video generation methods lack simultaneous support for sparse, physics-grounded, interactive, and scene-level control over multi-object dynamics.
  • Model: PhysStream: autoregressive image-to-video model with structured scene memory (positional and object-tracking maps) and velocity-increment control conditioning, trained in two stages (bidirectional then causal).
  • Code: not released

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

↑ 34 β˜… 0 Sep 16
  • Problem: Video diffusion models suffer from computational bottlenecks in attention mechanisms during long spatiotemporal token processing.
  • Model: Video DeltaNet (VDN): hybrid attention combining local Softmax attention with bidirectional linear memory and Video Delta Attention for efficient video generation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Week of

Graph Γ— LLM

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

↑ 40 πŸ“š 7000 β˜… 0 Sep 7
  • Problem: LLM agents lose track of objectives and repeat unproductive actions due to unconstrained generation over flat action histories lacking explicit procedural structure.
  • Model: "Procedural Graph: a directed, attributed graph organizing procedural knowledge into (procedure, relation, procedure) triplets with online guidance generation and offline self-evolution via LLM refiner.
  • Code: not released

DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat

↑ 22 β˜… 0 Sep 9
  • Problem: Multi-agent air combat requires structured relational modeling and explicit tactical role assignment, which flat MARL architectures lack.
  • Model: DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Proximal Policy Optimization, combining graph attention networks for battlefield relational modeling with dynamic role assignment (leader/supporter) and low-level maneuver execution.
  • Code: not released

Vision-Language Models

SenseNova-U1.5: Towards Native Unified Visual Intelligence

↑ 245 πŸ“š 46414 β˜… 0 Sep 9
  • Problem: Existing multimodal models lack unified visual understanding, reasoning, and generation capabilities within a single end-to-end architecture.
  • Model: SenseNova-U1.5: 8B-MoE native unified multimodal model with encoder-free and VAE-free architecture for visual understanding, reasoning, and generation at native resolutions up to 4K.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Memory as Plans: World-Action Modeling with Memory-Grounded Planning

↑ 36 πŸ“š 9812 β˜… 0 Sep 9
  • Problem: Robotic policies struggle with non-Markovian manipulation tasks requiring long-horizon memory while maintaining computational efficiency during execution.
  • Model: MaP-WAM: Memory-as-Plans framework decomposing memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution using a World-Action-Progress (WAP) model.
  • Code: not released

NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting

↑ 4 πŸ“š 96307 β˜… 0 Sep 7
  • Problem: Existing AI models struggle to capture complex irregular temporal dynamics and stochasticity of multimodal longitudinal patient records for representation and forecasting.
  • Model: NOAH: time-aware, task-agnostic, generative transformer model featuring bidirectional time integration and variational latent space for multimodal patient trajectory modeling and forecasting
  • Code: not released

World Models

Programmable World Model

↑ 101 πŸ“š 17364 β˜… 0 Sep 8
  • Problem: Video world models lack persistent state management and programmable rule enforcement for extended interactive experiences.
  • Model: "Programmable World Model: framework decoupling world-state evolution from visual generation via state-augmented 3D OBBs and a deterministic state compiler"
  • Code: AlayaLab/pwm

World in World: Explore the World with World Models

↑ 28 πŸ“š 27852 β˜… 0 Sep 9
  • Problem: Flexible control of frozen video world models for arbitrary viewpoint exploration while maintaining temporal consistency and plausible completion of unobserved regions.
  • Model: "World in World": training-free inference-time interface using correspondence-guided attention routing and evidence-wise attention CFG to convert heterogeneous visual evidence into controllable camera-conditioned video generation from a frozen causal video model.
  • Code: not released

Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs

↑ 28 πŸ“š 13372 β˜… 0 Sep 9
  • Problem: Reconstructing complex 3D worlds as executable code from single reference images requires a systematic construction approach.
  • Model: Recursive Code World Models (RCWM): framework coupling Recursive Scene Programs with a construction solver that recursively decomposes and refines scene structure through global-local-global perception loops.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Week of

Graph Γ— LLM

Evidence, Logic, and Compliance: Multi-Agent Structured Graph Reasoning with Expert Arbitration for Medical Referral

↑ 0 πŸ“š 49974 β˜… 0 Aug 31
  • Problem: LLMs struggle with medical referral decisions due to information overload and unstructured multi-agent collaboration causing semantic drift and confirmation bias.
  • Model: MASGR (Multi-Agent Structured Graph Reasoning): specialized agents extract evidence from distinct modalities and coordinate via explicit clinical reasoning graphs with expert knowledge arbitration.
  • Code: not released

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

↑ 0 β˜… 0 Sep 3
  • Problem: Medical concept encoders learn static representations across patients, but clinical meaning and predictive value depend on patient-specific context and trajectory.
  • Model: REFINE: KG-aware budgeted LLM graph refinement framework using sequential RL for adaptive KG expansion, heterogeneous GNN for structure, and frozen LLM with graph-aware soft prompts for semantic refinement.
  • Code: not released

Vision-Language Models

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

↑ 227 πŸ“š 13872 β˜… 0 Sep 2
  • Problem: Existing image generators rely heavily on paired image-text data; this work develops a fully open-source framework with transparent training recipes.
  • Model: LLaDA-Image: 6B Diffusion Transformer paired with frozen vision-language module from LLaDA2.0-Mini, trained via image-only pre-training followed by instruction-tuning.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

↑ 66 πŸ“š 85464 β˜… 0 Sep 3
  • Problem: Existing approaches lack unified multimodal models that jointly handle physics, geometry, and appearance for 3D world understanding and generation.
  • Model: Puffin-World: unified multimodal architecture integrating physical understanding, spatial simulation, and 3D world generation with native physics, geometry, and appearance states
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

↑ 29 πŸ“š 94571 β˜… 0 Sep 3
  • Problem: Streaming video understanding with MLLMs requires internalizing historical evidence into compact, evolving latent memory rather than treating it as external retrievable context.
  • Model: "LatentStream: progressive latent working memory framework shifting from store-and-retrieve to retrieve-and-internalize paradigm for streaming video understanding"
  • Code: not released

World Models

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

↑ 141 πŸ“š 13374 β˜… 0 Sep 2
  • Problem: Training scalable video world models requires handling heterogeneous datasets and video backbones with inconsistent supervision and difficult reproducibility.
  • Model: "SolarWM: reconfigurable multi-source data engine and backbone-native adaptation framework for interactive video world models"
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

↑ 33 πŸ“š 10312 β˜… 0 Sep 1
  • Problem: Joint audio-video generators synchronize video and audio but ignore script-specified timing, causing shot transitions and dialogue to occur at wrong times.
  • Model: "Temporal Context Routing" (TCR): routes script timing onto shared video-audio temporal axis via duration-normalized routing scores added to cross-attention logits.
  • Code: DAGroup-PKU/Temporal-Context-Routing

Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

↑ 0 πŸ“š 94571 β˜… 0 Sep 3
  • Problem: Existing world models for autonomous driving either separate future prediction from action generation or jointly predict them at the same temporal scale, limiting long-horizon anticipation and responsive decision-making.
  • Model: "Drive-HWM: Hierarchical slow-fast world modeling framework with dynamic-aware latents learned through optical-flow prediction for multi-scale autonomous driving.
  • Code: not released

Spatial Single-Cell Study

August 2026 55 papers

Week of

Graph Γ— LLM

CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

↑ 5 πŸ“š 12409 β˜… 0 Aug 25
  • Problem: LLM agents struggle to retrieve relevant skill subsets from large libraries without overwhelming context or losing workflow dependencies.
  • Model: CaSKG: counterfactual-causal skill graph framework that calibrates procedural relations using multi-signal candidate induction, textual counterfactual probes, and Bayesian state-gated publication for task-conditioned retrieval.
  • Code: ZhiyuanLi218/Caskg

Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning

↑ 0 πŸ“š 34060 β˜… 0 Aug 27
  • Problem: Message passing in GNNs is computationally expensive and sensitive to graph structure; unclear why neighborhood aggregation outperforms MLPs.
  • Model: "RTA" (Retrieve-Then-Aggregate): retrieval-augmented framework replacing structural message passing with label-aware retrieval and MLP-based aggregation on text-attributed graphs.
  • Code: not released

Are LLM-Enhanced GNNs Privacy-Safe?

↑ 0 πŸ“š 4963 β˜… 0 Aug 26
  • Problem: Privacy vulnerabilities in LLM-enhanced GNNs remain unexplored despite their widespread adoption for text-attributed graph learning.
  • Model: approach: systematic privacy risk evaluation framework covering link, label, and membership inference attacks on LLM-enhanced GNNs with differential privacy defenses
  • Code: not released

Vision-Language Models

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

↑ 255 πŸ“š 2391 β˜… 0 Aug 25
  • Problem: Native visual reasoning lacks scalable training tasks, reliable evaluation methods, and controlled comparisons across generative substrates.
  • Model: "VBVR-Pro: A closed-loop testbed for native visual reasoning through generation with 300 procedurally generated tasks, verifiable reward scorers, and support for image, video, and interleaved generators."
  • Code: not released

VGI-Bench: Probing Visual Intelligence in Video Generation Models

↑ 174 πŸ“š 30925 β˜… 0 Aug 25
  • Problem: Existing benchmarks for evaluating visual reasoning in video generation models use mismatched inputs, don't require valid intermediate trajectories, and lack calibrated difficulty levels.
  • Model: approach: VGI-Bench, a benchmark with 27 tasks across 810 instances organized by task domains and skill tags to evaluate visual reasoning in video generation models
  • Code: not released

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

↑ 105 πŸ“š 4084 β˜… 0 Aug 27
  • Problem: Evaluating whether multimodal large language models can reliably navigate and act in real-scale urban environments beyond local perception tasks.
  • Model: approach: UrbanGroundβ€”a 3D city sandbox built from geospatial data of Hong Kong to test MLLM agent navigation and spatial reasoning in closed-loop first-person interaction.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

World Models

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

↑ 182 πŸ“š 4181 β˜… 0 Aug 26
  • Problem: World model scaling relies on fuzzy reward signals like CLIP scores, lacking the executable verification that enables RL post-training for code agents.
  • Model: Agentic World Model (AWoMo): a world-building agent that proposes scene edits, observes human-engine verification signals, and learns from development trajectories using Reinforcement Learning with Human-Engine Verification (RLHEV).
  • Code: LanceZPF/cardinal-preview

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

↑ 137 πŸ“š 60168 β˜… 0 Aug 27
  • Problem: Existing video generation models lack evaluation of whether they reproduce correct distributions of possible physical behaviors, not just individual plausibility.
  • Model: approach: PAWBench benchmark with PAWEval protocol for evaluating video generators as stochastic samplers of world dynamics via distributional alignment
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Code World Model: Coding Agent as World Brain

↑ 32 πŸ“š 27731 β˜… 0 Aug 26
  • Problem: Existing video-based world models learn only visual outcomes, not the underlying rules and mechanisms needed for persistent, coherent open-ended world evolution.
  • Model: Code World Model: framework combining a coding agent (world brain) that maintains executable state through language model reasoning and generated code, with a video model for high-fidelity visual rendering via proxy representation.
  • Code: not released

Spatial Single-Cell Study

Learning Interpretable Tumor Microenvironment Representations by Fitting Pan-Cancer Cell State-Niche Correlation

↑ 0 πŸ“š 84 β˜… 0 Aug 26
  • Problem: Existing spatial transcriptomics models cannot jointly learn cell state-niche correlations and infer underlying ligand-receptor mechanisms with interpretability.
  • Model: GITIII-scale: hierarchical interpretable pan-cancer foundation model using graph transformers to decompose cell state changes by sender cell and gene to reveal ligand-receptor signaling.
  • Code: lugia-xiao/GITIII_scale

Week of

Graph Γ— LLM

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

↑ 13 β˜… 0 Aug 20
  • Problem: Single LLM agents cannot handle complex tasks requiring heterogeneous expertise, parallel execution, and persistent state across interdependent subtasks.
  • Model: approach: Graph Engineeringβ€”a paradigm using explicit, dynamic graph structures to organize tasks, coordinate heterogeneous agents, and manage runtime state for multi-agent system intelligence.
  • Code: DEEP-JLU/Awesome-Graph-Engineering

Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

↑ 0 πŸ“š 42977 β˜… 0 Aug 18
  • Problem: LLMs suffer from reasoning evidence perception drift when reasoning over knowledge graphs, selecting semantically plausible but structurally unsound evidence.
  • Model: Structure-Internalized Rule Language Model (SIRLM): LLM framework coupling structural rule generation with parametric learning via a Structure-Internalized Rule Generator, KG tokenizer, and neuro-symbolic reasoner.
  • Code: lazyloafer/SIRLM

Beyond Similarity Matching: Structured Reasoning for Open-Vocabulary Referring Segmentation in 3DGS

↑ 0 πŸ“š 2671 β˜… 0 Aug 17
  • Problem: Existing 3D Gaussian Splatting methods rely on global text similarity and fail on queries with attributes, spatial relations, and fine-grained parts.
  • Model: QAGaussian: query-adaptive neural reasoning framework using query-conditioned multi-scale Gaussian slots, relation-aware slot graphs, and granularity-adaptive routing for open-vocabulary referring segmentation in 3DGS.
  • Code: zqeslwyz/QAGaussian

Vision-Language Models

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

↑ 51 πŸ“š 2112 β˜… 0 Aug 17
  • Problem: VLMs struggle with embodied navigation due to misalignment with 2D pre-training, inefficient reasoning schedules, and poor memory management.
  • Model: TAMP-Nav: unified framework with Pixel-to-3D action formulation, selective reasoning and anchor-trajectory memory, and two-level Group Relative Policy Optimization (GRPO) alignment.
  • Code: not released

EXIMO: VLM Guided Exploration of VLA Policies

↑ 13 πŸ“š 56943 β˜… 0 Aug 19
  • Problem: Efficiently finetune vision-language-action robot policies for new long-horizon tasks without expensive teleoperation data.
  • Model: EXIMO: three-stage algorithm (explore via VLM orchestration, imitate via supervised finetuning, optimize via residual RL) for sample-efficient VLA policy adaptation
  • Code: not released

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

↑ 22 πŸ“š 2544 β˜… 0 Aug 17
  • Problem: Existing diffusion models cannot efficiently edit ultra-high-resolution images due to quadratic attention complexity and memory constraints, and two-stage pipelines suffer from information divergence and texture degradation.
  • Model: EditBridge: diffusion bridge framework using prior-guided block-wise sparse attention for efficient data-to-data translation from low-resolution edited results to high-resolution outputs conditioned on original HR source.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

World Models

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

↑ 155 πŸ“š 36972 β˜… 0 Aug 17
  • Problem: Video generation models lack evaluation of outcome achievement while preserving semantic grounding between reference images and generated results.
  • Model: approach: SemComp-Bench, a VLM-based evaluation protocol measuring Outcome Achievement and Generation Reliability through structured binary questions on outcome-centric video generation.
  • Code: not released

Spatial Single-Cell Study

scDNM-VAE enables directly inspectable deep clustering of single-cell RNA-seq data through signed dendritic gating

↑ 0 πŸ“š 5233 β˜… 0 Aug 18
  • Problem: Deep clustering models for scRNA-seq assign cells through opaque latent or centroid mechanisms, limiting interpretability of clustering decisions.
  • Model: scDNM-VAE: variational autoencoder with dendritic neuron-inspired clustering head using signed synaptic weights and thresholds for directly inspectable cell assignment.
  • Code: not released

Week of

Graph Γ— LLM

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

↑ 0 πŸ“š 45123 β˜… 0 Aug 11
  • Problem: Object-centric question answering in long egocentric video requires preserving persistent object identity and structured spatial change, which caption-based methods fail to capture.
  • Model: R4DSG: Relative 4D scene graph memory that indexes video by time, place, persistent objects, and anchor-relative transitions without requiring global world coordinates.
  • Code: https://dualtransparency.github.io/R4DSG/

Vision-Language Models

Intern-S2-Preview: Scientific Agentic Foundation Model

↑ 56 β˜… 0 Aug 12
  • Problem: AI systems need to reason over heterogeneous scientific modalities, interact with tools, and sustain progress over long task horizons for scientific discovery.
  • Model: "Intern-S2-Preview-397B": multimodal scientific foundation model with time series forecasting, agentic RL training, and modular Memory Decoder for domain specialization.
  • Code: InternLM/xtuner

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

↑ 26 πŸ“š 9561 β˜… 0 Aug 10
  • Problem: Static pretrained feature spaces used in FrΓ©chet distance losses enable FrΓ©chet hacking, where optimized metrics improve while visual quality degrades.
  • Model: "AdvFD: Adversarial FrΓ©chet Distance" β€” combines static pretrained encoders with an adversarially learned adaptive representation and real-feature whitening to expose residual distribution mismatches during generator post-training.
  • Code: not released

World Models

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

↑ 119 πŸ“š 14838 β˜… 0 Aug 13
  • Problem: Interactive world models struggle to balance persistent memory, low-latency response, and long-horizon generation due to conflicting computational demands.
  • Model: "Evoke": world model with externalized scene geometry bank and linear-attention teacher for long-horizon supervision enabling few-step student generation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

↑ 93 πŸ“š 16421 β˜… 0 Aug 13
  • Problem: Action-conditioned video world models struggle to faithfully predict robot manipulation while preserving arm identity, object state, and scene geometry.
  • Model: DreamX-Phi 1.0: action-conditioned video world model using PRoPE-style SE(3) geometric encoding, depth supervision, SAM3 masks, and V-JEPA teacher for faithful robotic manipulation prediction.
  • Code: AMAP-ML/DreamX-Phi

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

↑ 43 πŸ“š 37560 β˜… 0 Aug 13
  • Problem: Existing world model benchmarks use fixed action sequences that fail to enable fair comparison across models with different action granularities and response dynamics.
  • Model: PlayWorld: benchmark with multi-modal Agent Player that adaptively executes long-horizon objectives to evaluate world models on geometry consistency, interaction fidelity, and state evolution.
  • Code: kxding/PlayWorld

Spatial Single-Cell Study

Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

↑ 0 β˜… 0 Aug 14
  • Problem: Standard multimodal models for spatial transcriptomics compress H&E and gene expression into shared latent space without explicitly separating shared versus modality-specific variation.
  • Model: approach: Disentangled multimodal representation learning comparing VAE-based (MMVAE+, MMVAE+sg) and contrastive (disSSL) models that factorize shared and private latent components for H&E and spatial transcriptomics integration.
  • Code: not released

Program-space Diffusion for Morphology-to-Transcriptomics Prediction

↑ 0 β˜… 0 Aug 14
  • Problem: Predicting spatial gene expression from histology is computationally expensive and ignores coordinated transcriptional variation.
  • Model: approach: Conditional diffusion model operating in transcriptional program space learned via consensus non-negative matrix factorization (cNMF), with Pearson residual normalization for count data.
  • Code: not released

Week of

Graph Γ— LLM

NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning

↑ 0 πŸ“š 44134 β˜… 0 Aug 5
  • Problem: Designing JEPA-style latent prediction objectives for node-level graph self-supervised learning and determining which structural signals the predictor should condition on.
  • Model: NodeJEPA: joint-embedding predictive architecture that masks structure-aware k-hop ego-subgraphs and predicts latent node representations using a structure-conditioned predictor with spectral and centrality descriptors.
  • Code: OliverZ-dot/Node-Jepa

EvtGraph: Event-Adaptive Compression for Sparse Temporal Graph Learning in Multimodal Time Series

↑ 0 πŸ“š 4351 β˜… 0 Aug 5
  • Problem: Multimodal temporal data have non-uniform information density, but standard models allocate computation uniformly, causing inefficiency and node explosion in graph representations.
  • Model: "EvtGraph: Event-adaptive compression framework with node budget control and temporally constrained sparse graph reasoning for multimodal time series.
  • Code: not released

Vision-Language Models

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

↑ 55 πŸ“š 93935 β˜… 0 Aug 4
  • Problem: Text-to-image models lack unified agentic control over reasoning, tool use, and generation for complex open-world image synthesis tasks.
  • Model: ToolArtist: unified multimodal model trained with supervised fine-tuning and agentic RL (RAD-GRPO) to dynamically orchestrate reasoning, external tools, and image generation within a single policy
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

↑ 57 πŸ“š 30079 β˜… 0 Aug 4
  • Problem: Fundamental mechanisms governing how modalities interact during unified multimodal pretraining remain underexplored.
  • Model: approach: systematic empirical exploration of unified multimodal pretraining (language, visual understanding, visual generation) using controlled experiments to derive design principles and efficient recipes.
  • Code: not released

ChronoVision: Temporal Reasoning via Latent State Reconstruction

↑ 38 πŸ“š 21747 β˜… 0 Aug 5
  • Problem: Multimodal LLMs struggle with temporal reasoning tasks because language-based explanations fail to capture continuous visual transformations.
  • Model: ChronoVision: multimodal framework with Reconstructive Visual Head predicting latent final states, ROI Attention Locating module, and RL-based implicit process grounding.
  • Code: not released

World Models

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

↑ 38 πŸ“š 5338 β˜… 0 Aug 6
  • Problem: Training LLM agents for long-horizon tool use requires costly external environments or unreliable simulators; how can agents internalize environment dynamics without external interaction?
  • Model: EnvACE: agentic RL method where a policy alternates between acting (generating tool calls) and world rehearsal (playing the environment to produce action responses), jointly optimized end-to-end with task-success rewards.
  • Code: Within-yao/EnvACE

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

↑ 32 πŸ“š 6714 β˜… 0 Aug 6
  • Problem: Existing economic simulations lack generative mechanisms where heterogeneous agents interact to produce emergent aggregate outcomes with persistent empirical alignment.
  • Model: approach: Economic World Models (EWM) β€” a modular runtime system organizing heterogeneous agents, environment mechanics, co-evolution, and real-world alignment into a closed-loop generative engine for economic simulation.
  • Code: FreedomIntelligence/Awesome-Economic-World-Models

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

↑ 12 πŸ“š 31569 β˜… 0 Aug 5
  • Problem: Interactive video world models suffer from compounding errors in long-horizon rollouts with no ground-truth supervision to measure accumulated drift.
  • Model: "WorldCycle: Self-Verifiable Reinforcement Learning framework using reversible action cycles with spatial closure and temporal consistency rewards for long-horizon video world models"
  • Code: not released

Spatial Single-Cell Study

Control-Anchored Residual Flow Matching Conditioned on Gene Geometry for Virtual Cell Perturbation Modeling

↑ 0 β˜… 0 Aug 7
  • Problem: Existing graph-based virtual cell models conflate stable gene associations with intervention-specific response pathways, lacking separation between structural priors and learned dynamics.
  • Model: GeneGeoFlow: control-anchored residual flow conditioned on perturbation-specific gene geometry derived from multi-scale spectral coordinates of Gene Ontology and coexpression networks.
  • Code: not released

Week of

Graph Γ— LLM

Harness-G: A Graph-Structured Harness for Search Agents

↑ 3 πŸ“š 14432 β˜… 0 Jul 29
  • Problem: RL search agents suffer from retrieval-equivalence collapse where distinct queries yield overlapping evidence, preventing effective credit assignment.
  • Model: Harness-G: graph-structured retrieval framework that reformulates free-form query generation as finite action selection over a paragraph-sentence-entity graph, with Structured Non-myopic Credit (SNC) for credit assignment.
  • Code: 7HHHHH/Harness-G

GraphQAG: A Knowledge-Graph-Guided Visual Analytics Framework for Question-Answer Pairs Generation

↑ 0 πŸ“š 25208 β˜… 0 Jul 29
  • Problem: Existing QA generation methods fail to capture fragmented knowledge distributed across paragraphs and multi-entity relationships in long documents.
  • Model: "GraphQAG: knowledge-graph-guided visual analytics framework combining entity/relation extraction, graph-based generation space, and interactive visualization for LLM-based QA pair generation from documents.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Vision-Language Models

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

↑ 130 πŸ“š 14078 β˜… 0 Jul 28
  • Problem: LLM-centric vision-language-action models incur high computational and memory overhead, limiting real-time robotic control and deployment on resource-constrained platforms.
  • Model: TurboVLA: direct V+Lβ†’A vision-language-action model using independent vision and text encoders with lightweight cross-attention interaction and continuous action chunk prediction.
  • Code: H-EmbodVis/TurboVLA

HumanCLAW: Can Vision-Language Models Act Through a Body?

↑ 71 πŸ“š 5852 β˜… 0 Jul 28
  • Problem: Vision-language models lack embodied self-awareness to make moment-to-moment action decisions through a physical humanoid body in closed-loop control.
  • Model: approach: HumanCLAW framework decoupling VLM decision-making from motor execution via atomic skill commands translated to full-body motion with physics simulation
  • Code: not released

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

↑ 37 πŸ“š 6210 β˜… 0 Jul 29
  • Problem: Existing text-to-image editors fail at anatomically plausible multi-person interaction synthesis, producing fused limbs and interpenetrating bodies that VLM judges fail to detect.
  • Model: "MPIE-Bench and MPIE-Eval: benchmark and mesh-based evaluation protocol for multi-person interaction editing using 3D human mesh reconstruction to score anatomical completeness and contact geometry.
  • Code: AnnLin0628/mpie-bench

World Models

PhiZero: A World Model Built Around Physical Language

↑ 160 πŸ“š 11539 β˜… 0 Jul 30
  • Problem: Existing video world models predict pixels directly, leaving physical dynamics implicit; how can we learn explicit representations of state transitions?
  • Model: PhiZero: world model using discrete physical language tokenizer and reasoner to infer state transitions then render videos
  • Code: not released

QQWorld: Quantile-Quantile Matching for World Model Regularization

↑ 23 πŸ“š 4488 β˜… 0 Jul 30
  • Problem: Latent world models regularized with Epps-Pulley objective suffer from heavy-tailed latent distributions due to vanishing corrective gradients for tail samples.
  • Model: QQWorld: quantile-quantile matching regularizer for latent world models, replacing Epps-Pulley with rank-matched Gaussian quantile alignment
  • Code: not released

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

↑ 17 πŸ“š 4839 β˜… 0 Jul 30
  • Problem: Existing video world models lack unified, frame-level action control that transfers across diverse dynamics and appearances.
  • Model: "ShadowDancer": interactive video world model learning unified dynamics representations from shadow pairs (video pairs with identical motion but resampled appearance) via cross-shadow prediction.
  • Code: not released

Spatial Single-Cell Study

BayesClint: Bayesian Multi-Scale Clustering and Multi-Sample Integration With Feature Selection for Spatial Transcriptomics Data

↑ 0 πŸ“š 439 β˜… 0 Jul 27
  • Problem: Existing spatial transcriptomics clustering methods perform dimension reduction and clustering separately and cluster at only a single scale, rather than jointly identifying cell types and spatial domains.
  • Model: "BayesClint": Bayesian hierarchical model for joint factor analysis, multi-scale clustering, multi-sample integration, and nested feature selection in spatial transcriptomics.
  • Code: AlvinSheng/BayesClint
July 2026 42 papers

Week of

Graph Γ— LLM

HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology

↑ 0 πŸ“š 8811 β˜… 0 Jul 23
  • Problem: Predicting spatially resolved gene expression from H&E histology while quantifying prediction confidence and accounting for tissue heterogeneity.
  • Model: HierarchicalDAEW: dual-graph architecture with domain-aware edge-weighted graph convolution for spot-level predictions and gene-level graph decoder fusing protein-protein interactions with co-expression, plus evidential uncertainty estimation via Normal-Inverse-Gamma loss.
  • Code: not released

Attacking Graph Foundation Models Through Their Shared Representation

↑ 0 πŸ“š 364 β˜… 0 Jul 20
  • Problem: Graph foundation models share a representation bottleneck (alignment layer) across domains, creating a novel attack surface not studied in prior adversarial graph work.
  • Model: approach: representation-space and input-space adversarial attacks targeting alignment layers in six graph foundation models (spectral tokenizers, text embeddings, discrete codebooks) at inference time.
  • Code: not released

Vision-Language Models

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

↑ 50 πŸ“š 34877 β˜… 0 Jul 21
  • Problem: Mobile agents struggle to follow natural language-described targets in crowded scenes because vision-language-action models reason in abstract spatial latents misaligned with explicit image detections.
  • Model: ReferTrack: VLA model that selects targets from indexed bounding boxes then decodes tracking waypoints, using temporal-viewpoint-bbox indicator tokens to preserve motion cues.
  • Code: MedlarTea/referTrack

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

↑ 71 πŸ“š 1033 β˜… 0 Jul 20
  • Problem: Large-scale visual generators are expensive to train, fine-tune, and deploy; efficient compact alternatives are needed.
  • Model: Mage-Flow: 4B-scale generative stack combining Mage-VAE lightweight tokenizer with Native-Resolution Multimodal Diffusion Transformer for text-to-image generation and instruction-based editing.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Visual Contrastive Self-Distillation

↑ 44 πŸ“š 2274 β˜… 0 Jul 22
  • Problem: On-policy self-distillation requires asymmetric information between teacher and student, typically through privileged answers or visual evidence signals.
  • Model: Visual Contrastive Self-Distillation (VCSD): contrasts EMA teacher predictions on original and content-erased images to sharpen token distributions for on-policy self-distillation.
  • Code: not released

World Models

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

↑ 33 πŸ“š 18895 β˜… 0 Jul 22
  • Problem: Existing video diffusion transformers require quadratic attention complexity, limiting efficiency for long-sequence high-resolution video generation on single GPUs.
  • Model: SANA-Video 2.0: hybrid linear-softmax attention video diffusion transformer with block attention residuals for efficient 720p video generation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

↑ 56 πŸ“š 212 β˜… 0 Jul 20
  • Problem: Generating interactive, spatiotemporally consistent, long-horizon video worlds from text, images, or video with low latency response.
  • Model: AlayaWorld: 15B video diffusion transformer generating 24-fps video with bounded visual context (sink frame, temporal history, spatial memory) and discrete autoregressive distillation for efficient inference.
  • Code: AlayaLab/AlayaWorld

GraphVid: Interactive Graph-Controllable Video Generation

↑ 4 πŸ“š 1056 β˜… 0 Jul 22
  • Problem: Controllable video generation lacks precise multi-object interaction control; trajectory-based methods scale poorly and require dense annotations.
  • Model: GraphVid: graph-conditioned image-to-video generation model using directed interaction scene graphs with edge-aware graph reasoning and LoRA adaptation to a frozen video diffusion transformer.
  • Code: not released

Spatial Single-Cell Study

Auditing pretraining contamination in single-cell foundation model benchmarks

↑ 0 πŸ“š 593 β˜… 0 Jul 21
  • Problem: Single-cell foundation model benchmarks may contain pretraining data, conflating memorization with genuine generalization in zero-shot evaluations.
  • Model: "scContam": per-cell audit framework combining MinHash gene-set fingerprints and loss-based membership inference attack (MIA-scFM) to detect pretraining contamination.
  • Code: not released

Loom: Multi-Region Analysis of Spatial Transcriptomics with Local Neighborhoods and Global Trajectories

↑ 0 β˜… 0 Jul 24
  • Problem: Existing spatial transcriptomics platforms lack integrated support for pseudo-temporal trajectory analysis, multi-region comparison, and local microenvironment investigation.
  • Model: "Loom": visual computing system integrating spatial coordinate registration, dimensionality reduction, clustering, pseudotime simulation, and novel glyph-based encoding for spatiotemporal ST analysis
  • Code: ScheWann/Loom

Week of

Graph Γ— LLM

Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment

↑ 0 πŸ“š 13352 β˜… 0 Jul 13
  • Problem: Existing graph foundation models rely heavily on textual attributes or limited training data; enabling tabular foundation models to capture graph structural information remains unexplored.
  • Model: GTAlign: Graph-to-Table Alignment framework combining a universal graph encoder with community-guided continual pre-training and tabular foundation models for text-free graph representation learning.
  • Code: not released

Vision-Language Models

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

↑ 151 πŸ“š 27844 β˜… 0 Jul 15
  • Problem: Existing video MLLMs lack generalization across diverse video types, suffer from high computational demands, and remain partially closed-source, hindering reproducibility.
  • Model: "VideoChat3": fully open video MLLM with Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for efficient spatiotemporal encoding across general, long-form, and streaming video understanding.
  • Code: MCG-NJU/VideoChat3

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

↑ 127 πŸ“š 922 β˜… 0 Jul 13
  • Problem: Closed-source multimodal systems achieve strong text-to-image generation and editing performance, but their methods remain undisclosed; demonstrating comparable results in open-source models under constrained budgets.
  • Model: Boogu-Image-0.1: unified multimodal understanding and generation model family with Base, Turbo, Edit, and Edit-Turbo variants for text-to-image generation, fast inference, and instruction-based editing.
  • Code: Boogu-Project/Boogu-Image
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

↑ 24 πŸ“š 29151 β˜… 0 Jul 15
  • Problem: Existing embodied AI models separate textual reasoning from visual goal imagination; a unified approach coupling language planning with visual state prediction is needed.
  • Model: Hy-Embodied-RxBrain: multimodal Mixture-of-Transformers foundation model for embodied cognition integrating joint language-visual reasoning and imagination in a single planning sequence.
  • Code: Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0

World Models

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

↑ 37 πŸ“š 7690 β˜… 0 Jul 14
  • Problem: Existing keyframe-conditioned video generation models lack comprehensive evaluation frameworks to measure both keyframe fidelity and overall video quality.
  • Model: "KeyFrame-Compass: benchmark and automated evaluation framework for keyframe-conditioned video generation combining six keyframe execution metrics and MLLM-based quality assessment"
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

↑ 32 πŸ“š 4837 β˜… 0 Jul 14
  • Problem: Existing benchmarks lack comprehensive evaluation of multi-reference-to-audio-video generation, which requires jointly reasoning over multiple references while maintaining synchronized visual and audio content.
  • Model: approach: MultiRef-Compassβ€”a benchmark with 350 curated samples, asset-composition pipeline, and hybrid evaluation framework integrating automatic metrics with MLLM-as-a-Judge across four dimensions (Basic Quality, Reference Consistency, Audio-Visual Consistency, Instruction Following).
  • Code: not released

From Pixels to States: Rethinking Interactive World Models as Game Engines

↑ 31 πŸ“š 16 β˜… 0 Jul 15
  • Problem: Interactive world models lack the explicit state representation and rule-governed dynamics of conventional game engines, limiting coherence and real-time interactivity.
  • Model: approach: Unified framework organizing interactive game world modeling along four dimensions (player action control, game state dynamics, state-observation persistence, real-time generation) grounded in the action-state-observation loop.
  • Code: not released

Spatial Single-Cell Study

A vision foundation model for single-cell biology via spatial gene cartography

↑ 0 πŸ“š 1073 β˜… 0 Jul 15
  • Problem: Single-cell foundation models based on language models discard gene relationships and expression magnitude by tokenizing transcriptomes.
  • Model: scVision: vision transformer foundation model that renders cells as continuous gene-expression images using optimal transport for spatial gene placement, pretrained via masked image modeling on 72 million human cells.
  • Code: not released

Week of

Graph Γ— LLM

Canopy: A Heterograph Foundation Model for Metabolic Engineering

↑ 0 β˜… 0 Jul 7
  • Problem: Existing computational approaches for metabolic engineering either ignore experimental data or discard relational structure of biological knowledge.
  • Model: Canopy: heterogeneous graph foundation model integrating 6.9M nodes across 13 types with multi-modal encodings (ESM-2, MoLFormer, PubMedBERT) and pretraining on link prediction, masked node modelling, distance prediction, and contrastive clustering.
  • Code: not released

Vision-Language Models

Vision as Unified Multimodal Generation

↑ 45 πŸ“š 19660 β˜… 0 Jul 6
  • Problem: Computer vision tasks require task-specific architectures; can a unified multimodal model handle diverse vision tasks through text and image generation?
  • Model: SenseNova-Vision: unified multimodal model that formulates vision tasks as text and image generation without task-specific heads
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

↑ 51 πŸ“š 44 β˜… 0 Jul 7
  • Problem: Vision-language-action models struggle with long-horizon robotic tasks due to Markovian assumptions and memory mechanisms that remain external to latent reasoning space.
  • Model: LaMem-VLA: dual latent memory framework with curator, seeker, condenser, and weaver components that integrate short-term and long-term memory tokens directly into VLA reasoning.
  • Code: not released

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

↑ 30 πŸ“š 2058 β˜… 0 Jul 9
  • Problem: Existing agent benchmarks use sandboxed environments and single-turn evaluation, making it difficult to identify specific capability failures in real-world proactive agent tasks.
  • Model: UniClawBench: capability-driven benchmark with five foundational dimensions (Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, Cross-Platform Coordination) evaluated via three-role closed-loop assessment in live Docker containers.
  • Code: HKU-MMLab/UniClawBench

World Models

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

↑ 76 πŸ“š 7987 β˜… 0 Jul 7
  • Problem: Robot learning is bottlenecked by physical teleoperation; decoupling data collection from hardware constraints requires a scalable digital alternative.
  • Model: RynnWorld-Teleop: robot-centric action-conditioned world model using depth-aware skeletal conditioning, progressive human-to-robot diffusion transformer training, and streaming autoregressive distillation for real-time egocentric video synthesis.
  • Code: alibaba-damo-academy/RynnWorld-Teleop

AlayaWorld: Long-Horizon and Playable Video World Generation

↑ 85 πŸ“š 8 β˜… 0 Jul 7
  • Problem: Building interactive, playable video worlds requires solving control, consistency, stability, and real-time generation challenges in generative models.
  • Model: "AlayaWorld": autoregressive DiT with prompt-switching, camera-control module, 3D cache, history-compression, error bank, and few-step distillation for interactive world generation
  • Code: not released

Spatial Single-Cell Study

DriftST: One-Step Generative Inference of Spatial Transcriptomics from H\&E Histology

↑ 0 πŸ“š 11 β˜… 0 Jul 6
  • Problem: Inferring spatial transcriptomics from H&E histology requires balancing generative fidelity with inference speed while capturing gene dependencies and handling multiple resolutions.
  • Model: "DriftST": Cellular Drifting generative model with STransformer for one-step H&E-to-expression inference using co-expression attention and gene residual gating.
  • Code: yyh030806/DriftST

COAST: Context-Aware Differential Learning for Gene Expression Prediction in Spatial Transcriptomics

↑ 0 β˜… 0 Jul 10
  • Problem: Existing spatial gene expression prediction methods from histology supervise absolute expression but rarely exploit relative expression relationships between spots.
  • Model: COAST: context-aware differential learning framework using type-specific feature modulation and Transformer encoder with joint absolute and differential regression objectives
  • Code: not released

Score Distributions, Not Cells: Evaluating Single-Cell Perturbations Under Class Overlap

↑ 0 β˜… 0 Jul 6
  • Problem: Single-cell perturbation classification uses per-cell accuracy, which measures class overlap rather than model quality when perturbations produce overlapping cell distributions.
  • Model: "Classifier Discrimination Score (CDS): averages classifier probability vectors across all cells of a perturbation to form population profiles for ranking perturbations"
  • Code: not released

Week of

Graph Γ— LLM

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

↑ 9 πŸ“š 54647 β˜… 0 Jun 30
  • Problem: Standard LLMs produce fluent but weakly traceable responses to materials design problems, making it difficult to verify intermediate reasoning supports final answers.
  • Model: Graph-PRefLexOR: graph-native reasoning model fine-tuned with Group Relative Policy Optimization (GRPO) to organize multi-step scientific reasoning into explicit phases (brainstorm, graph construction, pattern extraction, synthesis) linking neural generation with symbolic relational structure.
  • Code: lamm-mit/graph-preflexor-grpo

NoPA: Non-Parametric Online 3D Scene Graph Generation

↑ 3 πŸ“š 9846 β˜… 0 Jun 30
  • Problem: Real-time 3D scene graph generation requires balancing geometric detail preservation with computational efficiency, as Gaussian-based object approximations lose geometric information and produce unstable merging.
  • Model: NoPA: Non-parametric representation using fixed-size particle sets per object with Maximum Mean Discrepancy-based merging and relationship propagation for online 3D scene graph generation.
  • Code: ShunChengWu/3DSSG

Vision-Language Models

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

↑ 22 πŸ“š 68732 β˜… 0 Jun 30
  • Problem: Multimodal LLMs suffer from language-space bottleneck in reasoning; variational latent reasoning introduces train-inference mismatch where the training posterior exploits answer information unavailable at test time.
  • Model: "Asymmetric Mutual Variational Learning (AMVL): bidirectional KL calibration framework combining forward prior alignment with reverse posterior regularization for continuous multimodal latent reasoning"
  • Code: not released

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

↑ 20 πŸ“š 16552 β˜… 0 Jun 29
  • Problem: Medical multimodal LLMs suffer from sparse credit assignment and cascading reasoning failures that propagate errors in clinical image reasoning tasks.
  • Model: Medical Reasoning-aware Policy Optimization (MRPO): GRPO-based RL algorithm with step-wise process rewards that assigns exponentially larger penalties to tokens in earlier invalid reasoning steps.
  • Code: dmis-lab/MRPO

Optimizing Visual Generative Models via Distribution-wise Rewards

↑ 14 πŸ“š 5953 β˜… 0 Jul 1
  • Problem: Sample-wise RL rewards for visual generative models cause reward hacking, degrading image diversity and introducing visual artifacts.
  • Model: approach: Distribution-wise RL framework using subset-replace strategy to efficiently optimize generative models against FID-based rewards, with post-hoc model merging coefficient optimization.
  • Code: not released

World Models

MemLearner: Learning to Query Context memory for Video World Models

↑ 25 πŸ“š 4014 β˜… 0 Jun 30
  • Problem: Video world models lack effective memory mechanisms, causing scene inconsistency over extended durations when dealing with occlusions and dynamic objects.
  • Model: MemLearner: learning-based adaptive context query method using query tokens to retrieve relevant historical frames for memory-augmented video generation.
  • Code: not released

Valdi: Value Diffusion World Models

↑ 12 πŸ“š 1721 β˜… 0 Jul 1
  • Problem: Diffusion models excel at modeling uncertain dynamics but their iterative inference is too slow for low-latency Model Predictive Control in latent space.
  • Model: "Valdi: Value Diffusion World Models" β€” latent diffusion dynamics model jointly trained with value and reward prediction for online MPC, using single-step inference.
  • Code: Kit115/ValueDiffusionWorldModels

Spatial Single-Cell Study

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

↑ 0 πŸ“š 2539 β˜… 0 Jun 30
  • Problem: Neural networks compress distinct biological concepts into superposed latent dimensions, corrupting representational geometry and hindering interpretability of high-dimensional patient-neuronal image data.
  • Model: "GW-map: Gromov-Wasserstein optimal transport method for aligning sparse autoencoder image representations with scRNA-seq data"
  • Code: jijihihi/Bio_superposition
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Building artificial intelligence virtual tissue (AIVT) for tissue state representation, feature prediction, and dynamic simulation

↑ 0 πŸ“š 1063 β˜… 0 Jun 29
  • Problem: Conventional computational models inadequately capture tissue complexity as spatially organized, multiscale biological systems.
  • Model: "AI Virtual Tissue (AIVT)": framework learning unified, spatially resolved tissue state representations from multimodal data for feature prediction and spatiotemporal dynamics simulation.
  • Code: not released
June 2026 48 papers

Week of

Graph Γ— LLM

Vision-Language Models

World Models

Spatial Single-Cell Study

Week of

Graph Γ— LLM

Vision-Language Models

World Models

Kairos: A Native World Model Stack for Physical AI

↑ 35 β˜… 0 Jun 15
  • Problem: World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world know
  • Model:
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Week of

Graph Γ— LLM

APEX: A Network-Native Time-Series Foundation Model for Forecasting and Anomaly Detection for Wireless Edge Operations

↑ 1 πŸ“š 150 β˜… 0 Jun 9
  • Problem: Generic time-series foundation models fail to capture wireless network telemetry characteristicsβ€”bursty, zero-inflated signals with cross-protocol dependencies.
  • Model: APEX: network-native decoder-only transformer for forecasting enterprise AP telemetry and anomaly detection, available in cloud (269M) and edge (10.5M) variants.
  • Code: not released

Detecting Differences Is Not Understanding Structure: Large Language Models Fail at Graph Isomorphism

↑ 0 β˜… 0 Jun 8
  • Problem: LLMs achieve high accuracy on graph isomorphism detection but fail to maintain permutation invariance when nodes are relabeled, suggesting pattern exploitation rather than genuine structural reasoning.
  • Model: approach: diagnostic evaluation protocol testing permutation invariance of LLMs (GPT-4o, Gemini, Llama) on graph isomorphism tasks across multiple serialization formats and prompting strategies.
  • Code: not released

Vision-Language Models

InterleaveThinker: Reinforcing Agentic Interleaved Generation

↑ 77 πŸ“š 42325 β˜… 0 Jun 10
  • Problem: Image generators cannot produce interleaved text-image sequences due to architectural constraints, limiting applications in visual narratives and embodied manipulation.
  • Model: InterleaveThinker: multi-agent framework with planner and critic agents that retrofits existing image generators for interleaved generation using GRPO-based trajectory optimization.
  • Code: zhengdian1/InterleaveThinker

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

↑ 53 πŸ“š 8682 β˜… 0 Jun 11
  • Problem: Existing vision-language-action models lack laboratory-specific training data and cannot accommodate diverse robot embodiments needed for scientific protocol execution.
  • Model: LabVLA: vision-language-action model combining Qwen3-VL-4B-Instruct backbone with FAST action token pretraining and flow matching posttraining via DiT action expert
  • Code: not released

World Models

Avatar V: Scaling Video-Reference Avatar Video Generation

↑ 4 β˜… 0 Jun 10
  • Problem: Existing avatar video generation methods condition on single static images, failing to capture dynamic behavioral patterns and identity nuances required for production-quality talking avatars.
  • Model: Avatar V: production-scale framework for video-reference-conditioned avatar generation using Diffusion Transformer with Sparse Reference Attention, motion representation stream, and identity-aware super-resolution refinement.
  • Code: not released

Scale Buys Interpolation, Structure Buys a Horizon: Certified Predictability for Equivariant World Models

↑ 0 πŸ“š 17151 β˜… 0 Jun 11
  • Problem: World models lack per-prediction trustworthiness certificates and prediction horizons; average error does not indicate whether specific predictions can be trusted or for how long.
  • Model: approach: Equivariant latent world models with computable multi-step certification via Lyapunov spectrum stratification, proving orbit-constant error under equivariance and horizon bounds T_j(Ξ΅)∼log(1/Ξ΅)/Ξ»_j
  • Code: not released

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

↑ 2 β˜… 0 Jun 11
  • Problem: Existing robot world models fail to simultaneously achieve high fidelity, long-horizon consistency, and efficient inference for manipulation tasks.
  • Model: WEAVER (World Estimation Across Views for Embodied Reasoning): a multi-view world model combining flow matching, diffusion forcing, pretrained encoders, and latent reward prediction for robot manipulation.
  • Code: Lightning-AI/torchmetrics

Spatial Single-Cell Study

OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction

↑ 0 πŸ“š 329 β˜… 0 Jun 11
  • Problem: Predicting single-cell transcriptional responses to perturbations requires models that capture population-level distribution shifts without relying on complex auxiliary encoders or specialized latent spaces.
  • Model: OCOO-T: flow-matching-based Transformer model that directly denoises continuous gene expression profiles conditioned on perturbation embeddings, dosage, and cell covariates via adaptive layer normalization.
  • Code: not released

Adaptive spatial blocking for scalable clustering inference with applications to high-throughput spatial proteomics

↑ 0 β˜… 0 Jun 10
  • Problem: Existing Ripley's K-function methods for spatial clustering are computationally prohibitive for large-scale spatial proteomics data due to O(nΒ²) complexity.
  • Model: "B-KAMP" (block-based KAMP): adaptive spatial blocking algorithm aggregating clustering evidence across disjoint rectangular blocks with asymptotic normal inference.
  • Code: mingyugo/B_KAMP

Week of

Graph Γ— LLM

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

↑ 2 πŸ“š 1201 β˜… 0 Jun 2
  • Problem: Graph Language Models may develop internal pathologies where graph tokens become activation outliers without meaningfully representing graph structure.
  • Model: approach: mechanistic interpretability analysis of Graph Language Models (LLaGA, TEA-GLM) through graph sink token detection and intervention experiments
  • Code: not released

The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning

↑ 0 πŸ“š 6199 β˜… 0 Jun 4
  • Problem: Standard benchmarks average model performance across heterogeneous datasets, obscuring geometry-dependent performance variations and misleading conclusions about generalization.
  • Model: "CurvBench": curvature-stratified evaluation framework partitioning datasets by intrinsic geometry (positive, negative, near-zero curvature) to reveal geometry-dependent model performance trade-offs.
  • Code: https://sirbabbage.github.io/CurvBench_HOME/

A Graph Foundation Model with Spectral Parsing and Prototype-Guided Spatial Propagation

↑ 0 πŸ“š 12963 β˜… 0 Jun 2
  • Problem: Graph foundation models struggle with cross-graph transfer due to diverse graph structures and entangled spectral components requiring different propagation behaviors.
  • Model: SPG: graph foundation model combining learnable Chebyshev spectral filters for feature decomposition with Gromov-Wasserstein prototype geometry for transferable structural knowledge.
  • Code: not released

Vision-Language Models

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

↑ 21 πŸ“š 27459 β˜… 0 Jun 3
  • Problem: Existing unified video generation and editing models are computationally expensive, relying on massive parameters and token concatenation that quadruples self-attention complexity.
  • Model: LoomVideo: 5B-parameter unified video generation and editing architecture using MLLM encoder, Deepstack injection, and zero-overhead Scale-and-Add conditioning.
  • Code: MSALab-PKU/LoomVideo

Personal AI Agent for Camera Roll VQA

↑ 19 πŸ“š 7113 β˜… 0 Jun 2
  • Problem: AI assistants cannot efficiently answer personalized questions over thousands of personal camera roll photos spanning years.
  • Model: camroll-agent: conversational AI agent with hierarchical memory and tools for efficient navigation over large personalized visual memory
  • Code: not released

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

↑ 9 πŸ“š 3973 β˜… 0 Jun 4
  • Problem: Vision-Language Models struggle with spatial reasoning beyond observed images, failing to infer unobserved layouts and reason from alternative viewpoints with limited egocentric observations.
  • Model: "Astra: agentic spatial reasoning framework coupling Astra-VL (RL-trained Qwen3-VL policy) with Astra-WM (Bagel-based world simulator for action-conditioned novel-view generation with view consistency tuning)"
  • Code: not released

World Models

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

↑ 16 πŸ“š 4009 β˜… 0 Jun 3
  • Problem: Standard video generation benchmarks measure visual quality, not whether generated robot manipulation videos produce executable physical behavior.
  • Model: "Dream.exe: evaluation framework for video-to-execution grounding in robotic manipulation, combining video assessment, trajectory extraction, and physics simulator execution.
  • Code: not released

Towards World Models in Biomedical Research

↑ 0 πŸ“š 38743 β˜… 0 Jun 4
  • Problem: Current biomedical AI focuses on static pattern recognition rather than simulating how biological systems evolve under interventions and perturbations.
  • Model: approach: biomedical world models that learn latent representations of biological states and intervention-conditioned dynamics to simulate future trajectories
  • Code: not released

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

↑ 0 πŸ“š 5409 β˜… 0 Jun 4
  • Problem: Existing world models for robot evaluation use open-loop prediction, but VLA policies operate in closed-loop feedback; a world model for policy-in-the-loop evaluation is needed.
  • Model: "PiL-World": chunk-wise world model for closed-loop VLA evaluation using action-derived visual control, latent multi-view history conditioning, and joint multi-view prediction.
  • Code: not released

Spatial Single-Cell Study

Do Foundation Models See Biology? Evaluating Attention Coherence with Spatial Transcriptomics in Glioblastoma

↑ 0 πŸ“š 3679 β˜… 0 Jun 3
  • Problem: Whether attention maps from pathology foundation models capture genuine biological signals remains unknown, hindering clinical trust and regulatory approval.
  • Model: approach: spatial transcriptomics-based framework for objective evaluation of attention coherence in pathology foundation models using attention-based multiple instance learning
  • Code: not released

GC-MoE: Genomics-Guided Cell-Type-Specific Mixture of Experts for Histology-Based Single-Cell Spatial Transcriptomics

↑ 0 πŸ“š 1102 β˜… 0 Jun 1
  • Problem: Predict gene expression for individual cells from histology images and cell locations, accounting for cell-type-dependent expression variability.
  • Model: GC-MoE: Genomics-Guided Cell-Type-Specific Mixture of Experts that routes cells to type-specific experts and incorporates cell-type co-expression priors from scRNA-seq data.
  • Code: not released

Week of

Graph Γ— LLM

Vision-Language Models

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

↑ 117 πŸ“š 674 β˜… 0 May 27
  • Problem: Embodied AI systems are fragmented across task families and robot embodiments, limiting generalization across manipulation, navigation, and diverse platforms.
  • Model: Qwen-VLA: unified vision-language-action model extending Qwen3.5-4B with DiT-based flow-matching action decoder for cross-task, cross-embodiment embodied control.
  • Code: QwenLM/Qwen-VLA

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

↑ 45 πŸ“š 9553 β˜… 0 May 27
  • Problem: Vision-language models achieve high spatial reasoning benchmark scores but may rely on statistical shortcuts rather than structured 3D understanding.
  • Model: approach: Representation-level analysis framework using minimal contrastive pairs to measure spatial axis organization and disentanglement in VLM embeddings, plus SpatialTunnel synthetic benchmark.
  • Code: not released

World Models

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

↑ 50 πŸ“š 36656 β˜… 0 May 27
  • Problem: Converting video diffusion foundation models into real-time interactive world models requires scattered techniques across data, training, and inference pipelines.
  • Model: minWM: full-stack framework converting T2V/TI2V bidirectional diffusion models into camera-controllable few-step autoregressive video world models via Causal Forcing/Forcing++ distillation
  • Code: shengshu-ai/minWM

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

↑ 42 πŸ“š 16919 β˜… 0 May 27
  • Problem: Video diffusion models may perceive temporal direction without understanding causality; existing benchmarks rely on synthetic data with limited real-world generalization.
  • Model: YoCausal: two-level benchmark with Reverse Surprise Index (RSI) and Causality Cognition Index (CCI) metrics for evaluating causal cognition in video diffusion models.
  • Code: genmoai/models

AdaState: Self-Evolving Anchors for Streaming Video Generation

↑ 6 πŸ“š 1390 β˜… 0 May 27
  • Problem: Autoregressive video diffusion models anchor to static first frames, suppressing dynamics and locking scene composition despite natural evolution during generation.
  • Model: "AdaState": replaces static first-frame anchor with adaptive latent state that denoises alongside content at each chunk, evolving with generated scenes via relative time formulation.
  • Code: not released
May 2026 22 papers

Week of

Graph Γ— LLM

S2Aligner: Pair-Efficient and Transferable Pre-Training for Sparse Text-Attributed Graphs

↑ 0 πŸ“š 2080 β˜… 0 May 18
  • Problem: Graph foundation models struggle with sparse text-attributed graphs where node texts are missing, noisy, or uneven, causing unreliable structure-semantics alignment and transfer bias.
  • Model: S2Aligner: sparsity-aware and structure-enhanced LLM-as-Aligner framework that decouples semantic alignment from structural modeling via content-structure factorization and sparsity-aware cross-domain risk balancing.
  • Code: not released

Deep Neural Sheaf Diffusion

↑ 0 πŸ“š 1013 β˜… 0 May 18
  • Problem: Scaling Graph Neural Networks to depth is hindered by representation collapse and vanishing signals in existing sheaf diffusion methods.
  • Model: Deep Neural Sheaf Diffusion (DNSD): sheaf-based GNN replacing sheaf Laplacian with adjacency operator, adding normalization, odd nonlinearities, and gating to maintain informative signals across layers.
  • Code: not released

Vision-Language Models

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

↑ 42 πŸ“š 15499 β˜… 0 May 20
  • Problem: Current multimodal LLMs struggle with audio-visual reasoning because text-based chain-of-thought compresses continuous signals into discrete tokens, losing temporal grounding.
  • Model: LatentOmni: cross-modal reasoning framework interleaving textual reasoning with audio-visual latent states, using feature-level supervision and Omni-Sync Position Embedding for temporal alignment.
  • Code: not released

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

↑ 9 πŸ“š 19747 β˜… 0 May 19
  • Problem: Discrete autoregressive text-to-image models suffer from latent covariate shift when optimizing only the policy with a frozen decoder, causing alignment-fidelity trade-offs.
  • Model: RankE: End-to-end post-training framework for discrete text-to-image generation that co-evolves the AR policy and VQ decoder through alternating ranking-based optimization.
  • Code: not released

World Models

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

↑ 19 πŸ“š 13800 β˜… 0 May 20
  • Problem: Quantizing autoregressive video diffusion models is unexplored; standard quantization schemes designed for bidirectional diffusion transformers perform suboptimally on ARVDs.
  • Model: Q-ARVD: quantization framework for autoregressive video diffusion models using final-quality-guided frame-weighting and outlier-aware adaptive dual-scale quantization
  • Code: not released

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

↑ 7 πŸ“š 41260 β˜… 0 May 21
  • Problem: Current agentic LLMs lack control over when and how to plan, causing inefficient token use without reliable accuracy gains.
  • Model: "SRΒ²AM" (Self-Regulated Simulative Reasoning Agentic LLM): decomposes decision-making into three systemsβ€”simulative reasoning via world model, self-regulation via learned configurator, and reactive executionβ€”implemented as distinct chain-of-thought stages within an LLM.
  • Code: sailing-lab/sr2am

Spatial Single-Cell Study

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

↑ 0 πŸ“š 20431 β˜… 0 May 19
  • Problem: Multi-agent workflow design in open-ended scientific settings lacks curated training sets, reliable metrics, and standardized interfaces between tools and agents.
  • Model: AgentCo-op: retrieval-based synthesis framework that composes reusable skills, tools, and external agents into executable workflows through typed artifact handoffs and bounded evidence-guided local repair.
  • Code: ma-compbio-lab/AgentCo-Op

Week of

Graph Γ— LLM

Vision-Language Models

MMSkills: Towards Multimodal Skills for General Visual Agents

↑ 99 πŸ“š 30599 β˜… 0 May 13
  • Problem: Visual agents need reusable multimodal procedural knowledge that binds actions to visual state recognition and decision-making, beyond text-only skills.
  • Model: "MMSkills: framework for representing, generating, and utilizing reusable multimodal procedures for visual agents. Each skill couples textual procedures with runtime state cards and multi-view keyframes, generated via trajectory-to-skill generator and consulted via branch loading.
  • Code: DeepExperience/MMSkills

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

↑ 71 πŸ“š 3089 β˜… 0 May 13
  • Problem: No benchmark systematically compares long-context LVLMs and memory-augmented agents on multimodal multi-session conversations requiring visual evidence.
  • Model: "MemLens": benchmark with 789 questions across five memory abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge update, answer refusal) at four context lengths (32K-256K tokens)
  • Code: xrenaf/MEMLENS

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

↑ 58 πŸ“š 1421 β˜… 0 May 13
  • Problem: Existing multimodal agent memory evaluations fail to assess whether agents preserve fine-grained visual evidence needed for reasoning over time.
  • Model: "MemEye: a visual-centric evaluation framework measuring visual evidence granularity and reasoning complexity in multimodal agent memory"
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

World Models

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

↑ 87 πŸ“š 36802 β˜… 0 May 14
  • Problem: Existing AR diffusion distillation methods require 4+ sampling steps; frame-wise 1–2 step generation needs efficient, scalable student initialization.
  • Model: "Causal Forcing++": AR diffusion distillation pipeline using causal consistency distillation for few-step student initialization, avoiding expensive full PF-ODE trajectory precomputation.
  • Code: thu-ml/Causal-Forcing

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

↑ 38 πŸ“š 18942 β˜… 0 May 13
  • Problem: Existing camera-controlled video generation methods require large-scale camera-annotated data or expensive test-time optimization; no simple way to leverage pretrained video models' latent camera-control capability.
  • Model: "Warp-as-History": converts camera-induced geometric warps into camera-warped pseudo-history fed through pretrained video models' native history pathway, with target-frame positional alignment and visible-token selection.
  • Code: yyfz/Warp-as-History

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

↑ 7 πŸ“š 19915 β˜… 0 May 13
  • Problem: Existing video rewards fail to reliably score human motion realism because they operate in 2D pixel space without explicitly modeling 3D body physics and constraints.
  • Model: "PhyMotion": physics-grounded motion reward that recovers SMPL meshes, retargets to MuJoCo simulator, and evaluates motion via three axes (kinematic plausibility, contact/balance, dynamic feasibility)
  • Code: not released

Spatial Single-Cell Study

DUET: Dual-Paradigm Adaptive Expert Triage with Single-cell Inductive Prior for Spatial Transcriptomics Prediction

↑ 0 πŸ“š 5699 β˜… 0 May 13
  • Problem: Existing methods for inferring spatial gene expression from histology images oversimplify morphology-to-expression mapping and underutilize large-scale single-cell data as biological constraints.
  • Model: DUET: dual-paradigm framework synergizing parametric regression and memory-based retrieval with cellular inductive priors and adaptive expert triage for spatial transcriptomics prediction.
  • Code: Junchao-Zhu/DUET

StateXDiff: Cell State-Contextualized Multimodal Diffusion for Single-Cell Perturbation Prediction

↑ 0 β˜… 0 May 15
  • Problem: Predicting drug-induced cellular state changes at single-cell resolution under out-of-distribution conditions with limited multimodal information.
  • Model: StateXDiff: cell State-contextualized multimodal Diffusion framework integrating transcriptomic and pseudo-protein representations with mechanism-aware drug templates via latent conditional diffusion.
  • Code: not released