Paper Tracking Β· Archive

Every weekly run is preserved. Click a month to expand the weeks inside. Latest weeks at the top. ← Back to current week

August 2026 33 papers

Week of

Graph Γ— LLM

R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

↑ 0 πŸ“š 45123 β˜… 0 Aug 11
  • Problem: Object-centric question answering in long egocentric video requires preserving persistent object identity and structured spatial change, which caption-based methods fail to capture.
  • Model: R4DSG: Relative 4D scene graph memory that indexes video by time, place, persistent objects, and anchor-relative transitions without requiring global world coordinates.
  • Code: https://dualtransparency.github.io/R4DSG/

Vision-Language Models

Intern-S2-Preview: Scientific Agentic Foundation Model

↑ 56 β˜… 0 Aug 12
  • Problem: AI systems need to reason over heterogeneous scientific modalities, interact with tools, and sustain progress over long task horizons for scientific discovery.
  • Model: "Intern-S2-Preview-397B": multimodal scientific foundation model with time series forecasting, agentic RL training, and modular Memory Decoder for domain specialization.
  • Code: InternLM/xtuner

AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss

↑ 26 πŸ“š 9561 β˜… 0 Aug 10
  • Problem: Static pretrained feature spaces used in FrΓ©chet distance losses enable FrΓ©chet hacking, where optimized metrics improve while visual quality degrades.
  • Model: "AdvFD: Adversarial FrΓ©chet Distance" β€” combines static pretrained encoders with an adversarially learned adaptive representation and real-feature whitening to expose residual distribution mismatches during generator post-training.
  • Code: not released

World Models

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

↑ 119 πŸ“š 14838 β˜… 0 Aug 13
  • Problem: Interactive world models struggle to balance persistent memory, low-latency response, and long-horizon generation due to conflicting computational demands.
  • Model: "Evoke": world model with externalized scene geometry bank and linear-attention teacher for long-horizon supervision enabling few-step student generation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

↑ 93 πŸ“š 16421 β˜… 0 Aug 13
  • Problem: Action-conditioned video world models struggle to faithfully predict robot manipulation while preserving arm identity, object state, and scene geometry.
  • Model: DreamX-Phi 1.0: action-conditioned video world model using PRoPE-style SE(3) geometric encoding, depth supervision, SAM3 masks, and V-JEPA teacher for faithful robotic manipulation prediction.
  • Code: AMAP-ML/DreamX-Phi

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

↑ 43 πŸ“š 37560 β˜… 0 Aug 13
  • Problem: Existing world model benchmarks use fixed action sequences that fail to enable fair comparison across models with different action granularities and response dynamics.
  • Model: PlayWorld: benchmark with multi-modal Agent Player that adaptively executes long-horizon objectives to evaluate world models on geometry consistency, interaction fidelity, and state evolution.
  • Code: kxding/PlayWorld

Spatial Single-Cell Study

Disentangled Shared Representations Improve Morpho-Transcriptomic Integration

↑ 0 β˜… 0 Aug 14
  • Problem: Standard multimodal models for spatial transcriptomics compress H&E and gene expression into shared latent space without explicitly separating shared versus modality-specific variation.
  • Model: approach: Disentangled multimodal representation learning comparing VAE-based (MMVAE+, MMVAE+sg) and contrastive (disSSL) models that factorize shared and private latent components for H&E and spatial transcriptomics integration.
  • Code: not released

Program-space Diffusion for Morphology-to-Transcriptomics Prediction

↑ 0 β˜… 0 Aug 14
  • Problem: Predicting spatial gene expression from histology is computationally expensive and ignores coordinated transcriptional variation.
  • Model: approach: Conditional diffusion model operating in transcriptional program space learned via consensus non-negative matrix factorization (cNMF), with Pearson residual normalization for count data.
  • Code: not released

Week of

Graph Γ— LLM

NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning

↑ 0 πŸ“š 44134 β˜… 0 Aug 5
  • Problem: Designing JEPA-style latent prediction objectives for node-level graph self-supervised learning and determining which structural signals the predictor should condition on.
  • Model: NodeJEPA: joint-embedding predictive architecture that masks structure-aware k-hop ego-subgraphs and predicts latent node representations using a structure-conditioned predictor with spectral and centrality descriptors.
  • Code: OliverZ-dot/Node-Jepa

EvtGraph: Event-Adaptive Compression for Sparse Temporal Graph Learning in Multimodal Time Series

↑ 0 πŸ“š 4351 β˜… 0 Aug 5
  • Problem: Multimodal temporal data have non-uniform information density, but standard models allocate computation uniformly, causing inefficiency and node explosion in graph representations.
  • Model: "EvtGraph: Event-adaptive compression framework with node budget control and temporally constrained sparse graph reasoning for multimodal time series.
  • Code: not released

Vision-Language Models

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation

↑ 55 πŸ“š 93935 β˜… 0 Aug 4
  • Problem: Text-to-image models lack unified agentic control over reasoning, tool use, and generation for complex open-world image synthesis tasks.
  • Model: ToolArtist: unified multimodal model trained with supervised fine-tuning and agentic RL (RAD-GRPO) to dynamically orchestrate reasoning, external tools, and image generation within a single policy
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

↑ 57 πŸ“š 30079 β˜… 0 Aug 4
  • Problem: Fundamental mechanisms governing how modalities interact during unified multimodal pretraining remain underexplored.
  • Model: approach: systematic empirical exploration of unified multimodal pretraining (language, visual understanding, visual generation) using controlled experiments to derive design principles and efficient recipes.
  • Code: not released

ChronoVision: Temporal Reasoning via Latent State Reconstruction

↑ 38 πŸ“š 21747 β˜… 0 Aug 5
  • Problem: Multimodal LLMs struggle with temporal reasoning tasks because language-based explanations fail to capture continuous visual transformations.
  • Model: ChronoVision: multimodal framework with Reconstructive Visual Head predicting latent final states, ROI Attention Locating module, and RL-based implicit process grounding.
  • Code: not released

World Models

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

↑ 38 πŸ“š 5338 β˜… 0 Aug 6
  • Problem: Training LLM agents for long-horizon tool use requires costly external environments or unreliable simulators; how can agents internalize environment dynamics without external interaction?
  • Model: EnvACE: agentic RL method where a policy alternates between acting (generating tool calls) and world rehearsal (playing the environment to produce action responses), jointly optimized end-to-end with task-success rewards.
  • Code: Within-yao/EnvACE

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

↑ 32 πŸ“š 6714 β˜… 0 Aug 6
  • Problem: Existing economic simulations lack generative mechanisms where heterogeneous agents interact to produce emergent aggregate outcomes with persistent empirical alignment.
  • Model: approach: Economic World Models (EWM) β€” a modular runtime system organizing heterogeneous agents, environment mechanics, co-evolution, and real-world alignment into a closed-loop generative engine for economic simulation.
  • Code: FreedomIntelligence/Awesome-Economic-World-Models

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

↑ 12 πŸ“š 31569 β˜… 0 Aug 5
  • Problem: Interactive video world models suffer from compounding errors in long-horizon rollouts with no ground-truth supervision to measure accumulated drift.
  • Model: "WorldCycle: Self-Verifiable Reinforcement Learning framework using reversible action cycles with spatial closure and temporal consistency rewards for long-horizon video world models"
  • Code: not released

Spatial Single-Cell Study

Control-Anchored Residual Flow Matching Conditioned on Gene Geometry for Virtual Cell Perturbation Modeling

↑ 0 β˜… 0 Aug 7
  • Problem: Existing graph-based virtual cell models conflate stable gene associations with intervention-specific response pathways, lacking separation between structural priors and learned dynamics.
  • Model: GeneGeoFlow: control-anchored residual flow conditioned on perturbation-specific gene geometry derived from multi-scale spectral coordinates of Gene Ontology and coexpression networks.
  • Code: not released

Week of

Graph Γ— LLM

Harness-G: A Graph-Structured Harness for Search Agents

↑ 3 πŸ“š 14432 β˜… 0 Jul 29
  • Problem: RL search agents suffer from retrieval-equivalence collapse where distinct queries yield overlapping evidence, preventing effective credit assignment.
  • Model: Harness-G: graph-structured retrieval framework that reformulates free-form query generation as finite action selection over a paragraph-sentence-entity graph, with Structured Non-myopic Credit (SNC) for credit assignment.
  • Code: 7HHHHH/Harness-G

GraphQAG: A Knowledge-Graph-Guided Visual Analytics Framework for Question-Answer Pairs Generation

↑ 0 πŸ“š 25208 β˜… 0 Jul 29
  • Problem: Existing QA generation methods fail to capture fragmented knowledge distributed across paragraphs and multi-entity relationships in long documents.
  • Model: "GraphQAG: knowledge-graph-guided visual analytics framework combining entity/relation extraction, graph-based generation space, and interactive visualization for LLM-based QA pair generation from documents.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Vision-Language Models

TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

↑ 130 πŸ“š 14078 β˜… 0 Jul 28
  • Problem: LLM-centric vision-language-action models incur high computational and memory overhead, limiting real-time robotic control and deployment on resource-constrained platforms.
  • Model: TurboVLA: direct V+Lβ†’A vision-language-action model using independent vision and text encoders with lightweight cross-attention interaction and continuous action chunk prediction.
  • Code: H-EmbodVis/TurboVLA

HumanCLAW: Can Vision-Language Models Act Through a Body?

↑ 71 πŸ“š 5852 β˜… 0 Jul 28
  • Problem: Vision-language models lack embodied self-awareness to make moment-to-moment action decisions through a physical humanoid body in closed-loop control.
  • Model: approach: HumanCLAW framework decoupling VLM decision-making from motor execution via atomic skill commands translated to full-body motion with physics simulation
  • Code: not released

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

↑ 37 πŸ“š 6210 β˜… 0 Jul 29
  • Problem: Existing text-to-image editors fail at anatomically plausible multi-person interaction synthesis, producing fused limbs and interpenetrating bodies that VLM judges fail to detect.
  • Model: "MPIE-Bench and MPIE-Eval: benchmark and mesh-based evaluation protocol for multi-person interaction editing using 3D human mesh reconstruction to score anatomical completeness and contact geometry.
  • Code: AnnLin0628/mpie-bench

World Models

PhiZero: A World Model Built Around Physical Language

↑ 160 πŸ“š 11539 β˜… 0 Jul 30
  • Problem: Existing video world models predict pixels directly, leaving physical dynamics implicit; how can we learn explicit representations of state transitions?
  • Model: PhiZero: world model using discrete physical language tokenizer and reasoner to infer state transitions then render videos
  • Code: not released

QQWorld: Quantile-Quantile Matching for World Model Regularization

↑ 23 πŸ“š 4488 β˜… 0 Jul 30
  • Problem: Latent world models regularized with Epps-Pulley objective suffer from heavy-tailed latent distributions due to vanishing corrective gradients for tail samples.
  • Model: QQWorld: quantile-quantile matching regularizer for latent world models, replacing Epps-Pulley with rank-matched Gaussian quantile alignment
  • Code: not released

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

↑ 17 πŸ“š 4839 β˜… 0 Jul 30
  • Problem: Existing video world models lack unified, frame-level action control that transfers across diverse dynamics and appearances.
  • Model: "ShadowDancer": interactive video world model learning unified dynamics representations from shadow pairs (video pairs with identical motion but resampled appearance) via cross-shadow prediction.
  • Code: not released

Spatial Single-Cell Study

BayesClint: Bayesian Multi-Scale Clustering and Multi-Sample Integration With Feature Selection for Spatial Transcriptomics Data

↑ 0 πŸ“š 439 β˜… 0 Jul 27
  • Problem: Existing spatial transcriptomics clustering methods perform dimension reduction and clustering separately and cluster at only a single scale, rather than jointly identifying cell types and spatial domains.
  • Model: "BayesClint": Bayesian hierarchical model for joint factor analysis, multi-scale clustering, multi-sample integration, and nested feature selection in spatial transcriptomics.
  • Code: AlvinSheng/BayesClint
July 2026 42 papers

Week of

Graph Γ— LLM

HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology

↑ 0 πŸ“š 8811 β˜… 0 Jul 23
  • Problem: Predicting spatially resolved gene expression from H&E histology while quantifying prediction confidence and accounting for tissue heterogeneity.
  • Model: HierarchicalDAEW: dual-graph architecture with domain-aware edge-weighted graph convolution for spot-level predictions and gene-level graph decoder fusing protein-protein interactions with co-expression, plus evidential uncertainty estimation via Normal-Inverse-Gamma loss.
  • Code: not released

Attacking Graph Foundation Models Through Their Shared Representation

↑ 0 πŸ“š 364 β˜… 0 Jul 20
  • Problem: Graph foundation models share a representation bottleneck (alignment layer) across domains, creating a novel attack surface not studied in prior adversarial graph work.
  • Model: approach: representation-space and input-space adversarial attacks targeting alignment layers in six graph foundation models (spectral tokenizers, text embeddings, discrete codebooks) at inference time.
  • Code: not released

Vision-Language Models

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

↑ 50 πŸ“š 34877 β˜… 0 Jul 21
  • Problem: Mobile agents struggle to follow natural language-described targets in crowded scenes because vision-language-action models reason in abstract spatial latents misaligned with explicit image detections.
  • Model: ReferTrack: VLA model that selects targets from indexed bounding boxes then decodes tracking waypoints, using temporal-viewpoint-bbox indicator tokens to preserve motion cues.
  • Code: MedlarTea/referTrack

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

↑ 71 πŸ“š 1033 β˜… 0 Jul 20
  • Problem: Large-scale visual generators are expensive to train, fine-tune, and deploy; efficient compact alternatives are needed.
  • Model: Mage-Flow: 4B-scale generative stack combining Mage-VAE lightweight tokenizer with Native-Resolution Multimodal Diffusion Transformer for text-to-image generation and instruction-based editing.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Visual Contrastive Self-Distillation

↑ 44 πŸ“š 2274 β˜… 0 Jul 22
  • Problem: On-policy self-distillation requires asymmetric information between teacher and student, typically through privileged answers or visual evidence signals.
  • Model: Visual Contrastive Self-Distillation (VCSD): contrasts EMA teacher predictions on original and content-erased images to sharpen token distributions for on-policy self-distillation.
  • Code: not released

World Models

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

↑ 33 πŸ“š 18895 β˜… 0 Jul 22
  • Problem: Existing video diffusion transformers require quadratic attention complexity, limiting efficiency for long-sequence high-resolution video generation on single GPUs.
  • Model: SANA-Video 2.0: hybrid linear-softmax attention video diffusion transformer with block attention residuals for efficient 720p video generation.
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report

↑ 56 πŸ“š 212 β˜… 0 Jul 20
  • Problem: Generating interactive, spatiotemporally consistent, long-horizon video worlds from text, images, or video with low latency response.
  • Model: AlayaWorld: 15B video diffusion transformer generating 24-fps video with bounded visual context (sink frame, temporal history, spatial memory) and discrete autoregressive distillation for efficient inference.
  • Code: AlayaLab/AlayaWorld

GraphVid: Interactive Graph-Controllable Video Generation

↑ 4 πŸ“š 1056 β˜… 0 Jul 22
  • Problem: Controllable video generation lacks precise multi-object interaction control; trajectory-based methods scale poorly and require dense annotations.
  • Model: GraphVid: graph-conditioned image-to-video generation model using directed interaction scene graphs with edge-aware graph reasoning and LoRA adaptation to a frozen video diffusion transformer.
  • Code: not released

Spatial Single-Cell Study

Auditing pretraining contamination in single-cell foundation model benchmarks

↑ 0 πŸ“š 593 β˜… 0 Jul 21
  • Problem: Single-cell foundation model benchmarks may contain pretraining data, conflating memorization with genuine generalization in zero-shot evaluations.
  • Model: "scContam": per-cell audit framework combining MinHash gene-set fingerprints and loss-based membership inference attack (MIA-scFM) to detect pretraining contamination.
  • Code: not released

Loom: Multi-Region Analysis of Spatial Transcriptomics with Local Neighborhoods and Global Trajectories

↑ 0 β˜… 0 Jul 24
  • Problem: Existing spatial transcriptomics platforms lack integrated support for pseudo-temporal trajectory analysis, multi-region comparison, and local microenvironment investigation.
  • Model: "Loom": visual computing system integrating spatial coordinate registration, dimensionality reduction, clustering, pseudotime simulation, and novel glyph-based encoding for spatiotemporal ST analysis
  • Code: ScheWann/Loom

Week of

Graph Γ— LLM

Surprisingly Simple and Effective Multi-Domain Graph Foundation Model through Graph-to-Table Alignment

↑ 0 πŸ“š 13352 β˜… 0 Jul 13
  • Problem: Existing graph foundation models rely heavily on textual attributes or limited training data; enabling tabular foundation models to capture graph structural information remains unexplored.
  • Model: GTAlign: Graph-to-Table Alignment framework combining a universal graph encoder with community-guided continual pre-training and tabular foundation models for text-free graph representation learning.
  • Code: not released

Vision-Language Models

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

↑ 151 πŸ“š 27844 β˜… 0 Jul 15
  • Problem: Existing video MLLMs lack generalization across diverse video types, suffer from high computational demands, and remain partially closed-source, hindering reproducibility.
  • Model: "VideoChat3": fully open video MLLM with Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for efficient spatiotemporal encoding across general, long-form, and streaming video understanding.
  • Code: MCG-NJU/VideoChat3

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

↑ 127 πŸ“š 922 β˜… 0 Jul 13
  • Problem: Closed-source multimodal systems achieve strong text-to-image generation and editing performance, but their methods remain undisclosed; demonstrating comparable results in open-source models under constrained budgets.
  • Model: Boogu-Image-0.1: unified multimodal understanding and generation model family with Base, Turbo, Edit, and Edit-Turbo variants for text-to-image generation, fast inference, and instruction-based editing.
  • Code: Boogu-Project/Boogu-Image
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

↑ 24 πŸ“š 29151 β˜… 0 Jul 15
  • Problem: Existing embodied AI models separate textual reasoning from visual goal imagination; a unified approach coupling language planning with visual state prediction is needed.
  • Model: Hy-Embodied-RxBrain: multimodal Mixture-of-Transformers foundation model for embodied cognition integrating joint language-visual reasoning and imagination in a single planning sequence.
  • Code: Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0

World Models

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

↑ 37 πŸ“š 7690 β˜… 0 Jul 14
  • Problem: Existing keyframe-conditioned video generation models lack comprehensive evaluation frameworks to measure both keyframe fidelity and overall video quality.
  • Model: "KeyFrame-Compass: benchmark and automated evaluation framework for keyframe-conditioned video generation combining six keyframe execution metrics and MLLM-based quality assessment"
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

↑ 32 πŸ“š 4837 β˜… 0 Jul 14
  • Problem: Existing benchmarks lack comprehensive evaluation of multi-reference-to-audio-video generation, which requires jointly reasoning over multiple references while maintaining synchronized visual and audio content.
  • Model: approach: MultiRef-Compassβ€”a benchmark with 350 curated samples, asset-composition pipeline, and hybrid evaluation framework integrating automatic metrics with MLLM-as-a-Judge across four dimensions (Basic Quality, Reference Consistency, Audio-Visual Consistency, Instruction Following).
  • Code: not released

From Pixels to States: Rethinking Interactive World Models as Game Engines

↑ 31 πŸ“š 16 β˜… 0 Jul 15
  • Problem: Interactive world models lack the explicit state representation and rule-governed dynamics of conventional game engines, limiting coherence and real-time interactivity.
  • Model: approach: Unified framework organizing interactive game world modeling along four dimensions (player action control, game state dynamics, state-observation persistence, real-time generation) grounded in the action-state-observation loop.
  • Code: not released

Spatial Single-Cell Study

A vision foundation model for single-cell biology via spatial gene cartography

↑ 0 πŸ“š 1073 β˜… 0 Jul 15
  • Problem: Single-cell foundation models based on language models discard gene relationships and expression magnitude by tokenizing transcriptomes.
  • Model: scVision: vision transformer foundation model that renders cells as continuous gene-expression images using optimal transport for spatial gene placement, pretrained via masked image modeling on 72 million human cells.
  • Code: not released

Week of

Graph Γ— LLM

Canopy: A Heterograph Foundation Model for Metabolic Engineering

↑ 0 β˜… 0 Jul 7
  • Problem: Existing computational approaches for metabolic engineering either ignore experimental data or discard relational structure of biological knowledge.
  • Model: Canopy: heterogeneous graph foundation model integrating 6.9M nodes across 13 types with multi-modal encodings (ESM-2, MoLFormer, PubMedBERT) and pretraining on link prediction, masked node modelling, distance prediction, and contrastive clustering.
  • Code: not released

Vision-Language Models

Vision as Unified Multimodal Generation

↑ 45 πŸ“š 19660 β˜… 0 Jul 6
  • Problem: Computer vision tasks require task-specific architectures; can a unified multimodal model handle diverse vision tasks through text and image generation?
  • Model: SenseNova-Vision: unified multimodal model that formulates vision tasks as text and image generation without task-specific heads
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

↑ 51 πŸ“š 44 β˜… 0 Jul 7
  • Problem: Vision-language-action models struggle with long-horizon robotic tasks due to Markovian assumptions and memory mechanisms that remain external to latent reasoning space.
  • Model: LaMem-VLA: dual latent memory framework with curator, seeker, condenser, and weaver components that integrate short-term and long-term memory tokens directly into VLA reasoning.
  • Code: not released

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

↑ 30 πŸ“š 2058 β˜… 0 Jul 9
  • Problem: Existing agent benchmarks use sandboxed environments and single-turn evaluation, making it difficult to identify specific capability failures in real-world proactive agent tasks.
  • Model: UniClawBench: capability-driven benchmark with five foundational dimensions (Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, Cross-Platform Coordination) evaluated via three-role closed-loop assessment in live Docker containers.
  • Code: HKU-MMLab/UniClawBench

World Models

RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

↑ 76 πŸ“š 7987 β˜… 0 Jul 7
  • Problem: Robot learning is bottlenecked by physical teleoperation; decoupling data collection from hardware constraints requires a scalable digital alternative.
  • Model: RynnWorld-Teleop: robot-centric action-conditioned world model using depth-aware skeletal conditioning, progressive human-to-robot diffusion transformer training, and streaming autoregressive distillation for real-time egocentric video synthesis.
  • Code: alibaba-damo-academy/RynnWorld-Teleop

AlayaWorld: Long-Horizon and Playable Video World Generation

↑ 85 πŸ“š 8 β˜… 0 Jul 7
  • Problem: Building interactive, playable video worlds requires solving control, consistency, stability, and real-time generation challenges in generative models.
  • Model: "AlayaWorld": autoregressive DiT with prompt-switching, camera-control module, 3D cache, history-compression, error bank, and few-step distillation for interactive world generation
  • Code: not released

Spatial Single-Cell Study

DriftST: One-Step Generative Inference of Spatial Transcriptomics from H\&E Histology

↑ 0 πŸ“š 11 β˜… 0 Jul 6
  • Problem: Inferring spatial transcriptomics from H&E histology requires balancing generative fidelity with inference speed while capturing gene dependencies and handling multiple resolutions.
  • Model: "DriftST": Cellular Drifting generative model with STransformer for one-step H&E-to-expression inference using co-expression attention and gene residual gating.
  • Code: yyh030806/DriftST

COAST: Context-Aware Differential Learning for Gene Expression Prediction in Spatial Transcriptomics

↑ 0 β˜… 0 Jul 10
  • Problem: Existing spatial gene expression prediction methods from histology supervise absolute expression but rarely exploit relative expression relationships between spots.
  • Model: COAST: context-aware differential learning framework using type-specific feature modulation and Transformer encoder with joint absolute and differential regression objectives
  • Code: not released

Score Distributions, Not Cells: Evaluating Single-Cell Perturbations Under Class Overlap

↑ 0 β˜… 0 Jul 6
  • Problem: Single-cell perturbation classification uses per-cell accuracy, which measures class overlap rather than model quality when perturbations produce overlapping cell distributions.
  • Model: "Classifier Discrimination Score (CDS): averages classifier probability vectors across all cells of a perturbation to form population profiles for ranking perturbations"
  • Code: not released

Week of

Graph Γ— LLM

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

↑ 9 πŸ“š 54647 β˜… 0 Jun 30
  • Problem: Standard LLMs produce fluent but weakly traceable responses to materials design problems, making it difficult to verify intermediate reasoning supports final answers.
  • Model: Graph-PRefLexOR: graph-native reasoning model fine-tuned with Group Relative Policy Optimization (GRPO) to organize multi-step scientific reasoning into explicit phases (brainstorm, graph construction, pattern extraction, synthesis) linking neural generation with symbolic relational structure.
  • Code: lamm-mit/graph-preflexor-grpo

NoPA: Non-Parametric Online 3D Scene Graph Generation

↑ 3 πŸ“š 9846 β˜… 0 Jun 30
  • Problem: Real-time 3D scene graph generation requires balancing geometric detail preservation with computational efficiency, as Gaussian-based object approximations lose geometric information and produce unstable merging.
  • Model: NoPA: Non-parametric representation using fixed-size particle sets per object with Maximum Mean Discrepancy-based merging and relationship propagation for online 3D scene graph generation.
  • Code: ShunChengWu/3DSSG

Vision-Language Models

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

↑ 22 πŸ“š 68732 β˜… 0 Jun 30
  • Problem: Multimodal LLMs suffer from language-space bottleneck in reasoning; variational latent reasoning introduces train-inference mismatch where the training posterior exploits answer information unavailable at test time.
  • Model: "Asymmetric Mutual Variational Learning (AMVL): bidirectional KL calibration framework combining forward prior alignment with reverse posterior regularization for continuous multimodal latent reasoning"
  • Code: not released

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

↑ 20 πŸ“š 16552 β˜… 0 Jun 29
  • Problem: Medical multimodal LLMs suffer from sparse credit assignment and cascading reasoning failures that propagate errors in clinical image reasoning tasks.
  • Model: Medical Reasoning-aware Policy Optimization (MRPO): GRPO-based RL algorithm with step-wise process rewards that assigns exponentially larger penalties to tokens in earlier invalid reasoning steps.
  • Code: dmis-lab/MRPO

Optimizing Visual Generative Models via Distribution-wise Rewards

↑ 14 πŸ“š 5953 β˜… 0 Jul 1
  • Problem: Sample-wise RL rewards for visual generative models cause reward hacking, degrading image diversity and introducing visual artifacts.
  • Model: approach: Distribution-wise RL framework using subset-replace strategy to efficiently optimize generative models against FID-based rewards, with post-hoc model merging coefficient optimization.
  • Code: not released

World Models

MemLearner: Learning to Query Context memory for Video World Models

↑ 25 πŸ“š 4014 β˜… 0 Jun 30
  • Problem: Video world models lack effective memory mechanisms, causing scene inconsistency over extended durations when dealing with occlusions and dynamic objects.
  • Model: MemLearner: learning-based adaptive context query method using query tokens to retrieve relevant historical frames for memory-augmented video generation.
  • Code: not released

Valdi: Value Diffusion World Models

↑ 12 πŸ“š 1721 β˜… 0 Jul 1
  • Problem: Diffusion models excel at modeling uncertain dynamics but their iterative inference is too slow for low-latency Model Predictive Control in latent space.
  • Model: "Valdi: Value Diffusion World Models" β€” latent diffusion dynamics model jointly trained with value and reward prediction for online MPC, using single-step inference.
  • Code: Kit115/ValueDiffusionWorldModels

Spatial Single-Cell Study

Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images

↑ 0 πŸ“š 2539 β˜… 0 Jun 30
  • Problem: Neural networks compress distinct biological concepts into superposed latent dimensions, corrupting representational geometry and hindering interpretability of high-dimensional patient-neuronal image data.
  • Model: "GW-map: Gromov-Wasserstein optimal transport method for aligning sparse autoencoder image representations with scRNA-seq data"
  • Code: jijihihi/Bio_superposition
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Building artificial intelligence virtual tissue (AIVT) for tissue state representation, feature prediction, and dynamic simulation

↑ 0 πŸ“š 1063 β˜… 0 Jun 29
  • Problem: Conventional computational models inadequately capture tissue complexity as spatially organized, multiscale biological systems.
  • Model: "AI Virtual Tissue (AIVT)": framework learning unified, spatially resolved tissue state representations from multimodal data for feature prediction and spatiotemporal dynamics simulation.
  • Code: not released
June 2026 48 papers

Week of

Graph Γ— LLM

Vision-Language Models

World Models

Spatial Single-Cell Study

Week of

Graph Γ— LLM

Vision-Language Models

World Models

Kairos: A Native World Model Stack for Physical AI

↑ 35 β˜… 0 Jun 15
  • Problem: World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must natively acquire world know
  • Model:
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

Week of

Graph Γ— LLM

APEX: A Network-Native Time-Series Foundation Model for Forecasting and Anomaly Detection for Wireless Edge Operations

↑ 1 πŸ“š 150 β˜… 0 Jun 9
  • Problem: Generic time-series foundation models fail to capture wireless network telemetry characteristicsβ€”bursty, zero-inflated signals with cross-protocol dependencies.
  • Model: APEX: network-native decoder-only transformer for forecasting enterprise AP telemetry and anomaly detection, available in cloud (269M) and edge (10.5M) variants.
  • Code: not released

Detecting Differences Is Not Understanding Structure: Large Language Models Fail at Graph Isomorphism

↑ 0 β˜… 0 Jun 8
  • Problem: LLMs achieve high accuracy on graph isomorphism detection but fail to maintain permutation invariance when nodes are relabeled, suggesting pattern exploitation rather than genuine structural reasoning.
  • Model: approach: diagnostic evaluation protocol testing permutation invariance of LLMs (GPT-4o, Gemini, Llama) on graph isomorphism tasks across multiple serialization formats and prompting strategies.
  • Code: not released

Vision-Language Models

InterleaveThinker: Reinforcing Agentic Interleaved Generation

↑ 77 πŸ“š 42325 β˜… 0 Jun 10
  • Problem: Image generators cannot produce interleaved text-image sequences due to architectural constraints, limiting applications in visual narratives and embodied manipulation.
  • Model: InterleaveThinker: multi-agent framework with planner and critic agents that retrofits existing image generators for interleaved generation using GRPO-based trajectory optimization.
  • Code: zhengdian1/InterleaveThinker

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

↑ 53 πŸ“š 8682 β˜… 0 Jun 11
  • Problem: Existing vision-language-action models lack laboratory-specific training data and cannot accommodate diverse robot embodiments needed for scientific protocol execution.
  • Model: LabVLA: vision-language-action model combining Qwen3-VL-4B-Instruct backbone with FAST action token pretraining and flow matching posttraining via DiT action expert
  • Code: not released

World Models

Avatar V: Scaling Video-Reference Avatar Video Generation

↑ 4 β˜… 0 Jun 10
  • Problem: Existing avatar video generation methods condition on single static images, failing to capture dynamic behavioral patterns and identity nuances required for production-quality talking avatars.
  • Model: Avatar V: production-scale framework for video-reference-conditioned avatar generation using Diffusion Transformer with Sparse Reference Attention, motion representation stream, and identity-aware super-resolution refinement.
  • Code: not released

Scale Buys Interpolation, Structure Buys a Horizon: Certified Predictability for Equivariant World Models

↑ 0 πŸ“š 17151 β˜… 0 Jun 11
  • Problem: World models lack per-prediction trustworthiness certificates and prediction horizons; average error does not indicate whether specific predictions can be trusted or for how long.
  • Model: approach: Equivariant latent world models with computable multi-step certification via Lyapunov spectrum stratification, proving orbit-constant error under equivariance and horizon bounds T_j(Ξ΅)∼log(1/Ξ΅)/Ξ»_j
  • Code: not released

$\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulation

↑ 2 β˜… 0 Jun 11
  • Problem: Existing robot world models fail to simultaneously achieve high fidelity, long-horizon consistency, and efficient inference for manipulation tasks.
  • Model: WEAVER (World Estimation Across Views for Embodied Reasoning): a multi-view world model combining flow matching, diffusion forcing, pretrained encoders, and latent reward prediction for robot manipulation.
  • Code: Lightning-AI/torchmetrics

Spatial Single-Cell Study

OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Response Prediction

↑ 0 πŸ“š 329 β˜… 0 Jun 11
  • Problem: Predicting single-cell transcriptional responses to perturbations requires models that capture population-level distribution shifts without relying on complex auxiliary encoders or specialized latent spaces.
  • Model: OCOO-T: flow-matching-based Transformer model that directly denoises continuous gene expression profiles conditioned on perturbation embeddings, dosage, and cell covariates via adaptive layer normalization.
  • Code: not released

Adaptive spatial blocking for scalable clustering inference with applications to high-throughput spatial proteomics

↑ 0 β˜… 0 Jun 10
  • Problem: Existing Ripley's K-function methods for spatial clustering are computationally prohibitive for large-scale spatial proteomics data due to O(nΒ²) complexity.
  • Model: "B-KAMP" (block-based KAMP): adaptive spatial blocking algorithm aggregating clustering evidence across disjoint rectangular blocks with asymptotic normal inference.
  • Code: mingyugo/B_KAMP

Week of

Graph Γ— LLM

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

↑ 2 πŸ“š 1201 β˜… 0 Jun 2
  • Problem: Graph Language Models may develop internal pathologies where graph tokens become activation outliers without meaningfully representing graph structure.
  • Model: approach: mechanistic interpretability analysis of Graph Language Models (LLaGA, TEA-GLM) through graph sink token detection and intervention experiments
  • Code: not released

The Post-GCN Decade Revisited: Curvature-Stratified Evaluation of Relational Learning

↑ 0 πŸ“š 6199 β˜… 0 Jun 4
  • Problem: Standard benchmarks average model performance across heterogeneous datasets, obscuring geometry-dependent performance variations and misleading conclusions about generalization.
  • Model: "CurvBench": curvature-stratified evaluation framework partitioning datasets by intrinsic geometry (positive, negative, near-zero curvature) to reveal geometry-dependent model performance trade-offs.
  • Code: https://sirbabbage.github.io/CurvBench_HOME/

A Graph Foundation Model with Spectral Parsing and Prototype-Guided Spatial Propagation

↑ 0 πŸ“š 12963 β˜… 0 Jun 2
  • Problem: Graph foundation models struggle with cross-graph transfer due to diverse graph structures and entangled spectral components requiring different propagation behaviors.
  • Model: SPG: graph foundation model combining learnable Chebyshev spectral filters for feature decomposition with Gromov-Wasserstein prototype geometry for transferable structural knowledge.
  • Code: not released

Vision-Language Models

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

↑ 21 πŸ“š 27459 β˜… 0 Jun 3
  • Problem: Existing unified video generation and editing models are computationally expensive, relying on massive parameters and token concatenation that quadruples self-attention complexity.
  • Model: LoomVideo: 5B-parameter unified video generation and editing architecture using MLLM encoder, Deepstack injection, and zero-overhead Scale-and-Add conditioning.
  • Code: MSALab-PKU/LoomVideo

Personal AI Agent for Camera Roll VQA

↑ 19 πŸ“š 7113 β˜… 0 Jun 2
  • Problem: AI assistants cannot efficiently answer personalized questions over thousands of personal camera roll photos spanning years.
  • Model: camroll-agent: conversational AI agent with hierarchical memory and tools for efficient navigation over large personalized visual memory
  • Code: not released

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

↑ 9 πŸ“š 3973 β˜… 0 Jun 4
  • Problem: Vision-Language Models struggle with spatial reasoning beyond observed images, failing to infer unobserved layouts and reason from alternative viewpoints with limited egocentric observations.
  • Model: "Astra: agentic spatial reasoning framework coupling Astra-VL (RL-trained Qwen3-VL policy) with Astra-WM (Bagel-based world simulator for action-conditioned novel-view generation with view consistency tuning)"
  • Code: not released

World Models

Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?

↑ 16 πŸ“š 4009 β˜… 0 Jun 3
  • Problem: Standard video generation benchmarks measure visual quality, not whether generated robot manipulation videos produce executable physical behavior.
  • Model: "Dream.exe: evaluation framework for video-to-execution grounding in robotic manipulation, combining video assessment, trajectory extraction, and physics simulator execution.
  • Code: not released

Towards World Models in Biomedical Research

↑ 0 πŸ“š 38743 β˜… 0 Jun 4
  • Problem: Current biomedical AI focuses on static pattern recognition rather than simulating how biological systems evolve under interventions and perturbations.
  • Model: approach: biomedical world models that learn latent representations of biological states and intervention-conditioned dynamics to simulate future trajectories
  • Code: not released

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

↑ 0 πŸ“š 5409 β˜… 0 Jun 4
  • Problem: Existing world models for robot evaluation use open-loop prediction, but VLA policies operate in closed-loop feedback; a world model for policy-in-the-loop evaluation is needed.
  • Model: "PiL-World": chunk-wise world model for closed-loop VLA evaluation using action-derived visual control, latent multi-view history conditioning, and joint multi-view prediction.
  • Code: not released

Spatial Single-Cell Study

Do Foundation Models See Biology? Evaluating Attention Coherence with Spatial Transcriptomics in Glioblastoma

↑ 0 πŸ“š 3679 β˜… 0 Jun 3
  • Problem: Whether attention maps from pathology foundation models capture genuine biological signals remains unknown, hindering clinical trust and regulatory approval.
  • Model: approach: spatial transcriptomics-based framework for objective evaluation of attention coherence in pathology foundation models using attention-based multiple instance learning
  • Code: not released

GC-MoE: Genomics-Guided Cell-Type-Specific Mixture of Experts for Histology-Based Single-Cell Spatial Transcriptomics

↑ 0 πŸ“š 1102 β˜… 0 Jun 1
  • Problem: Predict gene expression for individual cells from histology images and cell locations, accounting for cell-type-dependent expression variability.
  • Model: GC-MoE: Genomics-Guided Cell-Type-Specific Mixture of Experts that routes cells to type-specific experts and incorporates cell-type co-expression priors from scRNA-seq data.
  • Code: not released

Week of

Graph Γ— LLM

Vision-Language Models

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

↑ 117 πŸ“š 674 β˜… 0 May 27
  • Problem: Embodied AI systems are fragmented across task families and robot embodiments, limiting generalization across manipulation, navigation, and diverse platforms.
  • Model: Qwen-VLA: unified vision-language-action model extending Qwen3.5-4B with DiT-based flow-matching action decoder for cross-task, cross-embodiment embodied control.
  • Code: QwenLM/Qwen-VLA

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

↑ 45 πŸ“š 9553 β˜… 0 May 27
  • Problem: Vision-language models achieve high spatial reasoning benchmark scores but may rely on statistical shortcuts rather than structured 3D understanding.
  • Model: approach: Representation-level analysis framework using minimal contrastive pairs to measure spatial axis organization and disentanglement in VLM embeddings, plus SpatialTunnel synthetic benchmark.
  • Code: not released

World Models

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

↑ 50 πŸ“š 36656 β˜… 0 May 27
  • Problem: Converting video diffusion foundation models into real-time interactive world models requires scattered techniques across data, training, and inference pipelines.
  • Model: minWM: full-stack framework converting T2V/TI2V bidirectional diffusion models into camera-controllable few-step autoregressive video world models via Causal Forcing/Forcing++ distillation
  • Code: shengshu-ai/minWM

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

↑ 42 πŸ“š 16919 β˜… 0 May 27
  • Problem: Video diffusion models may perceive temporal direction without understanding causality; existing benchmarks rely on synthetic data with limited real-world generalization.
  • Model: YoCausal: two-level benchmark with Reverse Surprise Index (RSI) and Causality Cognition Index (CCI) metrics for evaluating causal cognition in video diffusion models.
  • Code: genmoai/models

AdaState: Self-Evolving Anchors for Streaming Video Generation

↑ 6 πŸ“š 1390 β˜… 0 May 27
  • Problem: Autoregressive video diffusion models anchor to static first frames, suppressing dynamics and locking scene composition despite natural evolution during generation.
  • Model: "AdaState": replaces static first-frame anchor with adaptive latent state that denoises alongside content at each chunk, evolving with generated scenes via relative time formulation.
  • Code: not released
May 2026 22 papers

Week of

Graph Γ— LLM

S2Aligner: Pair-Efficient and Transferable Pre-Training for Sparse Text-Attributed Graphs

↑ 0 πŸ“š 2080 β˜… 0 May 18
  • Problem: Graph foundation models struggle with sparse text-attributed graphs where node texts are missing, noisy, or uneven, causing unreliable structure-semantics alignment and transfer bias.
  • Model: S2Aligner: sparsity-aware and structure-enhanced LLM-as-Aligner framework that decouples semantic alignment from structural modeling via content-structure factorization and sparsity-aware cross-domain risk balancing.
  • Code: not released

Deep Neural Sheaf Diffusion

↑ 0 πŸ“š 1013 β˜… 0 May 18
  • Problem: Scaling Graph Neural Networks to depth is hindered by representation collapse and vanishing signals in existing sheaf diffusion methods.
  • Model: Deep Neural Sheaf Diffusion (DNSD): sheaf-based GNN replacing sheaf Laplacian with adjacency operator, adding normalization, odd nonlinearities, and gating to maintain informative signals across layers.
  • Code: not released

Vision-Language Models

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

↑ 42 πŸ“š 15499 β˜… 0 May 20
  • Problem: Current multimodal LLMs struggle with audio-visual reasoning because text-based chain-of-thought compresses continuous signals into discrete tokens, losing temporal grounding.
  • Model: LatentOmni: cross-modal reasoning framework interleaving textual reasoning with audio-visual latent states, using feature-level supervision and Omni-Sync Position Embedding for temporal alignment.
  • Code: not released

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

↑ 9 πŸ“š 19747 β˜… 0 May 19
  • Problem: Discrete autoregressive text-to-image models suffer from latent covariate shift when optimizing only the policy with a frozen decoder, causing alignment-fidelity trade-offs.
  • Model: RankE: End-to-end post-training framework for discrete text-to-image generation that co-evolves the AR policy and VQ decoder through alternating ranking-based optimization.
  • Code: not released

World Models

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

↑ 19 πŸ“š 13800 β˜… 0 May 20
  • Problem: Quantizing autoregressive video diffusion models is unexplored; standard quantization schemes designed for bidirectional diffusion transformers perform suboptimally on ARVDs.
  • Model: Q-ARVD: quantization framework for autoregressive video diffusion models using final-quality-guided frame-weighting and outlier-aware adaptive dual-scale quantization
  • Code: not released

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

↑ 7 πŸ“š 41260 β˜… 0 May 21
  • Problem: Current agentic LLMs lack control over when and how to plan, causing inefficient token use without reliable accuracy gains.
  • Model: "SRΒ²AM" (Self-Regulated Simulative Reasoning Agentic LLM): decomposes decision-making into three systemsβ€”simulative reasoning via world model, self-regulation via learned configurator, and reactive executionβ€”implemented as distinct chain-of-thought stages within an LLM.
  • Code: sailing-lab/sr2am

Spatial Single-Cell Study

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

↑ 0 πŸ“š 20431 β˜… 0 May 19
  • Problem: Multi-agent workflow design in open-ended scientific settings lacks curated training sets, reliable metrics, and standardized interfaces between tools and agents.
  • Model: AgentCo-op: retrieval-based synthesis framework that composes reusable skills, tools, and external agents into executable workflows through typed artifact handoffs and bounded evidence-guided local repair.
  • Code: ma-compbio-lab/AgentCo-Op

Week of

Graph Γ— LLM

Vision-Language Models

MMSkills: Towards Multimodal Skills for General Visual Agents

↑ 99 πŸ“š 30599 β˜… 0 May 13
  • Problem: Visual agents need reusable multimodal procedural knowledge that binds actions to visual state recognition and decision-making, beyond text-only skills.
  • Model: "MMSkills: framework for representing, generating, and utilizing reusable multimodal procedures for visual agents. Each skill couples textual procedures with runtime state cards and multi-view keyframes, generated via trajectory-to-skill generator and consulted via branch loading.
  • Code: DeepExperience/MMSkills

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

↑ 71 πŸ“š 3089 β˜… 0 May 13
  • Problem: No benchmark systematically compares long-context LVLMs and memory-augmented agents on multimodal multi-session conversations requiring visual evidence.
  • Model: "MemLens": benchmark with 789 questions across five memory abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge update, answer refusal) at four context lengths (32K-256K tokens)
  • Code: xrenaf/MEMLENS

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory

↑ 58 πŸ“š 1421 β˜… 0 May 13
  • Problem: Existing multimodal agent memory evaluations fail to assess whether agents preserve fine-grained visual evidence needed for reasoning over time.
  • Model: "MemEye: a visual-centric evaluation framework measuring visual evidence granularity and reasoning complexity in multimodal agent memory"
  • Code: not released
  • ⚠ Interested, but agent could not fetch the PDF β€” summary based on abstract only.

World Models

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

↑ 87 πŸ“š 36802 β˜… 0 May 14
  • Problem: Existing AR diffusion distillation methods require 4+ sampling steps; frame-wise 1–2 step generation needs efficient, scalable student initialization.
  • Model: "Causal Forcing++": AR diffusion distillation pipeline using causal consistency distillation for few-step student initialization, avoiding expensive full PF-ODE trajectory precomputation.
  • Code: thu-ml/Causal-Forcing

Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video

↑ 38 πŸ“š 18942 β˜… 0 May 13
  • Problem: Existing camera-controlled video generation methods require large-scale camera-annotated data or expensive test-time optimization; no simple way to leverage pretrained video models' latent camera-control capability.
  • Model: "Warp-as-History": converts camera-induced geometric warps into camera-warped pseudo-history fed through pretrained video models' native history pathway, with target-frame positional alignment and visible-token selection.
  • Code: yyfz/Warp-as-History

PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

↑ 7 πŸ“š 19915 β˜… 0 May 13
  • Problem: Existing video rewards fail to reliably score human motion realism because they operate in 2D pixel space without explicitly modeling 3D body physics and constraints.
  • Model: "PhyMotion": physics-grounded motion reward that recovers SMPL meshes, retargets to MuJoCo simulator, and evaluates motion via three axes (kinematic plausibility, contact/balance, dynamic feasibility)
  • Code: not released

Spatial Single-Cell Study

DUET: Dual-Paradigm Adaptive Expert Triage with Single-cell Inductive Prior for Spatial Transcriptomics Prediction

↑ 0 πŸ“š 5699 β˜… 0 May 13
  • Problem: Existing methods for inferring spatial gene expression from histology images oversimplify morphology-to-expression mapping and underutilize large-scale single-cell data as biological constraints.
  • Model: DUET: dual-paradigm framework synergizing parametric regression and memory-based retrieval with cellular inductive priors and adaptive expert triage for spatial transcriptomics prediction.
  • Code: Junchao-Zhu/DUET

StateXDiff: Cell State-Contextualized Multimodal Diffusion for Single-Cell Perturbation Prediction

↑ 0 β˜… 0 May 15
  • Problem: Predicting drug-induced cellular state changes at single-cell resolution under out-of-distribution conditions with limited multimodal information.
  • Model: StateXDiff: cell State-contextualized multimodal Diffusion framework integrating transcriptomic and pseudo-protein representations with mechanism-aware drug templates via latent conditional diffusion.
  • Code: not released