Paper Tracking

An automated weekly digest. Every Monday an agent scans arxiv and HuggingFace Daily Papers, downloads the PDFs, and summarizes the top 3 picks per topic.

Week of Β· View archive β†’

Graph Γ— LLM

Graph-structured reasoning, graph foundation models, and LLM-augmented GNNs.

Vision-Language Models

Multimodal VLMs, vision encoders, and image-text foundation models.

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

↑ 27 πŸ“š 14588 β˜… 0 Sep 29
  • Problem: Unified multimodal models either quantize images to discrete tokens or use separate objectives for language and vision; a fully continuous shared generative process remains underexplored.
  • Model: "Multimodal Flow (MF-1): fully continuous flow-based generative model using ordered hyperchunks and chunk-causal Flow Matching backbone for joint language-vision modeling in embedding spaces.
  • Code: hustvl/Multimodal-Flow

World Models

Predictive world models, JEPA-style learning, video generation, and embodied simulators.

World Observer: Joint Actor-Observer Generation for Persistent World Modeling

↑ 77 πŸ“š 3981 β˜… 0 Oct 1
  • Problem: Existing video world models fail to maintain object states and dynamics after they leave the agent's view, limiting persistent world modeling.
  • Model: World Observer: decoupled actor-observer generation using a pretrained video Diffusion Transformer that jointly generates perspective actor and panoramic observer streams sharing a single world state.
  • Code: not released

FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation

↑ 19 πŸ“š 86589 β˜… 0 Sep 29
  • Problem: Long-horizon video generation loses coherence when early visual details fall outside the model's input window; existing selection methods rely on current content rather than future information needs.
  • Model: FrameMorrow: prospective frame selector predicting compact prospective tokens representing future information needs to guide selection of relevant historical frames across diverse generators.
  • Code: not released

Spatial Single-Cell Study

Single-cell + spatial omics foundation models, virtual cells, biological world models, and biomedical digital twins.

GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomics

↑ 0 πŸ“š 23935 β˜… 0 Sep 30
  • Problem: Spatial gene expression prediction from histopathology images relies on expensive tests; text-based optimization for incorporating gene semantic information remains underexplored.
  • Model: GATE-ST: Gene-Aware Text-image Encoder for Spatial Transcriptomicsβ€”combines H&E image patches with gene text summaries via cross-attention layers to predict spatial gene expression.
  • Code: not released

Autonomous Labs & Lab-in-the-Loop

Self-driving labs, closed-loop ML-guided experimentation, and agents that run real experimental campaigns.

Experimental Experience Modeling for Autonomous Research

↑ 0 πŸ“š 21750 β˜… 0 Sep 30
  • Problem: Autonomous research agents lack systematic methods to decide which experiments warrant running when evidence is insufficient.
  • Model: Experimental Experience Modeling (EEM): framework for acquiring, reusing, and accumulating experimental experience to guide experimental decisions through historical retrieval, sufficiency assessment, targeted pilots, and iterative library growth.
  • Code: not released

esQueranto: Differentiable Structured Quantum Light for Automated Scientific Discovery

↑ 0 πŸ“š 1470 β˜… 0 Sep 30
  • Problem: Existing quantum-optical simulators lack a unified framework combining photon-number quantum optics with structured-light propagation for automated experimental design.
  • Model: esQueranto: differentiable JAX-based simulator integrating Fock-state quantum optics, spatial beam geometry, and light propagation with automatic differentiation for optimization.
  • Code: artificial-scientist-lab/esQueranto

A roadmap for polymer informatics super-intelligence

↑ 0 πŸ“š 9 β˜… 0 Sep 28
  • Problem: Polymer informatics lacks integrated systems for inverse design, causal reasoning across chemistry-processing-performance, and autonomous closed-loop experimentation.
  • Model: approach: modular agent-directed architecture with conversational LLM layer coordinating domain-specialized tools for neat polymers, composites, solvents, synthesis, and processing across molecular-to-product design hierarchy
  • Code: not released