Usage instructions: here
Table of Contents
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval | Xingtian Pei et.al. | 2605.22478 | null |
| 2026-05-21 | Reduced Dynamical Maps in Finite Temperature Vibronic Coupling Models via Choi Matrices: Numerical Methods and Applications | Raffaele Borrelli et.al. | 2605.22459 | null |
| 2026-05-21 | Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? | Luca Modica et.al. | 2605.22170 | null |
| 2026-05-21 | RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching | Jinhyeok Yang et.al. | 2605.22083 | null |
| 2026-05-21 | Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion | Gang Dai et.al. | 2605.21907 | null |
| 2026-05-20 | Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation | Xuan Du et.al. | 2605.20946 | null |
| 2026-05-20 | Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition | Vinicius Ribeiro et.al. | 2605.20920 | null |
| 2026-05-20 | Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech | Semin Kim et.al. | 2605.20830 | null |
| 2026-05-19 | NeSST: A Python Tool for Neutron Spectra and Synthetic Diagnostics in Inertial Confinement Fusion | Aidan Crilly et.al. | 2605.20432 | null |
| 2026-05-19 | Recombination Thickness as an Uncertainty in Inflationary Observables | V. K. Oikonomou et.al. | 2605.19336 | null |
| 2026-05-18 | Probing SMEFT Operators through |
Amir Subba et.al. | 2605.18382 | null |
| 2026-05-18 | The classical Yangian symmetry of Auxiliary Field Sigma Models | Daniele Bielli et.al. | 2605.18213 | null |
| 2026-05-18 | Bridging the Gap: Converting Read Text to Conversational Dialogue | Parshav Singla et.al. | 2605.18001 | null |
| 2026-05-17 | Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation | Yuheng Chen et.al. | 2605.17488 | null |
| 2026-05-16 | Taming Audio VAEs via Target-KL Regularization | Prem Seetharaman et.al. | 2605.17085 | null |
| 2026-05-16 | SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis | Huimeng Wang et.al. | 2605.16964 | null |
| 2026-05-15 | Linked Multi-Model Data on Russian Domestic and Foreign Policy Speeches | Daria Blinova et.al. | 2605.15886 | null |
| 2026-05-15 | Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction | Hidde Folkertsma et.al. | 2605.15854 | null |
| 2026-05-15 | Global dynamics of a supercritical wave equation in a large data regime | Shijie Dong et.al. | 2605.15662 | null |
| 2026-05-14 | Implicit Dynamical Tensor Train Approximation for Kinetic Equations with Stiff Fokker--Planck Collisions | Geshuo Wang et.al. | 2605.15382 | null |
| 2026-05-14 | From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents | Md Tahmid Rahman Laskar et.al. | 2605.15104 | null |
| 2026-05-14 | Dyonic black holes supporting nearly-black self-gravitating thin shells | Shahar Hod et.al. | 2605.14410 | null |
| 2026-05-14 | Boundary null-controllability for the beam equation with classical structural damping | Sergei Avdonin et.al. | 2605.14371 | null |
| 2026-05-14 | Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment | Yuchen Sun et.al. | 2605.14311 | null |
| 2026-05-13 | Superharmonically Weighted Dirichlet Spaces | H. Bahajji-El Idrissi et.al. | 2605.13787 | null |
| 2026-05-13 | Beyond Explained Variance: A Cautionary Tale of PCA | Gionni Marchetti et.al. | 2605.13520 | null |
| 2026-05-13 | Exploiting Pre-trained Encoder-Decoder Transformers for Sequence-to-Sequence Constituent Parsing | Daniel Fernández-González et.al. | 2605.13373 | null |
| 2026-05-12 | AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling | Yiming Ren et.al. | 2605.11866 | null |
| 2026-05-11 | Exploring Token-Space Manipulation in Latent Audio Tokenizers | Francesco Paissan et.al. | 2605.11192 | null |
| 2026-05-11 | AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling | Jiacheng Shi et.al. | 2605.11098 | null |
| 2026-05-11 | PoDAR: Power-Disentangled Audio Representation for Generative Modeling | Alejandro Luebs et.al. | 2605.10084 | null |
| 2026-05-10 | Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech | Dong Yang et.al. | 2605.09386 | null |
| 2026-05-10 | Test-Time Speculation | Avinash Kumar et.al. | 2605.09329 | null |
| 2026-05-08 | AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care | Preyash Yadav et.al. | 2605.08480 | null |
| 2026-05-08 | LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling | Tong Zheng et.al. | 2605.08083 | link |
| 2026-05-08 | The Cauchy problem for the improved Boussinesq equation with spatially quasi-periodic initial data | Zhiqiang Wan et.al. | 2605.07669 | null |
| 2026-05-07 | Many-body theory predictions of positron binding energies in five-membered heterocycles involving N, O, S and NH substituents | S. K. Gregg et.al. | 2605.06926 | null |
| 2026-05-07 | WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling | Guanrou Yang et.al. | 2605.06407 | null |
| 2026-05-08 | Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM | Wenqian Cui et.al. | 2605.05927 | null |
| 2026-05-09 | X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning | Rixi Xu et.al. | 2605.05611 | null |
| 2026-05-07 | BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models | Sai Babu Patarlapalli et.al. | 2605.05561 | null |
| 2026-05-06 | Primordial Magnetic Fields at Cosmic Dawn: 21-cm Forecasts with HERA and SKA | Keduse Worku et.al. | 2605.05323 | null |
| 2026-05-06 | Lifespan of Classical Solutions to One-Dimensional Quasilinear Wave Equations | Yuusuke Sugiyama et.al. | 2605.04976 | null |
| 2026-05-06 | Nearly universal CMB TT spectrum from pre-inflationary dynamics in a closed universe: KICI scenario, bouncing universe, and emergent universe | Qihong Huang et.al. | 2605.04755 | null |
| 2026-05-06 | VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models | Yukun Chen et.al. | 2605.04613 | null |
| 2026-05-06 | Online Riemannian Gradient Descent for Quantum State Tomography with Matrix Product Operators | Jian-Feng Cai et.al. | 2605.04533 | null |
| 2026-05-06 | Stream-T1: Test-Time Scaling for Streaming Video Generation | Yijing Tu et.al. | 2605.04461 | null |
| 2026-05-01 | LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation | Venkata Pushpak Teja Menta et.al. | 2605.00777 | link |
| 2026-05-01 | Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe | Gaofei Shen et.al. | 2605.00607 | null |
| 2026-04-30 | Deeply virtual pion production through two-loop order | Wen Chen et.al. | 2604.28164 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-20 | MMD-Balls as Credal Sets: A PAC-Bayesian Framework for Epistemic Uncertainty in Test-Time Adaptation | Ahanaf Hasan Ariq et.al. | 2605.21783 | null |
| 2026-05-19 | PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding | Masao et.al. | 2605.20414 | null |
| 2026-05-19 | GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation | Shyma Alhuwaider et.al. | 2605.19890 | null |
| 2026-05-18 | CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation | Gyubin Lee et.al. | 2605.18916 | null |
| 2026-05-18 | WavFlow: Audio Generation in Waveform Space | Feiyan Zhou et.al. | 2605.18749 | null |
| 2026-05-18 | Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models | Yanyun Wang et.al. | 2605.18168 | null |
| 2026-05-18 | Stable Audio 3 | Zach Evans et.al. | 2605.17991 | null |
| 2026-05-17 | Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech | Kaavya Chaparala et.al. | 2605.17652 | null |
| 2026-05-17 | VISTA: Variance-Gated Inter-Sequence Test-Time Adaptation for Multi-Sequence MRI Segmentation | Zhipeng Deng et.al. | 2605.17433 | null |
| 2026-05-17 | Towards Principled Test-Time Adaptation for Time Series Forecasting | Haochun Wang et.al. | 2605.17250 | null |
| 2026-05-16 | Taming Audio VAEs via Target-KL Regularization | Prem Seetharaman et.al. | 2605.17085 | null |
| 2026-05-14 | Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis | Shuyang Cui et.al. | 2605.14555 | null |
| 2026-05-12 | WildRelight: A Real-World Benchmark and Physics-Guided Adaptation for Single-Image Relighting | Lezhong Wang et.al. | 2605.11696 | null |
| 2026-05-12 | Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery | Chi-Nguyen Tran et.al. | 2605.11654 | null |
| 2026-05-13 | TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning | Seongah Kim et.al. | 2605.11572 | null |
| 2026-05-11 | Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs | Konstantinos Soiledis et.al. | 2605.10281 | null |
| 2026-05-12 | jina-embeddings-v5-omni: Geometry-preserving Embeddings via Locked Aligned Towers | Florian Hönicke et.al. | 2605.08384 | null |
| 2026-05-06 | Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization | Mohammed Aman Bhuiyan et.al. | 2605.08214 | null |
| 2026-05-05 | Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models | Wei-Ping Huang et.al. | 2605.08186 | null |
| 2026-05-08 | STEPS: A Temporal Smooth Error Propagation Solver on the Manifolds for Test-Time Adaptation in Time Series Forecasting | Jiaqi Liu et.al. | 2605.08005 | null |
| 2026-05-07 | Optimal Transport Audio Distance with Learned Riemannian Ground Metrics | Wonwoo Jeong et.al. | 2605.05554 | null |
| 2026-05-06 | Empirical Study of Pop and Jazz Mix Ratios for Genre-Adaptive Chord Generation | Jinju Lee et.al. | 2605.04998 | null |
| 2026-05-11 | Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition | Zishu Zhou et.al. | 2605.04617 | null |
| 2026-05-06 | Stage-adaptive audio diffusion modeling | Xuanhao Zhang et.al. | 2605.04547 | null |
| 2026-05-05 | MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model | Jingyao Gong et.al. | 2605.03937 | null |
| 2026-05-05 | GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning | Yujun Li et.al. | 2605.03403 | null |
| 2026-05-05 | FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection | Kaixiang Zhao et.al. | 2605.03294 | null |
| 2026-05-03 | Excited states engineering maximizes singlet generation by triplet fusion in conjugated systems | Alessandra Ronchi et.al. | 2605.01998 | null |
| 2026-05-02 | MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention | Yimeng Zhang et.al. | 2605.01235 | null |
| 2026-05-01 | MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video | Kazuya Tateishi et.al. | 2605.00495 | null |
| 2026-05-01 | Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation | Kuan-Po Huang et.al. | 2605.00329 | null |
| 2026-04-28 | PI-TTA: Physics-Informed Source-Free Test-Time Adaptation for Robust Human Activity Recognition on Mobile Devices | Changyu Li et.al. | 2604.25435 | null |
| 2026-04-26 | AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance | Benjamin Klein et.al. | 2604.23909 | null |
| 2026-04-26 | Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation | Chunyu Li et.al. | 2604.23632 | null |
| 2026-04-24 | UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions | Chunyu Qiang et.al. | 2604.22209 | null |
| 2026-04-23 | Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation | Yingkai Yang et.al. | 2604.21772 | null |
| 2026-04-23 | Prototype-Based Test-Time Adaptation of Vision-Language Models | Zhaohong Huang et.al. | 2604.21360 | null |
| 2026-04-23 | an interpretable vision transformer framework for automated brain tumor classification | Chinedu Emmanuel Mbonu et.al. | 2604.21311 | null |
| 2026-04-21 | Multi-modal Test-time Adaptation via Adaptive Probabilistic Gaussian Calibration | Jinglin Xu et.al. | 2604.19093 | null |
| 2026-04-21 | AeroBridge-TTA: Test-Time Adaptive Language-Conditioned Control for UAVs | Lingxue Lyu et.al. | 2604.19059 | null |
| 2026-04-20 | Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval | HaeJun Yoo et.al. | 2604.18360 | null |
| 2026-04-19 | Dual Strategies for Test-Time Adaptation | Nam Nguyen Phuong et.al. | 2604.17542 | null |
| 2026-04-18 | Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts | Gabriel Jason Lee et.al. | 2604.16926 | null |
| 2026-04-17 | Detecting Alarming Student Verbal Responses using Text and Audio Classifier | Christopher Ormerod et.al. | 2604.16717 | null |
| 2026-04-17 | Cross-Modal Bayesian Low-Rank Adaptation for Uncertainty-Aware Multimodal Learning | Habibeh Naderi et.al. | 2604.16657 | null |
| 2026-04-17 | Beyond One-Size-Fits-All: Adaptive Test-Time Augmentation for Sequential Recommendation | Xibo Li et.al. | 2604.16121 | null |
| 2026-04-17 | Adapting in the Dark: Efficient and Stable Test-Time Adaptation for Black-Box Models | Yunbei Zhang et.al. | 2604.15609 | null |
| 2026-04-16 | ProtoTTA: Prototype-Guided Test-Time Adaptation | Mohammad Mahdi Abootorabi et.al. | 2604.15494 | null |
| 2026-04-16 | ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling | Jianxuan Yang et.al. | 2604.15086 | null |
| 2026-04-16 | VoxSafeBench: Not Just What Is Said, but Who, How, and Where | Yuxiang Wang et.al. | 2604.14548 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | Quantifying Full-Body Immersion | Alihan Bakir et.al. | 2605.22521 | null |
| 2026-05-21 | LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning | Yifan Dai et.al. | 2605.22012 | null |
| 2026-05-21 | Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction | Dinithi Dissanayake et.al. | 2605.21869 | null |
| 2026-05-20 | Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition | Junghyun Lee et.al. | 2605.21417 | null |
| 2026-05-20 | OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization | Mark Boss et.al. | 2605.21226 | null |
| 2026-05-20 | A Survey of Audio Reasoning in Multimodal Foundation Models | Zhihan Guo et.al. | 2605.21008 | null |
| 2026-05-20 | FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching | Jangho Park et.al. | 2605.20910 | null |
| 2026-05-19 | MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation | Yujie Wei et.al. | 2605.20183 | null |
| 2026-05-19 | Stage-adaptive Token Selection for Efficient Omni-modal LLMs | Zijie Xin et.al. | 2605.20035 | null |
| 2026-05-19 | EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection | Aritra Marik et.al. | 2605.19630 | null |
| 2026-05-18 | CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation | Gyubin Lee et.al. | 2605.18916 | null |
| 2026-05-18 | WavFlow: Audio Generation in Waveform Space | Feiyan Zhou et.al. | 2605.18749 | null |
| 2026-05-18 | OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding | Ruixiang Zhao et.al. | 2605.18577 | null |
| 2026-05-18 | Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI | Daiqi Liu et.al. | 2605.18466 | null |
| 2026-05-18 | SIREM: Speech-Informed MRI Reconstruction with Learned Sampling | Md Hasan et.al. | 2605.18221 | null |
| 2026-05-18 | Beyond the Cartesian Illusion: Testing Two-Stage Multi-Modal Theory of Mind under Perceptual Bottlenecks | Yajing Zhou et.al. | 2605.18194 | null |
| 2026-05-17 | Audio-Image Cross-Modal Retrieval with Onomatopoeic Images | Keisuke Imoto et.al. | 2605.17509 | null |
| 2026-05-17 | Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation | Yuheng Chen et.al. | 2605.17488 | null |
| 2026-05-16 | HighSync: High-Quality Lip Synchronization via Latent Diffusion Models | Saeed Firouzi Daghigh et.al. | 2605.16918 | null |
| 2026-05-13 | When Vision Speaks for Sound | Xiaofei Wen et.al. | 2605.16403 | null |
| 2026-05-14 | Sound Sparks Motion: Audio and Text Tuning for Video Editing | AmirHossein Naghi Razlighi et.al. | 2605.15307 | null |
| 2026-05-14 | XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference | Thomas Witt et.al. | 2605.14844 | null |
| 2026-05-15 | IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments | Dinanath Padhya et.al. | 2605.14736 | null |
| 2026-05-12 | SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning | Xin Cheng et.al. | 2605.12179 | null |
| 2026-05-12 | OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models | Yuchen Deng et.al. | 2605.12056 | null |
| 2026-05-13 | Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation | Che Liu et.al. | 2605.12034 | null |
| 2026-05-12 | AuDirector: A Self-Reflective Closed-Loop Framework for Immersive Audio Storytelling | Yiming Ren et.al. | 2605.11866 | null |
| 2026-05-12 | Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs | Chaeyoung Jung et.al. | 2605.11605 | null |
| 2026-05-13 | TB-AVA: Text as a Semantic Bridge for Audio-Visual Parameter Efficient Finetuning | Seongah Kim et.al. | 2605.11572 | null |
| 2026-05-12 | Probing Cross-modal Information Hubs in Audio-Visual LLMs | Jihoo Jung et.al. | 2605.10815 | null |
| 2026-05-11 | FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries | Qijie You et.al. | 2605.10228 | null |
| 2026-05-11 | APEX: Audio Prototype EXplanations for Classification Tasks | Piotr Kawa et.al. | 2605.10153 | null |
| 2026-05-11 | Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization | Yeongtak Oh et.al. | 2605.09996 | null |
| 2026-05-11 | Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought | Xuanchen Li et.al. | 2605.09906 | null |
| 2026-05-11 | ChladniSonify: A Visual-Acoustic Mapping Method for Chladni Patterns in New Media Art Creation | Yakun Liu et.al. | 2605.09846 | null |
| 2026-05-10 | When Sounds Hurt and Voices Aren't Heard: An Experience Report on Misophonia, Sensory Trauma, and Trauma-Informed Design | Tawfiq Ammari et.al. | 2605.09796 | null |
| 2026-05-10 | Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition | Yifan Wang et.al. | 2605.09468 | null |
| 2026-05-10 | PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor Networks | Prasenjit Karmakar et.al. | 2605.09434 | null |
| 2026-05-10 | Towards Conversational Medical AI with Eyes, Ears and a Voice | Meet Shah et.al. | 2605.09272 | null |
| 2026-05-09 | PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment | Qiqi Li et.al. | 2605.08945 | null |
| 2026-05-08 | MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head Generation | Xinyan Ye et.al. | 2605.08050 | null |
| 2026-05-08 | TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos | Hengyi Feng et.al. | 2605.07593 | null |
| 2026-05-08 | Towards multi-modal forgery representation learning for AI-generated video detection and localization | Dat Le et.al. | 2605.07232 | null |
| 2026-05-08 | PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition | Yuchen He et.al. | 2605.07154 | null |
| 2026-05-08 | Do Joint Audio-Video Generation Models Understand Physics? | Zijun Cui et.al. | 2605.07061 | null |
| 2026-05-06 | Making AI Drafts Count: A Quality Threshold in Audio Description Workflows | Lana Do et.al. | 2605.05348 | null |
| 2026-05-08 | How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study | Junran Wang et.al. | 2605.05340 | null |
| 2026-05-06 | To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition | Yangchen Yu et.al. | 2605.04877 | null |
| 2026-05-05 | A foundation model of vision, audition, and language for in-silico neuroscience | Stéphane d'Ascoli et.al. | 2605.04326 | null |
| 2026-05-05 | Audio-Visual Intelligence in Large Foundation Models | You Qin et.al. | 2605.04045 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have) | Hisashi Miyashita et.al. | 2605.22005 | null |
| 2026-05-19 | Contradiction Graphs Determine VC Dimension | Jesse Campbell et.al. | 2605.20434 | null |
| 2026-05-20 | Voice ''Cloning'' is Style Transfer | Kaitlyn Zhou et.al. | 2605.16578 | null |
| 2026-05-14 | Real-time virtual circuits for plasma shape control via neural network emulators | Alasdair Ross et.al. | 2605.14939 | null |
| 2026-05-14 | Strategic PAC Learnability via Geometric Definability | Yuval Filmus et.al. | 2605.13426 | null |
| 2026-05-12 | Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling | Chen Geng et.al. | 2605.12310 | null |
| 2026-05-12 | The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive | Shaina Raza et.al. | 2605.12075 | null |
| 2026-05-11 | Exploring Token-Space Manipulation in Latent Audio Tokenizers | Francesco Paissan et.al. | 2605.11192 | null |
| 2026-05-11 | Charting the Diameter Computation Landscape on Intersection Graphs in the Plane | Timothy M. Chan et.al. | 2605.10692 | null |
| 2026-05-11 | NaiAD: Initiate Data-Driven Research for LLM Advertising | Yihang Zhang et.al. | 2605.09918 | null |
| 2026-05-11 | Some model-theoretic consequences of high-arity uniform convergence, part I | Leonardo N. Coregliano et.al. | 2605.09911 | null |
| 2026-05-10 | Online Set Learning from Precision and Recall Feedback | Lee Cohen et.al. | 2605.09565 | null |
| 2026-05-08 | Approximation Error Upper and Lower Bounds for Hölder Class with Transformers | Xin He et.al. | 2605.07463 | null |
| 2026-05-08 | Regret-Oracle Complexity Tradeoffs in Agnostic Online Learning | Idan Attias et.al. | 2605.07155 | null |
| 2026-05-08 | Every Feedforward Neural Network Definable in an o-Minimal Structure Has Finite Sample Complexity | Anastasis Kratsios et.al. | 2605.07097 | null |
| 2026-05-07 | From Specification to Deployment: Empirical Evidence from a W3C VC + DID Trust Infrastructure for Autonomous Agents | Lars Kersten Kroehl et.al. | 2605.06738 | null |
| 2026-05-06 | Information-theoretic Limits of Learning and Estimation | Abbas El Gamal et.al. | 2605.06710 | null |
| 2026-05-07 | Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML | Jai Moondra et.al. | 2605.06656 | null |
| 2026-05-07 | WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling | Guanrou Yang et.al. | 2605.06407 | null |
| 2026-05-07 | A Fine-Grained Understanding of Uniform Convergence for Halfspaces | Aryeh Kontorovich et.al. | 2605.06004 | null |
| 2026-05-07 | X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning | Rixi Xu et.al. | 2605.05611 | null |
| 2026-05-06 | VC-FeS: Viewpoint-Conditioned Feature Selection for Vehicle Re-identification in Thermal Vision | Yasod Ginige et.al. | 2605.04750 | null |
| 2026-05-05 | Do Venture Capitalists Beat Random Allocation? | Max Sina Knicker et.al. | 2605.03980 | null |
| 2026-05-05 | MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model | Jingyao Gong et.al. | 2605.03937 | null |
| 2026-05-04 | Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings | Vamshi Nallaguntla et.al. | 2605.03079 | null |
| 2026-05-04 | Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion | Amirhosein Javadi et.al. | 2605.02849 | null |
| 2026-05-04 | InsureConnect: Blockchain and Digital Identity for the Property Insurance Market | João Eduardo Travassos et.al. | 2605.02824 | null |
| 2026-05-01 | LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation | Venkata Pushpak Teja Menta et.al. | 2605.00777 | null |
| 2026-05-01 | An Unsupervised Machine Learning-based Framework for Wafer Scale Variability Analysis and Performance Prediction of Ferroelectric Hf0.5Zr0.5O2 Thin Film Capacitors | Anika Anu et.al. | 2605.00544 | null |
| 2026-04-30 | VC-Density in Divisible Oriented Abelian Groups and Their Pairs | Ebru Nayir et.al. | 2604.27816 | null |
| 2026-04-30 | Benchmarking virtual cell models for in-the-wild perturbation response | Xinjie Mao et.al. | 2604.27646 | null |
| 2026-04-30 | JaiTTS: A Thai Voice Cloning Model | Jullajak Karnjanaekarin et.al. | 2604.27607 | null |
| 2026-04-29 | Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes | Akshay Mittal et.al. | 2604.26997 | null |
| 2026-04-29 | The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation | Yun-Shao Tsai et.al. | 2604.26347 | null |
| 2026-04-28 | One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech | Amanuel Gizachew Abebe et.al. | 2604.26136 | null |
| 2026-04-28 | Robust Accent Identification via Voice Conversion and Non-Timbral Embeddings | Rayane Bakari et.al. | 2604.25332 | null |
| 2026-04-28 | AgentDID: Trustless Identity Authentication for AI Agents | Minghui Xu et.al. | 2604.25189 | null |
| 2026-04-28 | On rates of convergence for sample average approximations without smoothness | Hien Duy Nguyen et.al. | 2604.25153 | null |
| 2026-04-27 | Null Measurability at the Symmetrization Interface in VC Learning | Dhruv Gupta et.al. | 2604.25028 | null |
| 2026-04-27 | Subjective Portrait Region Cropping in Landscape Videos with Temporal Annotation Smoothing | Cheng-Han Lee et.al. | 2604.24947 | null |
| 2026-04-27 | The Optimal Sample Complexity of Multiclass and List Learning | Chirag Pabbaraju et.al. | 2604.24749 | null |
| 2026-04-26 | Talking Slide Avatars: Open-Source Multimodal Communication Approach for Teaching | Xinxing Wu et.al. | 2604.23703 | null |
| 2026-04-17 | Audio2Tool: Bridging Spoken Language Understanding and Function Calling | Ramit Pahwa et.al. | 2604.22821 | null |
| 2026-04-24 | Nature of point defects in bulk hexagonal diamond | Ling Zhu et.al. | 2604.22393 | null |
| 2026-04-22 | Critical Activation Voltage for Phonon-Mediated Field-Driven Phenomena | Ric Fulop et.al. | 2604.20089 | null |
| 2026-04-22 | Robust Uniform Recovery of Structured Signals from Nonlinear Observations | Pedro Abdalla et.al. | 2604.20075 | null |
| 2026-04-20 | Parameterized Capacitated Vertex Cover Revisited | Michael Lampis et.al. | 2604.18746 | null |
| 2026-04-20 | Horospherical Depth and Busemann Median on Hadamard Manifolds | Yangdi Jiang et.al. | 2604.18242 | null |
| 2026-04-19 | Homogeneous Network Caching is Fixed-Parameter Tractable Parameterized by the Number of Caches | József Pintér et.al. | 2604.17546 | null |
| 2026-04-19 | Optimal Phylogenetic Reconstruction from Sampled Quartets | Dionysis Arvanitakis et.al. | 2604.17461 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | MotiMotion: Motion-Controlled Video Generation with Visual Reasoning | Lee Hsin-Ying et.al. | 2605.22818 | null |
| 2026-05-21 | VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis | Jinho Park et.al. | 2605.22570 | null |
| 2026-05-21 | Cell Phantom Video Generation in Elliptical Fourier Descriptor Domain | Francesco Benedetto et.al. | 2605.22563 | null |
| 2026-05-21 | Bernini: Latent Semantic Planning for Video Diffusion | Bernini Team et.al. | 2605.22344 | null |
| 2026-05-21 | Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors | Jiahe Chen et.al. | 2605.22272 | null |
| 2026-05-21 | One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems | Yufei Shi et.al. | 2605.22144 | null |
| 2026-05-21 | Video as Natural Augmentation: Towards Unified AI-Generated Image and Video Detection | Zhengcen Li et.al. | 2605.21977 | null |
| 2026-05-20 | StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation | Guanlong Jiao et.al. | 2605.21466 | null |
| 2026-05-20 | Q-ARVD: Quantizing Autoregressive Video Diffusion Models | Siao Tang et.al. | 2605.21072 | null |
| 2026-05-20 | Dynamic Video Generation: Shaping Video Generation Across Time and Space | Shikang Zheng et.al. | 2605.21042 | null |
| 2026-05-20 | DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation | Bo Ye et.al. | 2605.21028 | null |
| 2026-05-20 | FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching | Jangho Park et.al. | 2605.20910 | null |
| 2026-05-20 | What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing | Hangyu Lin et.al. | 2605.20795 | null |
| 2026-05-20 | RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers | Yuxi Liu et.al. | 2605.20659 | null |
| 2026-05-19 | Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation | Matthew Bendel et.al. | 2605.20476 | null |
| 2026-05-19 | Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision | Runyuan Cai et.al. | 2605.20309 | null |
| 2026-05-19 | MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation | Yujie Wei et.al. | 2605.20183 | null |
| 2026-05-19 | CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition | Hongji Yang et.al. | 2605.19995 | null |
| 2026-05-19 | Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls | Abdul Mohaimen Al Radi et.al. | 2605.19728 | null |
| 2026-05-19 | Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention | Wenhu Zhang et.al. | 2605.19726 | null |
| 2026-05-19 | PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning | Qiran Zhang et.al. | 2605.19382 | null |
| 2026-05-19 | SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution | Yiren Song et.al. | 2605.19319 | null |
| 2026-05-19 | PhyWorld: Physics-Faithful World Model for Video Generation | Pu Zhao et.al. | 2605.19242 | null |
| 2026-05-18 | Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos | Yuqi Tang et.al. | 2605.18984 | null |
| 2026-05-18 | Actionable World Representation | Kunqi Xu et.al. | 2605.18743 | null |
| 2026-05-19 | LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation | Yukang Chen et.al. | 2605.18739 | null |
| 2026-05-18 | EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation | Rosario Leonardi et.al. | 2605.18214 | null |
| 2026-05-18 | Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving | Lijun Zhou et.al. | 2605.18137 | null |
| 2026-05-18 | AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training | Yucheng Guo et.al. | 2605.17923 | null |
| 2026-05-18 | Temporal Aware Pruning for Efficient Diffusion-based Video Generation | Sheng Li et.al. | 2605.17837 | null |
| 2026-05-17 | Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation | Yuheng Chen et.al. | 2605.17488 | null |
| 2026-05-17 | Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration | Yiren Song et.al. | 2605.17423 | null |
| 2026-05-17 | SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection | Zixi Wei et.al. | 2605.17311 | null |
| 2026-05-17 | Image-to-Video Diffusion: From Foundations to Open Frontiers | Xianlong Wang et.al. | 2605.17248 | null |
| 2026-05-16 | StreamingEffect: Real-Time Human-Centric Video Effect Generation | Yiren Song et.al. | 2605.17019 | null |
| 2026-05-16 | DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis | Yi Zuo et.al. | 2605.16937 | null |
| 2026-05-14 | EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation | Ruozhen He et.al. | 2605.15199 | null |
| 2026-05-14 | RefDecoder: Enhancing Visual Generation with Conditional Video Decoding | Xiang Fan et.al. | 2605.15196 | null |
| 2026-05-14 | Quantitative Video World Model Evaluation for Geometric-Consistency | Jiaxin Wu et.al. | 2605.15185 | null |
| 2026-05-14 | Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video | Yifan Wang et.al. | 2605.15182 | null |
| 2026-05-14 | Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation | Min Zhao et.al. | 2605.15141 | null |
| 2026-05-14 | DriveCtrl: Conditioned Sim-to-Real Driving Video Generation | Haonan Zhao et.al. | 2605.15116 | null |
| 2026-05-14 | EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration | Wuyang Li et.al. | 2605.15042 | null |
| 2026-05-14 | Compositional Video Generation via Inference-Time Guidance | Ariel Shaulov et.al. | 2605.14988 | null |
| 2026-05-14 | Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image | Ming Qian et.al. | 2605.14984 | null |
| 2026-05-14 | MechVerse: Evaluating Physical Motion Consistency in Video Generation Models | Rahul Jain et.al. | 2605.14843 | null |
| 2026-05-13 | RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data | Harold Haodong Chen et.al. | 2605.13775 | null |
| 2026-05-13 | AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation | Yuchao Gu et.al. | 2605.13724 | null |
| 2026-05-13 | Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation | Jiayu Chen et.al. | 2605.13111 | null |
| 2026-05-13 | CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation | Seonghyun Jin et.al. | 2605.12938 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation | Javad Parsa et.al. | 2605.22743 | null |
| 2026-05-21 | SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers | Javad Rajabi et.al. | 2605.22668 | null |
| 2026-05-21 | From Baseline to Follow-Up: Counterfactual Spine DXA Image Synthesis in UK Biobank Using a Causal Hierarchical Variational Autoencoder | Yilin Zhang et.al. | 2605.22649 | null |
| 2026-05-21 | AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters | Hanjun Luo et.al. | 2605.22645 | null |
| 2026-05-21 | S2ED: From Story to Executable Descriptions for Consistency-Aware Story Illustration | Sijing Yin et.al. | 2605.22448 | null |
| 2026-05-21 | Impact of Atmospheric Turbulence and Pointing Error on Earth Observation | Celia Sánchez-de-Miguel et.al. | 2605.22268 | null |
| 2026-05-21 | Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation | Yilan Gao et.al. | 2605.22060 | null |
| 2026-05-21 | Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness | Hangyeol Lee et.al. | 2605.22011 | null |
| 2026-05-21 | Noise Schedule Design for Diffusion Models: An Optimal Control Perspective | Seo Taek Kong et.al. | 2605.21911 | null |
| 2026-05-20 | UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation | Jiayun Wang et.al. | 2605.21611 | null |
| 2026-05-20 | One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration | Chaoyang Wang et.al. | 2605.21484 | null |
| 2026-05-20 | OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation | Ziye Li et.al. | 2605.21343 | null |
| 2026-05-20 | RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution | Siyong Jian et.al. | 2605.21195 | null |
| 2026-05-20 | Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models | Kesong Li et.al. | 2605.21123 | null |
| 2026-05-20 | GeoDiff-SAR II: 3D-Driven Foundation Diffusion Models for SAR Generation via Decoupled Control | Xuanting Wu et.al. | 2605.21116 | null |
| 2026-05-20 | TextSculptor: Training and Benchmarking Scene Text Editing | Yiheng Lin et.al. | 2605.21090 | null |
| 2026-05-20 | Spatial Gram Alignment for Ultra-High-Resolution Image Synthesis | Jinjin Zhang et.al. | 2605.20808 | null |
| 2026-05-20 | Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction | Hanzhong Guo et.al. | 2605.20807 | null |
| 2026-05-20 | TASTE: A Designer-Annotated Multi-Dimensional Preference Dataset for AI-Generated Graphic Design | Haonan Zhu et.al. | 2605.20731 | null |
| 2026-05-20 | Rethinking Cross-Layer Information Routing in Diffusion Transformers | Chao Xu et.al. | 2605.20708 | null |
| 2026-05-19 | PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset | Haojun Chen et.al. | 2605.20147 | null |
| 2026-05-19 | MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling | Zhiping Yu et.al. | 2605.20090 | null |
| 2026-05-19 | Probability-Conserving Flow Guidance | Parsa Esmati et.al. | 2605.20079 | null |
| 2026-05-19 | A Framework for Evaluating Zero-Shot Image Generation in Concept-based Explainability | Giacomo Astolfi et.al. | 2605.19855 | null |
| 2026-05-19 | CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models | Junhao Li et.al. | 2605.19750 | null |
| 2026-05-19 | FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models | Yi Sun et.al. | 2605.19739 | null |
| 2026-05-19 | Physics-informed simulation framework for realistic sonar image generation and statistical validation | Kamal Basha S et.al. | 2605.19712 | null |
| 2026-05-19 | Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation | Junjie Wang et.al. | 2605.19639 | null |
| 2026-05-19 | Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting | Yue Yu et.al. | 2605.19554 | null |
| 2026-05-19 | Multi-Scale Generative Modeling with Heat Dissipation Flow Matching | Jun Ma et.al. | 2605.19371 | null |
| 2026-05-18 | Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs | Junyu Pan et.al. | 2605.18172 | null |
| 2026-05-18 | Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework | Mengyu Sun et.al. | 2605.18150 | null |
| 2026-05-18 | Generation Navigator: A State-Aware Agentic Framework for Image Generation | Jinming Liu et.al. | 2605.17969 | null |
| 2026-05-18 | Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation | Xiao He et.al. | 2605.17834 | null |
| 2026-05-18 | Content-Style Identification via Differential Independence | Subash Timilsina et.al. | 2605.17827 | null |
| 2026-05-18 | Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation | Baoteng Li et.al. | 2605.17807 | null |
| 2026-05-18 | Divergence-Suppressing Couplings for Rectified Flow | Yimeng Min et.al. | 2605.17733 | null |
| 2026-05-17 | AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment | Kuei-Chun Kao et.al. | 2605.17602 | null |
| 2026-05-17 | A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation | Tseten Sherpa et.al. | 2605.17564 | null |
| 2026-05-17 | Accelerating Redshift-Conditioned Galaxy Image Synthesis with One-step Generative Modeling | Tianyue Yang et.al. | 2605.17546 | null |
| 2026-05-14 | Aligning Latent Geometry for Spherical Flow Matching in Image Generation | Tuna Han Salih Meral et.al. | 2605.15193 | null |
| 2026-05-14 | Does Synthetic Layered Design Data Benefit Layered Design Decomposition? | Kam Man Wu et.al. | 2605.15167 | null |
| 2026-05-14 | HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling | Jonathan Cederlund et.al. | 2605.14877 | null |
| 2026-05-14 | Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning | Hanbo Cheng et.al. | 2605.14876 | null |
| 2026-05-14 | Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation | Mamadou Keita et.al. | 2605.14799 | null |
| 2026-05-14 | TOPOS: High-Fidelity and Efficient Industry-Grade 3D Head Generation | Bojun Xiong et.al. | 2605.14594 | null |
| 2026-05-14 | LiWi: Layering in the Wild | Yu He et.al. | 2605.14552 | null |
| 2026-05-14 | AnyBand-Diff: A Unified Remote Sensing Image Generation and Band Repair Framework with Spectral Priors | Zuopeng Zhao et.al. | 2605.14341 | null |
| 2026-05-14 | InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation | Yang Yue et.al. | 2605.14333 | null |
| 2026-05-14 | D2-CDIG: Controlled Diffusion Remote Sensing Image Generation with Dual Priors of DEM and Cloud-Fog | Zuopeng Zhao et.al. | 2605.14326 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators | Zachary Novack et.al. | 2605.22717 | null |
| 2026-05-20 | Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods | Fang-Chih Hsieh et.al. | 2605.21538 | null |
| 2026-05-20 | Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches | Junyoung Koh et.al. | 2605.21433 | null |
| 2026-05-20 | Musical Attention Transformer: Music Generation Using a Music-Specific Attention Model | Shinnosuke Taksuka et.al. | 2605.21081 | null |
| 2026-05-18 | MusicDET: Zero-Shot AI-Generated Music Detection | Chaolei Han et.al. | 2605.18072 | null |
| 2026-05-17 | S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation | Huakang Chen et.al. | 2605.17414 | null |
| 2026-05-16 | Taming Audio VAEs via Target-KL Regularization | Prem Seetharaman et.al. | 2605.17085 | null |
| 2026-05-15 | ARIA: A Diagnostic Framework for Music Training Data Attribution | Changheon Han et.al. | 2605.16181 | null |
| 2026-05-15 | Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation | Yuqing Cheng et.al. | 2605.15831 | null |
| 2026-05-14 | Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music | Mohammad Hossein Sameti et.al. | 2605.14765 | null |
| 2026-05-13 | Text2Score: Generating Sheet Music From Textual Prompts | Keshav Bhandari et.al. | 2605.13431 | null |
| 2026-05-11 | SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements | Ege Erdem et.al. | 2605.10398 | null |
| 2026-05-11 | Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration | Haowen Li et.al. | 2605.10203 | null |
| 2026-05-03 | TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation | Xiaoda Yang et.al. | 2605.01809 | null |
| 2026-05-03 | Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation | Jiafeng Liu et.al. | 2605.01790 | null |
| 2026-05-02 | MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention | Yimeng Zhang et.al. | 2605.01235 | null |
| 2026-05-01 | CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval | Yawen Qin et.al. | 2605.00824 | null |
| 2026-04-28 | SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton | Xuzheng He et.al. | 2604.25498 | null |
| 2026-04-24 | UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions | Chunyu Qiang et.al. | 2604.22209 | null |
| 2026-04-20 | A novel LSTM music generator based on the fractional time-frequency feature extraction | Li Ya et.al. | 2604.17823 | null |
| 2026-04-22 | Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation | Vaibhavi Lokegaonkar et.al. | 2604.17656 | null |
| 2026-04-10 | MAGE: Modality-Agnostic Music Generation and Editing | Muhammad Usama Saleem et.al. | 2604.09803 | null |
| 2026-04-07 | Anchored Cyclic Generation: A Novel Paradigm for Long-Sequence Symbolic Music Generation | Boyu Cao et.al. | 2604.05343 | null |
| 2026-04-05 | Neurological Plausibility of AI-Generated Music for Commercial Environments: An In-Silico Cortical Investigation Using Wubble and TRIBE v2 | Shaad Sufi et.al. | 2604.04025 | null |
| 2026-04-03 | Composer Vector: Style-steering Symbolic Music Generation in a Latent Space | Xunyi Jiang et.al. | 2604.03333 | null |
| 2026-03-24 | Echoes: A semantically-aligned music deepfake detection dataset | Octavian Pascu et.al. | 2603.23667 | null |
| 2026-03-24 | MuQ-Eval: An Open-Source Per-Sample Quality Metric for AI Music Generation Evaluation | Di Zhu et.al. | 2603.22677 | null |
| 2026-03-22 | Fusing Memory and Attention: A study on LSTM, Transformer and Hybrid Architectures for Symbolic Music Generation | Soudeep Ghoshal et.al. | 2603.21282 | null |
| 2026-03-22 | SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing | Jianyi Chen et.al. | 2603.21073 | null |
| 2026-03-11 | V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation | Yan-Bo Lin et.al. | 2603.11042 | null |
| 2026-03-09 | Designing a Generative AI-Assisted Music Psychotherapy Tool for Deaf and Hard-of-Hearing Individuals | Youjin Choi et.al. | 2603.07963 | null |
| 2026-03-02 | ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models | Xiaoyu Yi et.al. | 2603.01984 | null |
| 2026-03-01 | SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation | Hongrui Wang et.al. | 2603.01101 | null |
| 2026-03-04 | CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction | Yinghao Ma et.al. | 2603.00610 | null |
| 2026-02-28 | Efficient Long-Sequence Diffusion Modeling for Symbolic Music Generation | Jinhan Xu et.al. | 2603.00576 | null |
| 2026-02-27 | SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away | Jiajia Li et.al. | 2602.24071 | null |
| 2026-02-27 | DashengTokenizer: One layer is enough for unified audio understanding and generation | Heinrich Dinkel et.al. | 2602.23765 | null |
| 2026-02-23 | SongEcho: Towards Cover Song Generation via Instance-Adaptive Element-wise Linear Modulation | Sifei Li et.al. | 2602.19976 | null |
| 2026-03-02 | Depth-Structured Music Recurrence: Budgeted Recurrent Attention for Full-Piece Symbolic Music Modeling | Yungang Yi et.al. | 2602.19816 | null |
| 2026-02-19 | MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions | Rebecca Salganik et.al. | 2602.17769 | null |
| 2026-02-19 | Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment | Ivan Rinaldi et.al. | 2602.17599 | null |
| 2026-02-15 | Evaluating Disentangled Representations for Controllable Music Generation | Laura Ibáñez-Martínez et.al. | 2602.10058 | null |
| 2026-02-10 | Stemphonic: All-at-once Flexible Multi-stem Music Generation | Shih-Lun Wu et.al. | 2602.09891 | null |
| 2026-02-05 | Video-based Music Generation | Serkan Sulun et.al. | 2602.07063 | null |
| 2026-02-06 | AI-Generated Music Detection in Broadcast Monitoring | David Lopez-Ayala et.al. | 2602.06823 | null |
| 2026-02-03 | Rethinking Music Captioning with Music Metadata LLMs | Irmak Bukey et.al. | 2602.03023 | null |
| 2026-02-06 | ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation | Junmin Gong et.al. | 2602.00744 | null |
| 2026-01-27 | EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding | Luca Cerovaz et.al. | 2601.17517 | null |
| 2026-01-22 | Pay (Cross) Attention to the Melody: Curriculum Masking for Single-Encoder Melodic Harmonization | Maximos Kaliakatsos-Papakostas et.al. | 2601.16150 | null |
| 2026-01-22 | PF-D2M: A Pose-free Diffusion Model for Universal Dance-to-Music Generation | Jaekwon Im et.al. | 2601.15872 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? | Luca Modica et.al. | 2605.22170 | null |
| 2026-05-19 | Codec-Robust Attacks on Audio LLMs | Jaechul Roh et.al. | 2605.20519 | null |
| 2026-05-19 | Stage-adaptive Token Selection for Efficient Omni-modal LLMs | Zijie Xin et.al. | 2605.20035 | null |
| 2026-05-19 | Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning | Junyi Wang et.al. | 2605.19541 | null |
| 2026-05-18 | SAME: A Semantically-Aligned Music Autoencoder | Julian D. Parker et.al. | 2605.18613 | null |
| 2026-05-16 | Taming Audio VAEs via Target-KL Regularization | Prem Seetharaman et.al. | 2605.17085 | null |
| 2026-05-15 | Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation | Yuqing Cheng et.al. | 2605.15831 | null |
| 2026-05-12 | OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models | Yuchen Deng et.al. | 2605.12056 | null |
| 2026-05-11 | Exploring Token-Space Manipulation in Latent Audio Tokenizers | Francesco Paissan et.al. | 2605.11192 | null |
| 2026-05-11 | AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling | Jiacheng Shi et.al. | 2605.11098 | null |
| 2026-05-11 | Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs | Konstantinos Soiledis et.al. | 2605.10281 | null |
| 2026-05-11 | PoDAR: Power-Disentangled Audio Representation for Generative Modeling | Alejandro Luebs et.al. | 2605.10084 | null |
| 2026-05-07 | VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing | Jiacheng Xu et.al. | 2605.06765 | null |
| 2026-05-07 | PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization | Adhiraj Banerjee et.al. | 2605.06582 | null |
| 2026-05-06 | Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization | Zheng Fang et.al. | 2605.04700 | null |
| 2026-05-05 | Assessing the Impact of Noise and Speech Enhancement on the Intelligibility of Speech Codecs | Lyonel Behringer et.al. | 2605.03776 | null |
| 2026-04-29 | SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding | Mingyu Zhao et.al. | 2604.26296 | null |
| 2026-04-26 | HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models | Peize He et.al. | 2604.23717 | null |
| 2026-04-22 | Sema: Semantic Transport for Real-Time Multimodal Agents | Jiaying Meng et.al. | 2604.20940 | null |
| 2026-04-22 | ATIR: Towards Audio-Text Interleaved Contextual Retrieval | Tong Zhao et.al. | 2604.20267 | link |
| 2026-04-21 | Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages | Girish et.al. | 2604.19949 | null |
| 2026-04-20 | LLM-Codec: Neural Audio Codec Meets Language Model Objectives | Ho-Lam Chung et.al. | 2604.17852 | link |
| 2026-04-20 | ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics | Heewon Oh et.al. | 2604.16254 | null |
| 2026-04-17 | Hierarchical Codec Diffusion for Video-to-Speech Generation | Jiaxin Ye et.al. | 2604.15923 | link |
| 2026-04-21 | Qwen3.5-Omni Technical Report | Qwen Team et.al. | 2604.15804 | null |
| 2026-04-19 | ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning | Junyi Wang et.al. | 2604.14654 | null |
| 2026-04-16 | Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection | Meng Chen et.al. | 2604.14604 | null |
| 2026-04-14 | An Ultra-Low Latency, End-to-End Streaming Speech Synthesis Architecture via Block-Wise Generation and Depth-Wise Codec Decoding | Tianhui Su et.al. | 2604.12438 | null |
| 2026-04-14 | TokenSE: a Mamba-based discrete token speech enhancement framework for cochlear implants | Hsin-Tien Chiang et.al. | 2604.12246 | null |
| 2026-04-13 | Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization | Xiangyu Zhang et.al. | 2604.12145 | null |
| 2026-04-13 | Efficient Training for Cross-lingual Speech Language Models | Yan Zhou et.al. | 2604.11096 | null |
| 2026-04-13 | LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation | Qi Wang et.al. | 2604.11052 | null |
| 2026-04-10 | Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models | Pengbo Lyu et.al. | 2604.09371 | null |
| 2026-04-08 | Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models | Bajian Xiang et.al. | 2604.06871 | null |
| 2026-03-31 | Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative | Ksenia Lysikova et.al. | 2604.02374 | null |
| 2026-03-27 | LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling | Xiaoyu Fan et.al. | 2603.26364 | null |
| 2026-03-27 | findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding | Héctor Javier Vázquez Martínez et.al. | 2603.26292 | null |
| 2026-03-27 | Distilling Conversations: Abstract Compression of Conversational Audio Context for LLM-based ASR | Shashi Kumar et.al. | 2603.26246 | null |
| 2026-04-06 | Voxtral TTS | Mistral-AI et.al. | 2603.25551 | null |
| 2026-03-21 | AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former | Liyun Zhang et.al. | 2603.20894 | null |
| 2026-03-21 | OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement | Jingbin Hu et.al. | 2603.20638 | null |
| 2026-03-19 | Listen First, Then Answer: Timestamp-Grounded Speech Reasoning | Jihoon Jeong et.al. | 2603.19468 | null |
| 2026-03-18 | Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information | Shih-Heng Wang et.al. | 2603.18359 | null |
| 2026-03-20 | MOSS-TTS Technical Report | Yitian Gong et.al. | 2603.18090 | null |
| 2026-03-10 | Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection | Jinyang Wu et.al. | 2603.16914 | null |
| 2026-03-17 | On the Emotion Understanding of Synthesized Speech | Yuan Ge et.al. | 2603.16483 | null |
| 2026-03-15 | CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents | Wen-Chin Huang et.al. | 2603.14328 | null |
| 2026-03-15 | Controllable Accent Normalization via Discrete Diffusion | Qibing Bai et.al. | 2603.14275 | null |
| 2026-03-14 | Probing neural audio codecs for distinctions among English nuclear tunes | Juan Pablo Vigneaux et.al. | 2603.14035 | null |
| 2026-03-12 | TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling | Liang-Hsuan Tseng et.al. | 2603.12350 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning | Yifan Dai et.al. | 2605.22012 | null |
| 2026-05-20 | Study of flutter instability using the actuator line method for wind energy harvesting devices | Vitor G. Kleine et.al. | 2605.21596 | null |
| 2026-05-20 | PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects | Ziang Cao et.al. | 2605.21572 | null |
| 2026-05-19 | PlanRAG-Audio: Planning and Retrieval Augmented Generation for Long-form Audio Understanding | Masao et.al. | 2605.20414 | null |
| 2026-05-18 | A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook | Kaiwen Luo et.al. | 2605.20266 | null |
| 2026-05-19 | Stage-adaptive Token Selection for Efficient Omni-modal LLMs | Zijie Xin et.al. | 2605.20035 | null |
| 2026-05-19 | AffectVerse: Emotional World Models for Multimodal Affective Computing | Bo Zhao et.al. | 2605.19950 | null |
| 2026-05-19 | Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation | Zhifei Xie et.al. | 2605.19833 | null |
| 2026-05-19 | OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond | Zunhai Su et.al. | 2605.19660 | null |
| 2026-05-18 | OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding | Ruixiang Zhao et.al. | 2605.18577 | null |
| 2026-05-18 | Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs | Junyu Pan et.al. | 2605.18172 | null |
| 2026-05-18 | Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models | Yanyun Wang et.al. | 2605.18168 | null |
| 2026-05-18 | OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models | Morunliu Yang et.al. | 2605.18041 | null |
| 2026-05-17 | Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation | Yuheng Chen et.al. | 2605.17488 | null |
| 2026-05-17 | Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades | Donghyuk Jung et.al. | 2605.17443 | null |
| 2026-05-17 | S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation | Huakang Chen et.al. | 2605.17414 | null |
| 2026-05-17 | Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction | Chaoqun He et.al. | 2605.17360 | null |
| 2026-05-17 | Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities | Heejoon Koo et.al. | 2605.17225 | null |
| 2026-05-16 | A preconditioned augmented Lagrangian method for solving semidefinite programming problems | Tianyun Tang et.al. | 2605.17089 | null |
| 2026-05-16 | TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents | Zhiqiang Liu et.al. | 2605.16909 | null |
| 2026-05-14 | From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents | Md Tahmid Rahman Laskar et.al. | 2605.15104 | null |
| 2026-05-14 | OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance | Yeo Jeong Park et.al. | 2605.14458 | null |
| 2026-05-13 | NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating | Zhongju Yuan et.al. | 2605.13651 | null |
| 2026-05-13 | Leveraging Multimodal Self-Consistency Reasoning in Coding Motivational Interviewing for Alcohol Use Reduction | Guangzeng Han et.al. | 2605.12987 | null |
| 2026-05-12 | OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation | Guohui Zhang et.al. | 2605.12480 | null |
| 2026-05-12 | Benchmarking and Resource Analysis for Augmented-Lagrangian Quantum Hamiltonian Descent | Zeguan Wu et.al. | 2605.12066 | null |
| 2026-05-12 | OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models | Yuchen Deng et.al. | 2605.12056 | null |
| 2026-05-13 | Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation | Che Liu et.al. | 2605.12034 | null |
| 2026-05-12 | Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs | Chaeyoung Jung et.al. | 2605.11605 | null |
| 2026-05-11 | Qwen-Image-2.0 Technical Report | Bing Zhao et.al. | 2605.10730 | null |
| 2026-05-11 | Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization | Yeongtak Oh et.al. | 2605.09996 | null |
| 2026-05-12 | Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse | Kuan Zhang et.al. | 2605.09965 | null |
| 2026-05-09 | Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology | Jucheng Hu et.al. | 2605.09152 | null |
| 2026-05-09 | MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production | Chunyu Xue et.al. | 2605.08962 | null |
| 2026-05-09 | ALM-MTA:Front-Door Causal Multi-Touch Attribution Method for Creator-Ecosystem Optimization | Yuguang Liu et.al. | 2605.08881 | null |
| 2026-05-09 | Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search | Tao Yu et.al. | 2605.08762 | null |
| 2026-05-09 | Omni-scale Learning-based Sequential Decision Framework for Order Fulfillment of Tote-handling Robotic Systems | Jiaxin Liu et.al. | 2605.08758 | null |
| 2026-05-08 | jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition | Florian Hönicke et.al. | 2605.08384 | null |
| 2026-05-08 | Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators | Tsuyoshi Okita et.al. | 2605.07600 | null |
| 2026-05-08 | TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos | Hengyi Feng et.al. | 2605.07593 | link |
| 2026-05-07 | FastOmniTMAE: Parallel Clause Learning for Scalable and Hardware-Efficient Tsetlin Embeddings | Ahmed K. Kadhim et.al. | 2605.06982 | null |
| 2026-05-07 | Task-Aware Answer Preservation under Audio Compression for Large Audio Language Models | Amir Ivry et.al. | 2605.06631 | null |
| 2026-05-07 | SNAPO: Smooth Neural Adjoint Policy Optimization for Optimal Control via Differentiable Simulation | Dmitri Goloubentsev et.al. | 2605.06570 | null |
| 2026-05-07 | X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction | Xiaoming Ren et.al. | 2605.05765 | null |
| 2026-05-06 | GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy | Bogdan Minko et.al. | 2605.05277 | null |
| 2026-05-06 | Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization | Zheng Fang et.al. | 2605.04700 | null |
| 2026-05-06 | VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models | Yukun Chen et.al. | 2605.04613 | null |
| 2026-05-05 | MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model | Jingyao Gong et.al. | 2605.03937 | null |
| 2026-05-04 | Closed-form Model for Radiation Pattern of Pinching Antennas | Muhammad Zubair et.al. | 2605.02578 | null |
| 2026-05-02 | Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection | Tianxiao Li et.al. | 2605.01638 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-18 | Stable Audio 3 | Zach Evans et.al. | 2605.17991 | null |
| 2026-05-13 | When Vision Speaks for Sound | Xiaofei Wen et.al. | 2605.16403 | null |
| 2026-05-11 | Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration | Haowen Li et.al. | 2605.10203 | null |
| 2026-04-18 | Anonymization, Not Elimination: Utility-Preserved Speech Anonymization | Yunchong Xiao et.al. | 2604.17000 | null |
| 2026-04-17 | AST: Adaptive, Seamless, and Training-Free Precise Speech Editing | Sihan Lv et.al. | 2604.16056 | null |
| 2026-04-13 | StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection | Zhentao Liu et.al. | 2604.11917 | null |
| 2026-04-26 | Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing | Zeyue Tian et.al. | 2604.10708 | null |
| 2026-03-29 | VoxAnchor: Grounding Speech Authenticity in Throat Vibration via mmWave Radar | Mingda Han et.al. | 2603.27562 | null |
| 2026-03-27 | LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling | Xiaoyu Fan et.al. | 2603.26364 | null |
| 2026-03-15 | Localizing and Editing Knowledge in Large Audio-Language Models | Sung Kyun Chung et.al. | 2603.14343 | null |
| 2026-02-04 | Audio ControlNet for Fine-Grained Audio Generation and Editing | Haina Zhu et.al. | 2602.04680 | null |
| 2026-01-31 | Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards | Yong Ren et.al. | 2602.00560 | null |
| 2026-04-08 | Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs | Jun Xue et.al. | 2601.21463 | null |
| 2026-01-18 | A Unified Neural Codec Language Model for Selective Editable Text to Speech Generation | Hanchen Pei et.al. | 2601.12480 | null |
| 2026-01-08 | CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models | Junyang Chen et.al. | 2601.05329 | null |
| 2026-01-04 | LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models | Zhiyuan Zhao et.al. | 2601.04233 | null |
| 2025-12-29 | MiMo-Audio: Audio Language Models are Few-Shot Learners | Xiaomi LLM-Core Team et.al. | 2512.23808 | null |
| 2026-01-19 | MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model | Ye Tao et.al. | 2512.20339 | null |
| 2025-12-23 | QuarkAudio Technical Report | Chengwei Liu et.al. | 2512.20151 | null |
| 2025-12-16 | MuseCPBench: an Empirical Study of Music Editing Methods through Music Context Preservation | Yash Vishe et.al. | 2512.14629 | null |
| 2025-12-08 | Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits | Masato Ishii et.al. | 2512.07209 | null |
| 2025-11-15 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | Zhisheng Zheng et.al. | 2511.12347 | null |
| 2025-11-18 | Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models | Yi Yang et.al. | 2511.08252 | null |
| 2025-11-18 | MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers | Ali Boudaghi et.al. | 2511.04376 | null |
| Publish Date | Title | Authors | Code | |
|---|---|---|---|---|
| 2026-05-21 | DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders | Tianhang Wang et.al. | 2605.22777 | null |
| 2026-05-21 | AesFormer: Transform Everyday Photos into Beautiful Memories | Tianxiang Du et.al. | 2605.22126 | null |
| 2026-05-20 | Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning | Dian Zheng et.al. | 2605.21487 | null |
| 2026-05-20 | StreamGVE: Training-Free Video Editing via Few-Step Streaming Video Generation | Guanlong Jiao et.al. | 2605.21466 | null |
| 2026-05-20 | Semantic Granularity Navigation in Image Editing | Liangsi Lu et.al. | 2605.21190 | null |
| 2026-05-20 | TextSculptor: Training and Benchmarking Scene Text Editing | Yiheng Lin et.al. | 2605.21090 | null |
| 2026-05-20 | Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning | Zhangchi Hu et.al. | 2605.20961 | null |
| 2026-05-20 | What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing | Hangyu Lin et.al. | 2605.20795 | null |
| 2026-05-20 | Conflict-Aware Additive Guidance for Flow Models under Compositional Rewards | Xuehui Yu et.al. | 2605.20758 | null |
| 2026-05-19 | Multi-axis Analysis of Image Manipulation Localization | Keanu Nichols et.al. | 2605.20174 | null |
| 2026-05-19 | Are Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing | Xiaodong Wu et.al. | 2605.19511 | null |
| 2026-05-19 | SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution | Yiren Song et.al. | 2605.19319 | null |
| 2026-05-18 | Aurora: Unified Video Editing with a Tool-Using Agent | Yongsheng Yu et.al. | 2605.18748 | null |
| 2026-05-18 | InstructAV2AV: Instruction-Guided Audio-Video Joint Editing | Haojie Zheng et.al. | 2605.18467 | null |
| 2026-05-17 | HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing | Yuyao Zhang et.al. | 2605.17294 | null |
| 2026-05-16 | StreamingEffect: Real-Time Human-Centric Video Effect Generation | Yiren Song et.al. | 2605.17019 | null |
| 2026-05-16 | Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing | Shaodong Xu et.al. | 2605.16951 | null |
| 2026-05-15 | VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation | Yan Luo et.al. | 2605.15661 | null |
| 2026-05-15 | Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance | Song Wu et.al. | 2605.15533 | null |
| 2026-05-14 | Sound Sparks Motion: Audio and Text Tuning for Video Editing | AmirHossein Naghi Razlighi et.al. | 2605.15307 | null |
| 2026-05-14 | RefDecoder: Enhancing Visual Generation with Conditional Video Decoding | Xiang Fan et.al. | 2605.15196 | null |
| 2026-05-14 | From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing | Anirudh Sundara Rajan et.al. | 2605.15181 | null |
| 2026-05-14 | ACE-LoRA: Adaptive Orthogonal Decoupling for Continual Image Editing | Yuehao Liu et.al. | 2605.14948 | null |
| 2026-05-14 | SEDiT: Mask-Free Video Subtitle Erasure via One-step Diffusion Transformer | Zheng Hui et.al. | 2605.14894 | null |
| 2026-05-14 | Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis | Mor Ventura et.al. | 2605.14842 | null |
| 2026-05-14 | MiVE: Multiscale Vision-language features for reference-guided video Editing | Tong Wang et.al. | 2605.14664 | null |
| 2026-05-13 | Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints | Haoyang Hu et.al. | 2605.13349 | null |
| 2026-05-13 | Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation | Jingxuan He et.al. | 2605.13122 | null |
| 2026-05-13 | Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling | Xuehai Bai et.al. | 2605.13062 | null |
| 2026-05-12 | FRAME: Forensic Routing and Adaptive Multi-path Evidence Fusion for Image Manipulation Detection | Kaixiang Zhao et.al. | 2605.12826 | null |
| 2026-05-12 | Inline Critic Steers Image Editing | Weitai Kang et.al. | 2605.12724 | null |
| 2026-05-12 | Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation | Yabo Zhang et.al. | 2605.12305 | null |
| 2026-05-11 | FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry | Elias B. Krey et.al. | 2605.11203 | null |
| 2026-05-11 | Count Anything at Any Granularity | Chang Liu et.al. | 2605.10887 | null |
| 2026-05-11 | Masked Generative Transformer Is What You Need for Image Editing | Wei Chow et.al. | 2605.10859 | null |
| 2026-05-11 | Qwen-Image-2.0 Technical Report | Bing Zhao et.al. | 2605.10730 | null |
| 2026-05-10 | Towards Robust Sequential Decomposition for Complex Image Editing | Zilai Zeng et.al. | 2605.09233 | null |
| 2026-05-09 | Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination | Adrian Azzarelli et.al. | 2605.09024 | null |
| 2026-05-09 | FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence | Xinyu Yan et.al. | 2605.08820 | null |
| 2026-05-09 | RewardHarness: Self-Evolving Agentic Post-Training | Yuxuan Zhang et.al. | 2605.08703 | null |
| 2026-05-09 | EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics | Van-Loc Nguyen et.al. | 2605.08695 | null |
| 2026-05-08 | Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria | Juanxi Tian et.al. | 2605.08354 | null |
| 2026-05-08 | Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision | Jiacheng Chen et.al. | 2605.07940 | null |
| 2026-05-08 | BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing | Peilin Xiong et.al. | 2605.07846 | null |
| 2026-05-08 | SIMI: Self-information Mining Network for Low-light Image Enhancement | Xuanshuo Fu et.al. | 2605.07767 | null |
| 2026-05-08 | OphEdit: Training-Free Text-Guided Editing of Ophthalmic Surgical Videos | Ritul Jangir et.al. | 2605.07695 | null |
| 2026-05-08 | ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning | Honghua Chen et.al. | 2605.07477 | null |
| 2026-05-08 | EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement | Zitong Xu et.al. | 2605.07457 | null |
| 2026-05-08 | EditTransfer++: Toward Faithful and Efficient Visual-Prompt-Guided Image Editing | Lan Chen et.al. | 2605.07455 | null |
| 2026-05-08 | InsHuman: Towards Natural and Identity-Preserving Human Insertion | Jie Li et.al. | 2605.07402 | null |