Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Lar...
Fan XU, Luis A. Leiva
Action editor: Xiao Luo

Training-Free Pseudo-Fusion for Composed Image Retrieval with Diffusion Models and Multimodal Lar...
Fan XU, Luis A. Leiva
Action editor: Xiao Luo
🤖 Cloudflare Unveils Multimodal Decision Model Clef-omni
Cloudflare's announcement expands its open weight decision model family by adding audio and video input alongside text and images, inside a unified sequence with video and audio...
Simple is Better than Complex: A Representation-centric Perspective for Prompting-based Vision--L...
Yujia Yin, Jinhong Ni, Renjie Wu, HONGJI LI, Tianxin Wei, Zhong Li, Yifan Chen
Action editor: Jaeho Lee
🤖 Sharpa's D01 Robot Adds Touch to Its World Model
The description is a catalog of specifications and then a single sentence underneath: the robot's arm is built to carry its own weight, it can move at 10.5 metres per second, and it...
Didn't log #WeekWithoutDriving activities on the daily while I'm on vacation. I've been car & air passenger, rode bus & Montréal Metro, walked pedestrianized streets in Montréal. I get to see so much more at human speed! Public art, shops, people enjoying life. Zero parking hassle/cost. #multimodal
Discrete Diffusion in Large Language and Multimodal Models: A Survey
Runpeng Yu, Qi Li, Xinchao Wang
Action editor: Shuangfei Zhai
EmbeddingGemma 2:テキスト・画像・音声・動画を統合するオープンな軽量マルチモーダル埋め込みモデル
Google DeepMindが発表した「EmbeddingGemma 2」は、テキスト・画像・音声・動画を単一の空間にマッピングする7.4億パラメータのオープンな軽量マルチモーダル埋め込みモデルです。
🤖 Google's Compact EmbeddingGemma 2 Model Outperforms Larger Rivals
EmbeddingGemma 2 puts 740 million parameters into a model that converts text, images, video, audio and code into vectors, and claims it is the...
#RAGEmbeddings #Multimodal #BenchmarksEvaluation #AI #AIPulse
1/ #OpenCID : Ensuring total sovereignty for research data!
🔹 Local import & #multimodal management 🔹 #Python modularity (easy new formats) 🔹 100% Web-based & remote access 🔹 Auto-metadata extraction ( #FAIR / #OpenData) 🔹 Background imports for seamless workflow ⚙️
New #TMLR-Paper-with-Video:
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
Tobias Jan Wieczorek, Nathalie Daun, Mohammad Emtiyaz Khan, Marcus Rohrbach
Google DeepMind Releases EmbeddingGemma 2, a Local Multimodal Embedding for Text, Image, Audio and Video
🤖 Reka's Rho-1 Model Unifies Multimodal Processing
The argument is a practical one about how multimodal systems work. Most current systems are pipelines, where a central model plans and hands work off to specialists for images, video or...
Just saw Reka AI’s new Rho‑1 model—one neural net that juggles text, images, video and even robot control in a shared context window. Think inverse dynamics meets video generation. This could change multimodal AI forever. #RekaAI #Rho1 #Multimodal
🤖 Local Corrections Beat Global Safety Signals in AI Image Generation
The finding is a useful geometry rather than an argument against global safety. A compact unsafe subspace covers little of the diverse unsafe...
#SafetyAlignment #Multimodal #InferenceOptimization #AI #AIPulse
Unifying Understanding and Generation in Vision-Language Models: Advances, Challenges, and Opport...
Xiaocheng Lu, Ziyue Ma, Jie ZHANG, Jian Liu, Song Guo
Action editor: Yu-Xiong Wang
🤖 CellART Advances Single-Cell Analysis in Spatial Transcriptomics
Spatial transcriptomics platforms capture the spatial distribution of gene expression, but they differ in resolution, staining, and the number of genes measured...
"AI Agents on AWS: From Chatbots to Autonomous Digital Workers" by Nitya Angadi
🤖 New Vulnerability Research Model Outperforms Claude Opus at Fraction of Cost
The argument is that a defender needs a model that can be run locally, not an orchestrator, and the benchmark is therefore 60 tasks from 20 held out...
🤖 Proactive AI Agents Change the Game with Strategic Interruptions
Three agents that speak first, one pattern: a personal assistant that books and emails while the app is closed, a reading agent that reaches in chat and teams, and a driver...
🤖 NASA-IBM Lunar Model Unlocks Decades of Orbiter Data for Machine Learning
The dataset is the claim: nearly two million tile bundles from seventeen years of observation, including images from a narrow angle camera at one meter...