India and New Zealand finalize FTA, granting zero-duty access on all goods exports, boosting IT, education, and mobility.
Abstract: Automatic speech recognition (ASR) is the major human–machine interface in many intelligent systems, such as intelligent homes, autonomous driving, and servant robots. However, its ...
In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in ...
SAM Audio is the first unified AI model that can segment sound from complex audio mixtures using text, visual, and time span ...
Abstract: Audio-visual zero-shot learning (ZSL) leverages both video and audio information for model training, aiming to classify new video categories that were not seen during the training. However, ...
This paper will be presented as an oral paper at the ICASSP Audio for Multimedia and Multimodal Processing Session at 6/6/2023 10:50:00 (Eastern European Summer Time). Please cite our paper if you ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results