Ref-AVS

Ref-AVS seeks to segment objects within the visual domain based on expressions containing multimodal cues.