Hi,
Thanks for sharing this work! I ran a simple benchmarking on multiple CFP classification datasets, and found that RET-CLIP beats some other fundus photography foundation models (e.g. EyeCLIP, RetiZero, etc) in terms of classification accuracy. Congrats on that!
Apart from that, I am wondering if RET-CLIP still leads in multimodal tasks like VQA (since it is trained in a CLIP-like manner). Would be grateful if there are some pilot validation results that can share with us!
Hi,
Thanks for sharing this work! I ran a simple benchmarking on multiple CFP classification datasets, and found that RET-CLIP beats some other fundus photography foundation models (e.g. EyeCLIP, RetiZero, etc) in terms of classification accuracy. Congrats on that!
Apart from that, I am wondering if RET-CLIP still leads in multimodal tasks like VQA (since it is trained in a CLIP-like manner). Would be grateful if there are some pilot validation results that can share with us!