Google Gemini 1.5 Pro Critiques a Video Generated by OpenAI Sora

Recently, there have been exciting developments in the field of generative AI from both Google and OpenAI. Google revealed its upgraded multimodal Gemini 1.5 Pro while OpenAI launched Sora, a new text-to-video generation tool. In a fascinating turn of events, Gemini 1.5 Pro analyzed and critiqued a video generated by Sora, pointing out inconsistencies. In […]
MoE-LLaVA Outperforms LLaVA-1.5-7B with Only 3B Parameters

As vision-language models continue growing in size and capabilities, developing more efficient training methods is crucial. The standard approach scales performance by increasing parameters, but this exponentially drives up costs. Researchers at Peking University proposed a novel technique called MoE-LLaVA that can significantly expand model capacity while maintaining constant computation. This model outperforms LLaVA-1.5-7B with […]