In the ever-evolving landscape of Artificial Intelligence, the ability to generate high-quality videos using text, images, or video clips is fast becoming a pivotal tool for creative expression. Today, we dive into an in-depth comparison between two trailblazers in this realm: Runway Gen-2, a revolutionary multi-modal AI system from Runway ML, and Potat1, the pioneering open-source text-to-video model.
While both present their own special features, they are distinct in many important aspects. These differences show up notably in areas like their price, the quality they offer, and how simple they are to use. Let’s unwrap the features, benefits, and limitations of each, offering insights that can help you make an informed choice between these innovative tools. Welcome to our exploration of the future of video generation.

Table of contents
Runway Gen-2
Runway ML has developed Runway Gen-2, a multi-use AI system, that creates unique videos from text, images, or video clips. This system is expertly built to produce realistic and steady new videos. First off, it can take the look and style from a picture or some written words and use them in an existing video. This is called Video to Video transformation. But that’s not all. It can also create a video just from words you type, without needing an existing video. This is known as Text to Video transformation.
Modes
Gen-2 comes with multiple modes:
- Speaking of Text to Video, this mode lets you create videos in just about any style you fancy. And the best part? You only need a simple text prompt to bring your vision to life.
- Text + Image to Video: This mode generates a video using a driving image and a text prompt.
- Image to Video: This mode generates video using just a driving image.
- Stylization: This mode transfers the style of any image or prompt to every frame of your video.
- Storyboard: This mode turns mockups into fully stylized and animated renders.
- Mask: This mode isolates subjects in your video and modifies them with simple text prompts.
- Render: This mode turns untextured renders into realistic outputs by applying an input image or prompt.
- Customization: This mode unleashes the full power of Gen-2 by customizing the model for even higher fidelity results.
Runway Gen-2 Research and Pricing
Gen-2 is noted as a standard for video generation based on user studies, with results preferred over existing methods for image-to-image and video-to-video translation. The development of this tool is part of Runway Research’s mission to build multimodal AI systems that enable new forms of creativity.
However, Gen-2 is not perfect. As of now, you can only generate clips of 4 seconds long. It’s a paid tool, costing a minimum of $15 per month to generate 125 seconds of video every month. If you want to generate more, you will have to pay more. This tool is truly a boon for artists or creators looking to craft engaging content. Yet, the catch is in the cost. It may not fit everyone’s budget. Interestingly, this issue often surfaces among those content creators who are big in the game but are working with a rather tight budget.
Potat1
Potat1 is the first open-source 1024×576 Text To Video model. It’s a prototype model trained with 1xA100 (40GB) using 2197 clips and 68388 tagged frames from the salesforce/blip2-opt-6.7b-coco dataset. The model was trained for 10,000 steps.
Open Source and pricing
Potat1 is completely free and open source. It was trained on thousands of video clips to generate high-resolution videos. It doesn’t generate a watermark on their videos. To run this tool, you need a powerful GPU. It uses around 15 gigabytes of VRAM to generate a one-second clip. The quality of the generation is not as good as the Gen 2 of Runway ML. It’s slower in generating videos compared to Runway Gen-2. Potat1 requires a lot of VRAM to generate even one second of video, making it less accessible to the general public.
Information related to the dataset & config, finetuning, and the base model used can be found in the links provided on the Potat1 Hugging Face page.
Pros and Cons:
Runway Gen-2
- Pros: High-quality video generation, can take inspiration from a base image, useful for artists or creators, multiple modes for different needs.
- Cons: Only generates clips of 4 seconds long, expensive, requires multiple generations to get the desired result which uses more of your monthly quota.
Potat1
- Pros: Completely free and open source, doesn’t generate a watermark on their videos, trained on a large dataset.
- Cons: Requires a powerful GPU, slower in generating videos, the quality of the generation is not as good as Runway Gen-2, requires a lot of VRAM to generate even one second of video.
We trust that you found this blog post enlightening. If you’re on the lookout for more thrilling content we’ve curated, navigate to digialps.com. This is your ultimate destination for the freshest insights and developments in AI, technology, and marketing.
Our skilled team can efficiently handle everything from inception to execution.
If you’re eager to discover more, visit our website, or feel free to send an email to hello@digialps.com
We appreciate your time and interest in reading!






