Table of contents
- Introduction: Simplifying Video Editing
- The Basics: Representing Videos with Neural Atlases
- The New Era: The INVE System
- Overcoming the Speed Challenge
- Adding Advanced Effects with Bidirectional Mappings
- Incorporating Flexibility with a Layered Editing Model
- Enabling Direct Sketching on Frames
- Notable Results
- Conclusion: A Giant Leap in Video Editing

Introduction: Simplifying Video Editing
Video editing, which typically requires a great deal of expertise and hard work, is on the brink of a transformation. A new trend in AI research is looking to change it into a more user-friendly creative process. An innovative method known as Interactive Neural Video Editing (INVE), developed by Adobe and the University of British Columbia, is spearheading this transformation. INVE builds on the previous work of Layered Neural Atlases (LNA) by optimizing neural networks and graphic techniques, making real-time video editing more flexible and direct.
The Basics: Representing Videos with Neural Atlases
Firstly, we must understand the underlying technology. LNA uses implicit neural networks to represent videos. It turns scenes into 2D texture maps, known as atlases, one for each moving object and the background. It then employs multilayer perceptrons (MLPs) based on coordinates to map 3D video coordinates to these 2D atlas locations. Additional MLPs predict colors at atlas coordinates and blend the layers together.
The training process for this method is self-supervised. It involves reconstructing frames and promoting rigidity, sparsity, and temporal consistency through regularization losses. The actual editing occurs by modifying the atlases. Although LNA managed to allow spatial edits to spread over time, it fell short in terms of speed, control, and creative possibilities.
The New Era: The INVE System
INVE has made crucial enhancements to transform neural video editing. It introduces:
Speeding Up Neural Network
Highly optimized network architectures to make training and inference 5 times faster.
Advancing Effects with Bidirectional Mappings
Inverse mapping networks that enable advanced effects like texture tracking.
Adding Flexibility with Layered Editing
A layered editing model that supports separate adjustment layers.
Making Editing Direct with Vectorized Sketching
The ability to draw directly on frames through vectorized sketching.
Overcoming the Speed Challenge
One of the main challenges with LNA was its sluggish performance, which limited essential interactivity for creative workflows. INVE has overcome this by adapting multi-resolution hash encoding techniques from graphics to greatly speed up coordinate-based MLPs. This method, known as hash grids, significantly enhances encoding frequency and reduces the load on subsequent layers.
Along with this, fused MLP implementations further optimize computation. The combined impact of these strategies significantly improves the speed of convergence during training and responsiveness during editing.
Adding Advanced Effects with Bidirectional Mappings
While LNA only supported mapping pixels from frames to atlases (forward mapping), it did not allow the reverse. To tackle this, INVE has introduced inverse mapping networks, enabling bidirectional mapping between frames and atlases. This innovation powers advanced effects like texture tracking, which allows graphics to follow objects in 3D, as opposed to simple 2D distortion.
Incorporating Flexibility with a Layered Editing Model
To add more flexibility, INVE allows for layered editing with separate adjustment layers. This includes one layer for sketches, one for textures, and one for local color edits. With this feature, users can easily access and modify each layer individually, paving the way for sophisticated compositions and localization.
Enabling Direct Sketching on Frames
Earlier techniques necessitated the direct editing of atlases. However, the nonlinear neural mappings made it hard to visualize and control edits. INVE has solved this problem by allowing users to draw vectorized sketches directly on video frames. These are automatically mapped to atlas layers, thereby avoiding resampling artifacts and enabling intuitive frame-level editing.
Notable Results
Experiments have shown that INVE can propagate diverse edits such as sketches, graphics, and color changes consistently at a speed of 25 frames per second. Compared to LNA, INVE shows significant advantages in training and inference speed, reconstruction quality, editing flexibility and control, and the potential for creative possibilities.
Conclusion: A Giant Leap in Video Editing
By optimizing neural networks through techniques like encoding, fusion, and bidirectional mapping, and integrating vector graphics techniques, INVE paves the way for more interactive and direct video editing than was previously possible. Instead of struggling with abstract atlases, users can now easily experiment and iterate on frames, similar to editing photos. This research is a thrilling step forward towards democratizing cinematic creativity.
We also make great websites. Don’t wait! Talk to us today. Our friendly team is ready to help you in the best way possible!
Check out our other awesome articles:







One Response