p-video-2
Generate premium AI videos from text, images, or audio with native-speech lip-sync, sharp close-ups, and strong identity consistency.
P-Video-2 is Pruna’s quality-focused video generation model and the successor to P-Video.
How it works
The model takes:
- A prompt describing the subject, action, scene, and optionally camera, lighting, style, and audio.
- An optional image for image-to-video generation, or audio for audio-conditioned generation.
- A resolution, frame rate, duration, aspect ratio, and draft setting.
It returns a generated video at up to 1080p and 48 fps, with optional native audio.
Tips
- Start with the subject, action, and scene. Add camera movement, lighting, style, and dialogue when you need more precise and repeatable results.
- Turn
drafton for faster, lower-cost iteration, then turn it off for final-quality generations. - For native speech and lip-sync, write the dialogue directly in the prompt, keep
save_audioon, and don’t provide anaudioinput. - Use an
audioinput when generation needs to follow a specific track. The output duration follows the audio, sodurationis ignored. - For image-to-video, use a high-quality, well-lit reference image and keep the requested motion consistent with the starting frame.
aspect_ratiois ignored when an image is provided. - Use
last_frame_imagewhen you need control over the end frame. - Leave
durationempty to let the model choose the video length from the prompt, or set it explicitly from 1 to 20 seconds. - Prefer 24 fps at 1080p for stability, or 48 fps at 720p for smoother motion.
Resources
Docs: https://docs.pruna.ai/en/stable/docs_pruna_endpoints/index.html
Model created
Model updated