gener8-sam3 — SAM 3 video segmentation that keeps every object’s ID
Segment every instance of a text prompt (default person) across a whole video with Meta’s
SAM 3, and get each object separately, with an ID that stays the same for the whole clip.
The same person keeps the same number from the frame they appear until they leave, including
people who walk in later. Pick the ones you need and ignore the rest.
Most SAM 3 video wrappers return one merged mask of everything that matched. This one keeps the identities, so you can choose which people to replace, blur, relight or cut out.
Inputs
| Input | Default | Meaning |
|---|---|---|
video |
— | The video. Up to 2000 frames. Frames are read as stored (rotation metadata is not applied), so upload upright video. |
prompt |
person |
What to segment. Every matching instance gets its own ID. |
preview |
false |
Also return an MP4 with every object tinted in its own color and labeled with its number. |
Output
A list of files:
….zip— the bundle (always first)….mp4— the preview (only withpreview: true)
The bundle contains:
labels/labels_00000.png,labels_00001.png, … — one 16-bit PNG per frame. A pixel is0for background andLfor the object whoselabelisL. SAM 3’s masks are non-overlapping, so one image per frame holds every object without losing pixels.objects.json—width,height,frames,fps,promptand one entry per object:label,sam_id,frames(how many frames it appears in),first_frame,last_frame,mean_score,best_frame(the frame where it is largest),best_box(XYXY pixels on that frame),best_area,best_area_fraction,best_score.
Keep only the people you choose (Python)
import json
import zipfile
import numpy as np
from PIL import Image
keep = {6, 12} # labels chosen from objects.json or the preview
with zipfile.ZipFile("bundle.zip") as bundle:
meta = json.loads(bundle.read("objects.json"))
for index in range(meta["frames"]):
with bundle.open(f"labels/labels_{index:05d}.png") as handle:
labels = np.asarray(Image.open(handle))
mask = np.isin(labels, list(keep)) # True only on those people
To show a picker, crop each object’s best_box from its best_frame.
Run time and cost
Runs on an Nvidia L40S. Time grows with the number of frames and the number of tracked objects. Measured on a crowded 24 s clip (576 frames):
- 720×1280, 29 people, with preview: 212 s tracking + 66 s preview — 291 s of GPU, about $0.28
- 540×960, 26 people, no preview: 200 s of GPU, about $0.20
How it works
Hugging Face transformers Sam3VideoModel (SAM 3’s detector + tracker) in bfloat16, with the
kernels-community/cv-utils kernel for non-maximum suppression, hole filling and sprinkle
removal. Without that kernel, transformers silently skips those steps, so this model refuses to
start without it.
License
SAM 3 code and weights are Meta’s, under the SAM License.