8digit/gener8-sam3

SAM 3 video segmentation that keeps every object's ID across the clip: one 16-bit label map per frame, plus an optional colored preview.

Public
8 runs

gener8-sam3 — SAM 3 video segmentation that keeps every object’s ID

Segment every instance of a text prompt (default person) across a whole video with Meta’s SAM 3, and get each object separately, with an ID that stays the same for the whole clip. The same person keeps the same number from the frame they appear until they leave, including people who walk in later. Pick the ones you need and ignore the rest.

Most SAM 3 video wrappers return one merged mask of everything that matched. This one keeps the identities, so you can choose which people to replace, blur, relight or cut out.

Inputs

Input Default Meaning
video — The video. Up to 2000 frames. Frames are read as stored (rotation metadata is not applied), so upload upright video.
prompt person What to segment. Every matching instance gets its own ID.
preview false Also return an MP4 with every object tinted in its own color and labeled with its number.

Output

A list of files:

  1. ….zip — the bundle (always first)
  2. ….mp4 — the preview (only with preview: true)

The bundle contains:

  • labels/labels_00000.png, labels_00001.png, … — one 16-bit PNG per frame. A pixel is 0 for background and L for the object whose label is L. SAM 3’s masks are non-overlapping, so one image per frame holds every object without losing pixels.
  • objects.json — width, height, frames, fps, prompt and one entry per object: label, sam_id, frames (how many frames it appears in), first_frame, last_frame, mean_score, best_frame (the frame where it is largest), best_box (XYXY pixels on that frame), best_area, best_area_fraction, best_score.

Keep only the people you choose (Python)

import json
import zipfile

import numpy as np
from PIL import Image

keep = {6, 12}  # labels chosen from objects.json or the preview

with zipfile.ZipFile("bundle.zip") as bundle:
    meta = json.loads(bundle.read("objects.json"))
    for index in range(meta["frames"]):
        with bundle.open(f"labels/labels_{index:05d}.png") as handle:
            labels = np.asarray(Image.open(handle))
        mask = np.isin(labels, list(keep))  # True only on those people

To show a picker, crop each object’s best_box from its best_frame.

Run time and cost

Runs on an Nvidia L40S. Time grows with the number of frames and the number of tracked objects. Measured on a crowded 24 s clip (576 frames):

  • 720×1280, 29 people, with preview: 212 s tracking + 66 s preview — 291 s of GPU, about $0.28
  • 540×960, 26 people, no preview: 200 s of GPU, about $0.20

How it works

Hugging Face transformers Sam3VideoModel (SAM 3’s detector + tracker) in bfloat16, with the kernels-community/cv-utils kernel for non-maximum suppression, hole filling and sprinkle removal. Without that kernel, transformers silently skips those steps, so this model refuses to start without it.

License

SAM 3 code and weights are Meta’s, under the SAM License.

Model created
Model updated