sprited/kimodo

Kimodo text-to-motion: SOMA and G1 models, sequential prompts, motion constraints, multi-sample generation, GLB/BVH/NPZ and Mixamo FBX export, with animated previews. Unofficial NVIDIA Kimodo deployment.

Public
2 runs

Run time and cost

This model runs on Nvidia L40S GPU hardware. We don't yet have enough runs of this model to provide performance information.

Readme

Kimodo

Generate 3D skeletal motion from text, with optional motion constraints. Unofficial community deployment of NVIDIA Kimodo, hosted by Sprited.

Preview and downloads

The default output panel shows one motion video and one ZIP download. The ZIP contains all selected motion formats, metadata, PNG and interactive HTML viewer. Set output_delivery=individual for separate API file URLs.

Enable preview to get a browser-playable MP4 with front and side views, a PNG, and a self-contained interactive HTML skeleton viewer. The MP4 lets you inspect motion immediately; the HTML adds orbit, zoom, playback and frame seeking.

Choose output_format for GLB, NPZ, BVH, or FBX, or use all:

  • GLB contains the original skeleton hierarchy and animation. It has no character mesh or skin; use it as motion data or with a skeleton viewer/retargeting tool.
  • NPZ is the native Kimodo motion data.
  • BVH is available for SOMA models.
  • FBX requires your uploaded Mixamo-rigged character in custom_fbx and a SOMA model. sample_index, yaw_offset, and scale control the retargeted export.
  • A metadata JSON records the model, seed, skeleton, timing and generated filenames.

Motion controls

A period separates sequential motion segments: A person walks. A person waves. duration is seconds per segment, not total length. Optional segment_durations (for example [2, 3]) specifies each segment separately.

Generate 1–16 samples, using 0.5–30 seconds per segment and 10–500 diffusion steps. Larger requests take longer and may exceed available GPU memory. The default is one sample with 100 steps. Motion is finite, not guaranteed to loop.

constraints_json accepts the upstream Kimodo JSON format for root paths, full-body keyframes, hands, feet and other end-effectors. postprocess reduces foot skating; root_margin controls root correction. G1 skips postprocessing.

Available weights

Bundled: SOMA RP v1/v1.1, SOMA SEED v1/v1.1, G1 RP and G1 SEED. Default: Kimodo-SOMA-RP-v1.1. These use the original Llama-based LLM2Vec encoder. API callers do not need Hugging Face credentials.

The SMPL-X interface is retained for compatibility, but its separately gated weights are not available in this public deployment. Choose SOMA or G1.

Attribution and limitations

Built with Meta Llama 3. Kimodo source, NVIDIA model weights, Llama, LLM2Vec adapters, and the FBX dependencies retain their respective licenses and terms. See the Kimodo model card and ComfyUI-Kimodo, whose FBX retargeter is used here.

This renders skeleton previews, not a textured human or robot. Generated motion may miss parts of the prompt. Real character rigs need visual inspection after retargeting. Cold starts include loading the text encoder and motion model.

Model created
Model updated