zf-kbot/md-me-image-llada

Fast LLaDA Image Turbo FP8 generation and editing with an optimized text encoder

Public
33 runs

Run time and cost

This model runs on Nvidia L40S GPU hardware. We don't yet have enough runs of this model to provide performance information.

Readme

LLaDA-Image-Turbo-FP8

Text-to-image generation and single-image editing based on inclusionAI/LLaDA-Image-Turbo-FP8.

The latest version requires an access key in apikey (Secret) and a prompt. Provide input_image for single-image editing. aspect_ratio defaults to auto: editing follows the EXIF-corrected source ratio; text-to-image is square. Manual ratios center-crop the reference without stretching. Available ratios: auto, 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16, 21:9, 9:21. Auto ratios must be between 1:4 and 4:1.

resolution_mode=longest_edge (default) produces a 1024px longest edge. Select megapixels mode for native 0.5–2.0MP output; the default area is 1.0MP. Requests above 2MP are rejected. Inference padding is cropped from the returned image. High resolutions use encoder CPU offload and take longer.

steps supports 2, 4 (default), and 6. Omit seed or pass null for randomness, or use 0–2147483647. file_format defaults to jpg (quality 90); png is lossless. Outputs are RGB.

Grouped FP8 expert kernels reduce warm text encoding to about 63 ms on L20, with default four-step generation around 2.96 seconds before saving. Hosted latency includes platform overhead. Higher-resolution edits can be softer and warmer; small text and precise typography may contain errors.

Version 947ea276170d970e6ad1575d51e2ca7aa1d83ac2c8fb7971832690d78a8d8cce uses this interface. Older immutable versions retain their original parameters.

Model revision: 664a975e4b4980fb750fd47b2f87e1874bb14bd7. Upstream code, revision 6bac7e7c1618e3ed9ec075ba75ff78cedfca6b38.

Model created