Concept engine

z-image-turbo, the Model That Makes Text to 3D Possible

The 3D engine needs a picture, so z-image-turbo draws one from your sentence in about six seconds. You approve it before anything slow happens — and if you do not like it, redrawing costs nothing.

~6s per imageOutputs one concept image$0, no accountUpdated August 2026

What z-image-turbo actually does

It is a fast text-to-image model, and on this site it has exactly one job: turn the sentence you typed into a single clear picture of a single object, on a plain background, lit evenly. That picture is what the 3D engine builds from.

Speed is the whole point of choosing it. Six seconds is short enough that redrawing feels free, which matters because the concept image is the step that decides whether the mesh will be any good. A slower, prettier model would make you reluctant to try again.

Your words are not sent to the model unchanged: a short instruction is appended asking for one centred subject on a plain ground, in frame. That is not editorialising your idea — it is the framing the mesh stage needs, and it is the single biggest lever on the final result.

A concept image drawn by z-image-turbo from a text prompt on 3D AI Studio
A concept image drawn from one sentence. The mesh engine reads this picture, not your words.

z-image-turbo at a glance

How the model behaves as we run it, rather than what it can do in general.

z-image-turbo as deployed on 3D AI Studio
DeveloperTongyi, released as z-image-turbo
TaskText to image, optimised for speed
Input hereYour prompt, plus a short framing instruction we append
OutputOne concept image, shown to you before anything else runs
Typical timeAbout 6 seconds
Role in the pipelineStage one of text to 3D. Image to 3D skips it entirely
Price here$0 — no account, no credits, no watermark
Commercial useThe image and the mesh built from it are yours to ship

Six seconds is the typical case on an ordinary prompt. The figure moves with queue depth, not with how complicated your sentence is.

Why this engine

Why z-image-turbo sits in front of the 3D engine

The extra step looks like a detour. It is the reason text to 3D works at all here.

The 3D engine needs a picture

Hunyuan3D-2 is image-conditioned. Without a concept image there is nothing for it to read, so a text prompt has to become a picture somewhere.

You get a veto

Six seconds in, you see what the mesh will be built from. Redraw it, edit the prompt, or go ahead — the expensive stage only runs once you say so.

Fast enough to retry

A model that took a minute per image would make the approval step feel costly. At six seconds, trying a different phrasing is free in every sense.

Framing is enforced here

One centred subject on a plain ground is what makes the mesh stage reliable, so that instruction is added at this step rather than left to your prompt.

From your prompt to a finished file

Four stages. The middle two are the only ones that touch a GPU, and you approve the hand-over between them.

  1. Text to 3D or image to 3D

    Type a sentence, or drop a photo. Image to 3D skips straight to the mesh stage; text to 3D needs a picture first.

  2. Concept image

    z-image-turbo draws your prompt in about six seconds. You keep it, redraw it, or edit it with flux.2.

  3. Mesh and texture

    Hunyuan3D-2 turns that one image into geometry and bakes the texture onto it. This is the slow step, and the progress you see is real job state rather than an animation.

  4. Download the .glb

    Take the file. Convert it to STL, OBJ or USDZ in your browser if the next tool needs something else.

The honest part

What z-image-turbo is good at, and what it is not

Judge it as the first half of a 3D pipeline, not as a general-purpose image generator.

Good at

  • Single objects, cleanly framed. A prop, a figurine, a tool or a piece of furniture on a plain ground is exactly its brief.
  • Material and style cues. Ceramic, brushed metal, low-poly, hand-painted — these come through clearly enough for the mesh to follow.
  • Speed. Six seconds makes iteration genuinely free, which is worth more here than a slightly better single shot.
  • Being thrown away. The image is a means to a mesh. Redrawing three times to get the silhouette right is the intended workflow.

Not good at

  • Scenes and compositions. Two objects, a background or an environment will confuse the stage that follows.
  • Text and logos. Small lettering is unreliable in the image and disappears entirely at mesh resolution.
  • Photographic realism. It is fast rather than photoreal, and a photoreal concept would not make a better mesh anyway.
  • Precise composition control. There is no inpainting or region control here — for a targeted change, use flux.2 on the concept instead.

Write the prompt like a product photo brief. Name the object, the material and the style, and stop. Scenes, backgrounds and camera language make a picture the mesh stage will struggle with.

The generator

Make a 3D model right here

You do not have to go back to the home page. Type a prompt or drop an image, and take the .glb — free, anonymous, no watermark.

Free · anonymous · no queue
Try:
Turnstile — human check, no account

z-image-turbo — frequently asked questions

A fast text-to-image model. On 3D AI Studio it draws the concept image that text to 3D is built from — it does not build geometry itself, and it never touches your uploaded photos.

Want to see the output before you commit?

Generate one, then open the .glb in the free viewer and read the triangle count yourself.

One sentence, six seconds, then a mesh

Type a prompt and watch the concept appear before anything slow happens. Free, anonymous, no watermark.

FreeNo signupNo watermarkDownloads as .glb