Esc
Studio

Your own studio
instead of five
subscriptions

Images, video with sound, music and voice are made inside VibeBot. You ask in plain words — the bot picks the model, queues the job and brings the result into the conversation.

Images

An article cover, a banner, a post illustration, a product shot. Edits are words too: “make it darker”, “drop the caption”, “add steam above the cup” — the bot changes what you asked and leaves the rest alone.

  • Draw a calm, light cover about a coffee shop»
  • Same one, but at dusk»

Video with sound

A short clip arrives with its own sound — room tone, footsteps, hum, clicks. The film on this site’s home page was made exactly this way: each scene took about seven minutes on a single machine with a graphics card.

  • Make a five-second clip: evening, a person with a laptop»
  • Frame by frame, with sound, no outside services

Music

A bed for a clip, a jingle, a calm ninety-second track. Name the mood and the tempo; the bot handles the rest. The score in our film took three minutes of server time.

  • A calm instrumental, about ninety seconds»

Voice

The bot reads text aloud or unpacks your voice message. The same engine handles narration for clips and transcripts of long meeting recordings.

  • Read this out and send it as a file»
  • Transcribe the call and pull out the decisions»
Honestly

Where pictures
and clips are rendered

Models and privacy

Your server

A machine with a GPU and ComfyUI: enter the address and VibeBot polls the station itself and shows your models. No subscriptions, no limits, no one else looking.

Your computer

ComfyUI is on your laptop — no ports to open: the VibeBot desktop app passes the job to the station for you.

A cloud by key

No hardware at all — plug in a cloud key. It renders, you pay, and the price is shown before the run.

The built-in recipes work on a plain ComfyUI install too: the model is taken from yours, not ours — any SDXL or SD 1.5 will do.

How it looks in a conversation

1

Ask in plain words

“Make a cover”, “I need a five-second clip”, “find music for this video”.

2

The bot decides

Model, size, length and queue are its job. It only asks when the fork really matters.

3

Result in the chat

The image, clip or track lands right in the conversation. Not right? Say what to change.

Next