
The PixVerse AI image to video API uses separate identifiers for an uploaded image and a generated video task. Store the image ID, submit one generation request, then poll the returned video ID until it completes or fails. Validate the model settings before adding effects or increasing the queue.
This guide covers the PixVerse Platform API. Its operations are separate from the consumer app interface and its subscription benefits.
If you want browser-based character media without building an API integration, Cherrypop’s creator lets you start with an original character and continue into media. The technical workflow below serves a different need; no shared API or matching provider controls are implied.
Map the PixVerse AI image to video request
PixVerse’s image-to-video guide describes uploading an image, creating a generation task and checking its status before downloading. The upload supplies an img_id; the generation response supplies a video_id. A trace identifier is a third concept used for requests.
| Identifier | Meaning in the workflow | Mistake to avoid |
|---|---|---|
| img_id | Uploaded image used as input | Treating it as a finished video |
| video_id | Video generation task to query | Replacing it with a new task when polling |
| Ai-trace-id | Request trace value | Reusing one constant for every unique request |
| template_id | Optional selected effect template | Adding an effect unintentionally to a baseline |
Keep these values in a persistent job record alongside the input file, prompt, model and settings. A page refresh or worker restart should resume tracking the same task. Keep API keys in server configuration and out of the record, browser code and diagnostic logs.
Prepare an ordinary image first
The upload reference lists JPEG/JPG, PNG and WebP, a file size below 20 MB and maximum dimensions of 10,000 pixels. The guide recommends at least 1024 by 1024 pixels and a clear subject; that recommendation is not a minimum accepted size.
For an initial creative brief, use an original photograph of a brass desk bell. Ask for one restrained camera movement while the bell remains on the table. This gives the review a clear object outline, a stable base and a simple intended change.
Keep the source image separate from any preview thumbnail. If an integration accidentally submits the thumbnail, a successful upload can conceal the wrong input. Save the asset identifier and a non-sensitive checksum with the job so you can confirm what was actually submitted.
For a file upload, send multipart form data to POST /openapi/v2/image/upload on https://app-api.pixverse.ai, with the file in the image field. The reference also shows an image_url option. Supply the documented API-KEY and Ai-trace-id headers. On success, save Resp.img_id; the upload response also includes Resp.img_url.
Establish a plain image-to-video baseline
Send JSON to POST /openapi/v2/video/img/generate on the same API host, using those headers plus Content-Type: application/json. The image-to-video reference returns Resp.video_id in an envelope containing ErrCode and ErrMsg. Check the HTTP response and that envelope before accepting the task. The documented success examples use ErrCode: 0.
Start with the smallest configuration that expresses the task. An illustrative prompt for the bell is: “The camera moves slowly closer to the brass bell while the bell remains still on the desk.”
The V6 model page makes the baseline concrete. For image-to-video it requires model: "v6", an integer img_id, a prompt of at most 5,000 characters, integer duration from 1 to 15 seconds, and a quality string of 360p, 540p, 720p or 1080p. A five-second, 720p bell shot falls within those documented settings. It has not been executed here.
V6 lists optional generate_audio_switch and generate_multi_clip_switch booleans. Set them deliberately when audio or multiple clips affect the brief. Its image-to-video column does not support aspect_ratio, although text-to-video does. Do not copy the full text-to-video configuration into this endpoint.
Add effects only when the transformation is the task
PixVerse’s Video Effects documentation describes choosing an effect and adding its template ID to the generation parameters. For multiple-image templates it specifies img_ids, with the required count shown in Effect Center. A list of images is a different input shape from the single img_id baseline.
For the desk bell, decide whether you want a faithful object shot or a stylized transformation. Both can be legitimate creative projects, but they should have different acceptance criteria. A transformation that changes the bell is not a defect if changing it is the purpose of the effect.
The examples need care: the effects guide uses v4.5, while the generation reference uses v6. The generic reference also shows separate sound-effect fields, whereas V6 documents generate_audio_switch. These examples do not establish every template or optional-field combination for V6. Verify the selected template’s supported configuration before combining them. The published samples also contain comments or trailing commas; assemble valid JSON rather than pasting them unchanged.
Make effects opt-in in your application. Keep the template ID in the visible job summary so that someone reviewing a surprising result can see whether the transformation was requested.
Treat pending, completed and failed as different outcomes
Use GET /openapi/v2/video/result/{video_id} with the saved task ID and the documented headers. The status reference returns the job state in Resp.status. A successful status-query response only means the query worked; inspect the nested state before showing a result.
Use the documented states to drive the next action:
| Resp.status | Documented meaning | Application action |
|---|---|---|
| 1 | Generation successful | Read Resp.url and retrieve the file |
| 5 | Waiting or generating | Poll the same video_id |
| 6 | Deleted | Stop polling; show that the task was deleted |
| 7 | Moderation failed | Stop; show a distinct failure |
| 8 | Generation failed | Stop; retain diagnostics for review |
Keep unexpected states visible as unresolved. Only the documented successful state should unlock the result path; do not interpret an unrecognized number as completion. A missing result URL should trigger investigation instead of a download button.
The status guide mentions both 3–5-second polling and intervals of no less than five seconds. Five seconds satisfies both statements. Add an application timeout and a later-resume path; reaching your timeout does not establish that the provider task failed.
Diagnose the failing stage before retrying
Separate transport failures, API errors, terminal job states and unwanted creative results. Save the request time, endpoint, trace ID, model, HTTP status and returned error fields so a support report identifies the actual attempt.
- 400013 or 400017: the image-to-video guide identifies invalid request values or parameters. Check JSON syntax, value types, the input ID and the selected model’s supported fields before resubmitting.
- 500044: the guide identifies a concurrency limit. Pause new submissions while existing work finishes. PixVerse’s rate-limit page defines concurrency as simultaneously generating tasks, so increasing the number of workers does not remove the account limit.
- A task stays pending: confirm you are polling its saved video ID and using distinct trace IDs for unique requests. The guide flags reused trace IDs as a troubleshooting issue; it does not establish automatic deduplication.
- A create request times out: its outcome may be unknown. Check your response logs and account records before authorizing a replacement. If no task ID was received, there is nothing to insert into the status URL yet.
A moderation failure is a terminal outcome, not a trigger for automated attempts to evade safeguards. A completed clip with unwanted motion needs a creative revision, not the same retry treatment as a connection error.
Budget for the API configuration you selected
The official guide requires appropriate API access and available or purchased credits. This article does not establish that a consumer plan includes API usage, or that a displayed promotional benefit applies to your chosen endpoint and model.
For a batch, record the charge for the selected configuration and the number of requested outputs. Reserve budget for a small number of explicit corrections rather than an open-ended retry loop. Separate transport retries, failed tasks and creative alternatives in the record; they answer different questions about reliability and cost.
Check actual account usage after an authorized trial. This guide does not establish current charges, consumer-plan entitlements or how a refund appears in an account.
Verify the file and the creative result
Download the successful result and confirm that it plays, has the intended framing and fits the project’s duration requirement. Then review the bell’s outline, base and background at several points. A returned URL proves less than a file you have opened, and a playable file proves less than a clip that meets the brief.
Keep the output beside its input and generation record. The completed workflow should answer which image, model, settings and optional template produced that file. Before increasing the batch size, confirm that one job can be recovered after a restart and traced through to a playable file.
When you want to create with a character instead
An API integration is useful when you need to manage submissions, task records and files in your own application. If your immediate goal is an original companion character, Cherrypop offers a browser workflow: define appearance, personality and a scenario, then continue into chat and supported media creation.
The Cherrypop video entry includes image-to-video and continuation modes. Treat that as an alternative task path, not an implementation of the API described above or a promise of identical duration, resolution or output quality.
Cherrypop is free to start with limits; relevant media may require Premium or Cherries. Check the displayed operation cost. Create the character and opening scenario when the interaction matters as much as the clip.