For footwear

AI model video for
footwear.

Shoes are not a drape problem. Nothing about a boot falls or flows; it is a rigid object whose whole appeal is its silhouette and how it sits against a leg. That makes footwear a different generation problem from clothing, with a different source photo and a different set of things to check.

Proportion on a leg, not on a board Shaft height where it actually lands

What does footwear video show that a product shot cannot?

Proportion. A shoe photographed alone has no scale: a chunky sole reads the same at any size, and a boot shaft could end anywhere. On a leg, all of that resolves at once. Where the shaft lands against the calf, how high the platform really is, how wide the toe box looks in context. Those are the questions a footwear shopper actually has.

The second thing is stance. A heel changes how a foot sits and how the leg lines up above it, and that is a large part of what someone is buying when they buy a heel. A flat product shot on a white background shows the object. A model shot shows the effect of wearing it, which is a different piece of information and the one closer to the purchase.

Which source angle works best for shoes?

A three-quarter view, not a top-down flat lay. Clothing photographs well laid flat because the garment's cut is the information. A shoe laid flat and shot from above loses its profile, and the profile is the product. Shoot it at the angle a shoe is normally shown: slightly from the side and front, with the sole line visible.

What that angle needs to carry: the full silhouette from toe to heel, the sole profile including any platform or heel height, the upper's material, and any strap, buckle or lace structure in its natural position. If the heel is hidden behind the shoe in your photograph, the generation will invent a heel, and it will invent an ordinary one.

AI model video for footwear: a tan leather ankle boot photographed at a three-quarter angle beside the same pair worn with cropped cream trousers, showing where the boot shaft meets the ankle above a stacked block heel.
The angle on the left is the one to shoot, because it keeps the sole line and the heel readable. The frame on the right is the one that answers where the shaft actually lands.

Does the model's body type matter for footwear?

Less than it does for clothing, and this is worth saying plainly rather than overselling the gallery. A dress needs to be seen on a build near the shopper's own, because fit is the question. A shoe mostly needs a leg for scale, and the difference between builds matters much less to the result.

Where it does matter is with tall boots, because the shaft's relationship to the calf changes with the leg, and with anything worn against bare skin. For most trainers, flats and low boots, picking one model and staying consistent across the range is a better use of credits than generating the same shoe on several builds.

Which footwear is hardest to generate?

Heels and tall boots, for the same underlying reason: both put the shoe's structure in a precise relationship with the ankle and calf, and small errors there are immediately visible. A heel at slightly the wrong angle reads as wrong even to someone who cannot say why. Knee-high boots are the hardest single item in this category.

After those, it is fine hardware and branding. Buckles, eyelets, lace patterns and side logos sit at the scale where 1080p output and generative detail both start to struggle, so check them on the still. Technical soles with complex tread are the other common failure: the tread will render as plausible rather than as yours.

How do you keep a footwear range consistent?

Same model, same framing, across the whole range. Footwear is bought comparatively, with a customer flicking between four pairs, and a range where each shot uses a different body and a different crop reads as chaotic even when each individual render is good. Pick one setup and apply it to everything.

The cheap way to enforce that is to generate every still first, at 1 credit each, lay them side by side, and only animate the ones that match. Approving stills before spending on video is the whole point of the two-stage flow, described on check the still before rendering.

What does one pair cost to put on a model?

One credit for the try-on still, and three more for a five second video built from it, so 4 credits for a finished clip. A ten second video is 6 credits. The free plan's ten credits a month is three finished videos or ten try-on stills, which is enough to cover a small range in stills and animate the best of them.

Output is 1080p, and there is no 2K or 4K on any plan, which is worth knowing if you are producing footwear assets for large-format display rather than a product page. One generation gives you 9:16, 1:1 and 16:9. Getting the result onto the product page is covered in publishing to the product page, and the underlying two-stage pipeline in the two-stage generation.

Questions

About shoes.

Should I photograph shoes as a pair or singly?

A single shoe at a three-quarter angle gives the clearest source, because a pair introduces a second silhouette and often some overlap that hides the sole line. If your catalogue photography is already a pair shot it will usually work, but where you are shooting specifically for generation, one shoe with its full profile visible is the better input.

Does it handle knee-high and over-the-knee boots?

It will attempt them, and they are the hardest item in this category. The shaft has to meet the calf at a believable height and follow the leg's shape, and errors there are obvious. Generate the still first at 1 credit and look specifically at where the shaft lands and how it sits against the leg before spending 3 credits animating it.

Will the sole tread and branding render accurately?

Not reliably at fine scale. Tread patterns, small side logos, eyelets and stamped branding sit where generative detail and the 1080p ceiling both start to give out, so treat them as approximations rather than as documentation of your product. If a specific detail is a selling point, keep a real photograph of it in the gallery alongside the video.

Can I use the same model across my whole range?

Yes, and for footwear that is usually the right call. Shoes are browsed comparatively, so consistent framing and a consistent model make a range legible in a way that varied casting does not. Base AI models are available on every plan, and all AI models including premium ones are on the paid tiers.

Does it work for trainers and technical footwear?

Yes for the silhouette and the general look, with the same caveat about fine detail. Trainers are among the easier items here because the shape is familiar and the shoe sits flat on the foot, so proportion errors are less likely than with a heel. Complex tread and panelled uppers are where you should expect approximation.

Three videos, free

Try it on the boot with the tallest shaft.

Ten credits a month, which is three finished videos or ten try-on stills. Start with the hardest pair in the range, because the easy ones were always going to work.