
Product Video Cards for Marketplaces: Assembling in One Hour
How an online store can quickly turn product descriptions into video cards for marketplaces without shooting and an editor — an analysis from Telematic.

How an online store can quickly turn product descriptions into video cards for marketplaces without shooting and an editor — an analysis from Telematic.
Most online stores have a product card with photos, specifications, and descriptions — and almost never a video. The reason is not that videos don’t work: a product video almost always boosts conversion and holds attention longer than a static photo. The reason is that shooting, editing, and voicing a video for each item in the assortment is a separate profession, a separate budget, and a separate person, who is often simply not available in the staff of a small store.
If there are fifteen products in the catalog, you can hire a freelance editor for a one-time project. If there are one and a half thousand, and the assortment is updated every week, manual editing ceases to be a viable solution — there won’t be enough time, money, or patience. This is the scenario we at Telematic encounter most often with e-commerce clients: it’s not that “we don’t want video,” but rather “we physically have no way to produce video at the required volume.”
Next, we will discuss how we solved this problem for ourselves and our clients: not by replacing the operator and editor with an expensive studio, but by eliminating the need to shoot and edit manually, leaving as input what the store already has — the product card text and photos.
The most common source for a video card is not shooting from scratch, but what is already written in the product card: specifications, advantages, composition, and packaging. This text can be entered into Telematic manually or pulled from the store's website as a source — and then AI processes it into a script, rather than summarizing it in three general sentences. Numbers, specific characteristics, and formulations of advantages are preserved — because for a product card, invented general phrases like “high-quality and stylish” do not work, while “bowl capacity 1.5 liters, six programs, timer up to 24 hours” — do work.
Here, a hybrid script mode comes in handy: AI transforms a dry description into lively voiceover, as if the product is being presented by a sales consultant. For products where the accuracy of wording is important (medical devices, equipment with warranty conditions), the full text mode is more suitable — then the voiceover text is taken from the description almost verbatim, without paraphrasing. Which mode to choose depends on the niche, but this can be switched with one setting on the project, rather than being manually adjusted for each product.
A separate setting is the narrative pace: for a short video about a low-cost product, a dynamic pace is usually needed, while for expensive equipment with a long list of specifications — a standard or calm pace, so that the viewer has time to read the numbers.
Before generating frames in Telematic, there is a separate editing phase — a timeline where the scenes of the future video can be rearranged, trimmed, and assembled in the desired order even before images are generated for them. This is convenient for a video card: for example, a scene with the price or promotion can intentionally be placed at the beginning, rather than at the end, where half of the viewers may not watch it.
For visuals, it is not necessary to generate images from scratch — there is a styling mode for the store's own product photos: an existing photo can be “dressed” in one of the visual styles (there are 29 in the library, plus separate operator “looks” — camera effects, transitions, and colors), without losing the product and its real appearance. This is important for cards: the buyer should recognize in the video exactly the item they saw in the photo, not an abstract generated image.
The frame format is also set separately from the video format itself: for example, you can create a vertical video for stories while generating product images in a wide angle — this way, the detail is better seen in its entirety, rather than a cropped close-up.
A video card almost always has at least two recipients: the card on the marketplace (often horizontal or square format) and stories or shorts on social media (vertical). Doing this as two separate projects is inefficient — in Telematic, the same script can be expanded into two orientations at once, and for social media, a short vertical teaser can be additionally created — a condensed version of the main video without a call to action, simply “watch the full video.” Such a teaser works well as an announcement in stories that leads to the full card.
Subtitles for the video card are not an option, but a necessity: most people browse the marketplace feed without sound. In Telematic, the subtitle area is set with a draggable frame for a specific platform — this way, the text does not cover the price, the “add to cart” button, or other interface elements that the platform overlays on the video.
For the voiceover, there is no need to hire a narrator for each batch of products — you can choose a voice from the library or clone your own once, and then it is used in all videos of the series, maintaining the recognizable brand intonation. A separate detail that practically solves more problems than it seems: text normalization before dubbing. Product specifications are full of numbers, abbreviations, and units of measurement — “1.5 l,” “220 V,” “up to -20°C” — and without normalization, the voice engine reads them unnaturally. We specifically tested videos without this step and with it: the difference in perception is noticeable even to an untrained ear, especially for equipment and products with numerical specifications.
The true value of automation is not revealed with one product, but with a batch. When there are fifty or two hundred cards in the queue, manually starting the text processing, editing, and publishing for each one is also a full day's work. For this, Telematic has an autopilot: one click launches the full cycle from text processing to finished video and publication, and it can be applied in bulk to several materials at once, rather than one by one.
Photos and already generated videos go into a common media library — this is convenient when the same product needs to be reused in a new promotion or seasonal selection: there is no need to search for the originals again, they are already indexed and available for reuse.
When we tested this scenario on our internal demo product catalog, the longest part of the process turned out to be not the video generation, but the initial preparation of product texts — the actual assembly of the video after that indeed took just a few minutes per item, not hours. For a store with a large and frequently updated assortment, this is the main win: video ceases to be a luxury for a handful of bestsellers and becomes a standard part of every card.
How to Connect Telegram, YouTube, and VK in One Go
Car Dealerships: Video of Every Car Without Operator On-Site