We are incorporating AI into the creation of creative content for ARVAL – the market leader in operational leasing
Many products and services are based primarily on rational decision-making. Customers do not choose them on impulse, but based on price, specifications, long-term value and a comparison of alternatives. Operating leases are a prime example of this. The decision concerns the financing and the entire ecosystem of services surrounding the vehicle, where, at first glance, it might seem that it is purely a matter of figures.

In reality, however, emotion plays a crucial role here too. The first impression of a car, the idea of how it might be used, or the life situation the customer associates with it, often determines whether they even begin to consider the rational aspects of their choice. It was precisely this tension between the rationality of the product and the emotional aspect of decision-making that, following our success in the tender for Arval, led to us being tasked with designing the creative element of the campaigns for 2026. The aim was to find a format capable of bridging these two worlds – precise, fact-based communication combined with a visual and emotional layer capable of capturing attention in the digital environment. Alongside traditional banners, we therefore decided to work extensively with short videos generated using AI.
Brief: How to approach operational leasing customers in a different way
The basic brief was fairly straightforward – how to communicate the concept of an operating lease visually in a way that would appeal to those interested in both new and used cars, whilst avoiding a generic or indistinguishable appearance.
Traditional formats have clearly defined boundaries in this segment: vehicle specifications, price, benefits, comparisons. In an environment where users switch between content in a matter of seconds, it is not enough simply to provide information. It is essential to utilise visual dynamics and capture the viewer’s attention immediately. We therefore decided to split the video component of the campaigns into two distinct concepts, which differ both in visual style and in the way they engage the viewer’s attention.
Two different approaches to creativity
The initial concept was based on the simple principle of showcasing the product in its natural environment. After researching the communication strategies of ARVAL’s overseas branches, we drew inspiration from the Slovakian branch, which used a photorealistic model of the car against a stylised, almost sci-fi backdrop. We were particularly struck by the camera movement and the clarity of the message. We took a slightly different approach. We placed each car in its natural context, so that the focus was not on an abstract background, but on situations in which a specific car makes sense.
The Mercedes E-Class appears in a modern urban neighbourhood, where it visually emphasises its premium character. We’ve set the Hyundai i30 against the backdrop of a country road and a weekend getaway, where we’re focusing on family use. With the Kia brand, on the other hand, we drew on the carmaker’s official materials, which feature a modern urban setting, and for selected models we went even further into a futuristic world that matched the design of the car itself.
The second concept is based on a completely different principle. It is not about a static placement within the environment, but about transformation. We drew inspiration from short-form content on social media, which visually references the theme of transformation familiar from the *Transformers* film series. An old car is driving along the road and, within a few seconds, visually transforms into a new model. This moment is complemented by sound design, which plays a crucial role in the overall impression. The aim is not merely to show the change, but to create a brief, dynamic visual moment that holds the viewer’s attention right to the end.
Dead ends: where AI ceases to be a trivial tool
Initial experiments have shown that the mere availability of tools does not automatically result in functional and visually realistic output.
For this project, we chose the openart.ai platform, which offers dozens of models for generating both images and videos. However, it quickly became clear that the problem wasn’t simply ‘generating something’, but rather generating consistent and accurate output.
The first major dead end was the choice of model. There are 33 different models available for generating static image data and a further 24 for video. Each reacts differently to the prompt provided, handles realism differently and interprets the reference photographs provided in a different way. Selecting the right combination of models and prompts proved to be a task in its own right, involving days of testing and hundreds of unsuccessful outputs.
A second dead end arose in relation to the consistency of the cars. The AI models tended to alter the car’s design, mix different generations of cars or combine different makes into a single hybrid object. In other words, precisely what is unacceptable in automotive design.

The third issue arose in the animation itself. Although the output looked good visually, the details often revealed the limitations of generative models. A typical example was static logos on a moving vehicle, or the mismatch between the movement of the wheels and their reflections. At first glance, the video worked, but on closer inspection it was clear that the physics were inconsistent.
The process: from the prompt to the final video
We eventually streamlined the entire production process into a few steps that are repeated for each output.
The process begins with the creation of a static visual. This is generated using a precisely defined prompt, which is then adjusted by the LLM to match a specific model within the tool. Working with reference photographs plays a major role. We include a real photograph of the specific car in the prompt, thereby significantly limiting the model’s creative deviations. This step has proved to be crucial. When the AI has only a text description, there is a high degree of variability in the results. However, when it has a visual reference, the accuracy of the generated model increases significantly.

Video generation builds on this static foundation. Once again, using an LLM, we refine the prompt, this time focusing on camera movement, scene dynamics and the desired type of animation. The result is a short video clip based on a precisely defined visual.
The final stage of editing is carried out manually in DaVinci Resolve Studio, where our editors combine the individual clips, add text overlays and finalise them for specific campaigns.
Iteration and debugging
The most difficult part of the whole process was not the production itself, but stabilising the workflow. Selecting a suitable model and combining it with prompt engineering took approximately two days of pure working time. At this stage, we tested the individual models one by one, always using the same prompt, in order to understand their behaviour.
The LLM assisted us in this process as an analytical tool. We outlined the entire scope of the project to it, including a list of available models and the required outputs. Based on this, we drew up a shortlist of candidates, which we then tested manually. Further iterations focused on the consistency of the cars. Here, we gradually fine-tuned the prompt to include the exact make, model and year of manufacture. It was only the combination of textual accuracy and visual references that yielded consistent results. We use AI to generate a video from the static model, which we then refine manually.
An editor’s perspective: a change in working methods
An interesting side effect of the whole process was a change in the very approach to video production. From the perspective of our editor, Antonín Pulkrab, who has been involved in video production for over 15 years – including through his collaboration with Czech Television – the role is undergoing a fundamental shift. The traditional workflow begins with a vision and working directly with the footage. In this case, however, the footage is not created in advance. It is created solely on the basis of a description.
This means that the work shifts from the realm of editing to that of accurately describing the result. The editor becomes more of a conductor of the brief than a creator of visual material. The change is fundamental, primarily in that creative control is partly replaced by error checking and iterations of the output.
Compared with traditional banner formats, we recorded almost double the click-through rate for these AI-generated videos across all campaigns, whilst the cost per click was roughly half. Furthermore, for some performance campaigns, there was also a 25–30 per cent reduction in the cost per conversion.
Conclusion: rapid production with a high degree of variability
The resulting system makes it possible to create short marketing videos in a matter of hours, often in less than two. Production costs are also significantly lower than with traditional video production. The tool stack used, including openart.ai, operates on a credit-based model, where the cost of a single video is in the region of a few hundred crowns. From a scalability perspective, this represents a major step forward. It allows for the rapid testing of various creative concepts, message variations and visual styles without the need for high production costs.
When searching for new creative formats for our campaigns, we needed a partner who could combine marketing thinking with modern technologies. The TRITON IT team introduced the use of AI-generated videos, which allowed us to quickly test various creative concepts and at the same time significantly speed up production. I particularly appreciate their ability to find functional solutions, work with data, and continuously optimize the entire process. The result was not only interesting creatives, but above all better campaign performance and room for the further development of this area.
The main benefit of this approach is significant time savings, low production costs and the ability to iterate creative concepts quickly. At the same time, however, the system also has its limitations. The greatest of these is the limited creative control available to a professional editor. Instead of working with the image, most of the effort is channelled into working with the text and ensuring its accuracy. Creative decision-making is partly delegated to models, which have their own limitations when it comes to interpreting reality.
But this also opens up new possibilities. It is no longer simply a matter of ‘knowing how to edit video’, but of being able to define precisely what the end result should look like, and iteratively bringing it closer to reality.
Want to create effective campaigns using genAI?
Related articles
Development at TRITON IT has long been based on a Linux environment and infrastructure, which we build to handle the entire lifecycle of digital...
Content marketing is undergoing one of the biggest changes of the last twenty years. Until recently, most companies focused primarily on traditional...
Czech Travel Agency, the operator of the Lázně Travel portal, is one of the leading players in the Czech market for spa and wellness breaks....