Create Engaging Videos with OpenClaw
I used OpenClaw to generate video and audio, stitch clips, and return landscape and square versions through Telegram.
Tencent · Singapore · 27 March 2026
Overview
I wanted to move from asking for a video prompt to receiving the finished clip in chat. OpenClaw handled my requests through a preconfigured skill connected to Tencent Cloud Media Processing Service.
In the live private beta demo, I worked through a stylized video concept, generated audio, and asked for changes to the finished output. I received clips with audio, a landscape version, and a square crop of the stitched portrait video. The transcript covers 26-27 March.
Media Workflows
Generation
I generated video segments and audio, including an animated interlude between the opening and closing clips.
Editing and delivery
I asked the agent to stitch the clips, add a consistent audio bed, convert a video to landscape, and make a 1:1 square crop of the assembled portrait video.
Architecture and Tools
I kept media processing in Tencent Cloud and used OpenClaw to submit jobs, track progress, and retrieve the outputs. Generation took time, so the agent waited for segments to finish before assembling them and sending the result into Telegram.
For the remake, the agent generated the opening, middle, and closing segments sequentially to avoid concurrency issues, then stitched them together.