In-house product · Language & audio

Verbison turns selected website content into multilingual audio.

Verbison creates editorially controlled audio versions for websites. Operators choose which content is spoken, how it is phrased, and which languages and voices are published. Visitors listen directly within the embedded website experience.

Public websiteverbisonai.com
ProWebSolutions role

Product concept, website integration, translation and text-to-speech pipeline, and ongoing operation by ProWebSolutions.

  • LLM
  • Text-to-Speech
  • Multilingual Content
  • AI Infrastructure
verbisonai.com
Public Verbison homepage with an audio player and orange wave artwork
Public, logged-out view captured on 16 September 2026.

An editorial listening version, not generic read-aloud

Verbison is not designed as a generic screen reader. Website operators decide which passages become part of the listening version. An administration mode inside the website allows them to select, review and edit the content. The published audio therefore remains a deliberate editorial decision.

This control matters when written page elements need to be shortened, explained or structured differently for listening. The audio version can be closer to an edited publication than a word-for-word rendering of every visible element.

Translation and speech synthesis as separate stages

For multilingual versions, approved content is translated first. Verbison uses a self-hosted language model with around 80 billion parameters in infrastructure located in Frankfurt. A second stage generates the audio through controlled text-to-speech infrastructure using the selected voice.

Longer material is divided into suitable segments. A broker, workers and queues distribute processing, allowing languages to run in parallel and as a pipeline. Separating editorial state, translation and speech generation means an individual stage can be repeated without restarting the complete workflow in an uncontrolled way.

Synchronised playback inside the website

Verbison is embedded into a website through JavaScript. During playback, the translated spoken text can be displayed while the corresponding section of the original page is highlighted. The listener can therefore see which part of the page belongs to the current audio segment.

This synchronisation connects the player, text segments and page elements. It is useful for people who are less confident reading a language, prefer listening or benefit from using text and audio together. The website remains the publishing surface; Verbison adds a controlled listening layer.

Publishing new versions without interruption

When an editor changes content, voice or language, the existing publication does not need to disappear immediately. The currently approved audio remains available until the new version has been processed completely and is ready. Only then does it replace the previous version.

Verbison combines content selection, multilingual AI processing, controlled speech synthesis and synchronised web integration. The complete publishing workflow is the key capability: editorial teams retain control while compute-intensive processing is organised transparently in the background and the existing listening version remains available.

Relevant engineering services

The capabilities demonstrated by this project.