How DeepL Built a Translation Powerhouse with AI with CEO Jarek Kutylowski
Summary
DeepL’s opening came from entering translation during the 2017 neural reset, when prior approaches had to be broadly discarded. Jarek Kutylowski says specialization enabled architectures that balance source fidelity with native-sounding output. DeepL uses pretrained models, adding substantial compute and curated multilingual data rather than training everything from scratch.
DeepL’s advantage is as much systems engineering as research, a split Kutylowski estimates at “maybe like 50/50.” It began building data centers and software frameworks in 2017 because suitable GPU compute was unavailable; compute costs have since grown substantially with NVIDIA’s DGX generation and Blackwell. Company growth and revenue streams correlated with compute needs, while using additional GPUs would eventually also require “more researchers, more brains basically.”
Translation quality remains valuable because each improvement can unlock a more valuable and risk-sensitive workflow. A casual colleague email can tolerate imperfections, but contracts or terms and conditions published in 20 languages carry legal exposure; reducing a paralegal’s post-editing creates “a really big return on investment.” DeepL therefore injects document context, terminology and customer information rather than training separate models for its hundreds of thousands of customers.
AI will “severely reduce” content translated exclusively by human translators, though Kutylowski expects humans to remain in compliance-heavy life sciences and financial workflows. DeepL uses thousands of translators for training, feedback and quality assurance, but “not for production, not for inference.” Models are more reliable and accurate in a sense than people at avoiding ordinary slips, yet still lack human-level understanding of intent and can fail on ambiguous, broken or unusually short source text.
Speech translation is DeepL’s newer market, with latency—not preservation of vocal style—the immediate product priority. Kutylowski found translated customer conversations in Japan “pretty damn near” direct participation compared with waiting for an interpreter. Speech-recognition errors and conversational grammar make the input harder, while company terminology and proper names remain important for quality; real-time translation could broaden participation in international business.
As general-purpose LLMs improve, DeepL’s defense must move up-stack from sentence conversion into enterprise workflow. Kutylowski wants models to understand review processes, earlier AI translations and subsequent human edits, then feed those signals back into the model. The strategic challenge is selecting the highest-value workflows from a highly horizontal product rather than deepening every translation use case equally.
Deep dive
Not yet available upstream; scheduled sync will retry.