After shaping tools at OpenArt and Fish Audio, Zhang is focused on a harder question: how AI can sound more human without losing the human
A polished image, a smooth video, or a realistic voice can still feel empty if it does not carry the intent of the person behind it. That gap sits at the center of Helena Zhang’s current work. After helping build creative AI products at OpenArt and Fish Audio, Zhang is now focused on one of the harder problems in the field: helping AI understand what people mean, how they speak, and why a piece of content should feel specific rather than generic.
The problem is becoming more visible as creative AI becomes more capable. Models can generate images that look real. They can create videos with increasing accuracy. They can clone voices from short audio clips and produce speech across languages. Yet much of the output still carries the same flaw. It can look finished without feeling personal.
Zhang sees that limitation clearly.
“A model can produce something that resembles the training data,” Zhang says. “That does not mean it understands the person asking for it, the audience receiving it, or the emotion the creator wants to reach.”
That distinction has shaped her career in creative AI. At OpenArt, where Zhang joined as the second hire and head of growth, she worked during a period when image and video generation were moving from curiosity to serious creative infrastructure. The company grew from zero to 60M ARR and worked with leading models including Seedream, Nano Banana 2, GPT Image 2, Seedance 2.0, Google Veo 3, and Kling. For Zhang, the work was about more than putting advanced models in front of users. It was about helping people move from impressive outputs to intentional creative decisions.
Her tutorials at OpenArt made that philosophy practical. Zhang grew the company’s YouTube channel from zero to 200K subscribers in a year by showing people how to use fast-changing AI capabilities in real creative situations. She taught users how to train Flux Dev LoRAs to capture a style or character, how to use AI image editing and inpainting, and how to create consistent characters with Kling 1.6 Images to Video without training a LoRA.
Those examples matter because they point to a larger belief. Zhang does not treat creativity as an accident that happens when someone gets a good output. She treats it as a process shaped by context, judgment, and the creator’s sense of what belongs.
“Good creative tools should not flatten everyone into the same voice,” Zhang says. “They should help people make work that carries their own timing, language, and point of view.”
That idea became even more important during her work at Fish Audio, where she was a cofounder. Fish Audio launched open-weight AI text-to-speech models, including Fish Audio S1 and S2, that supported more than 80 languages. Users created more than 5 million voice clones on the platform. The work gave Zhang a close view of how voice technology can expand creative expression, but it also sharpened her awareness of what sound alone cannot solve.
A voice clone can reproduce tone. It can capture pace, texture, and recognizable qualities in speech. What it cannot automatically know is context. It does not know why one phrase feels natural in one community and awkward in another. It does not understand when a joke lands, when a hook feels forced, or when a piece of language carries the wrong social signal.
Zhang’s current focus is aimed at that deeper layer. She is working on AI that listens to real user and customer conversations, studies intent and psychology, and learns from word choice. Her goal is to move creative AI beyond the surface of realism and toward content that feels rooted in the people it is meant to serve.
“Language carries culture,” Zhang says. “The words people use tell you what group they belong to, what moment they are living in, and how they connect with each other. If AI ignores that, the content may be technically correct and still feel wrong.”
She points to vernacular as one example. Phrases can spread quickly through a group, generation, or online space. Terms may not make sense when separated from the people who use them, but inside a particular context, they carry meaning. For Zhang, that is not a minor detail. It is part of the work required to build AI that can support real creative expression.
Her next phase also involves agents that think on longer horizons. Instead of generating one piece of content in isolation, Zhang is interested in systems that execute multiple steps and learn from sparse real-world goals. One example is studying which types of opening hooks get genuine reactions from a specific target audience. The point is not to chase shallow engagement. It is to help AI understand that creative success often depends on the audience, the setting, and the emotional response a person is trying to create.
That focus separates her work from the louder argument that AI will simply replace human creativity. Zhang’s position is more demanding. If AI is going to become useful in creative work, it has to become better at understanding humans, not merely better at generating content that looks familiar.
“The future I care about is not AI speaking over people,” Zhang says. “It is AI helping people sound more like themselves, with more clarity and more power.”
Her credibility comes from building at the center of fast-moving creative AI companies, but also from the way she has translated technical tools for actual users. She judged at the 2025 MIT AI Filmmaking Hackathon, was invited to speak at the 2025 AI User Conference, and was invited to speak at SendPulse and AiMe Academy’s AI Marketing Day 2025. She also had four Top 5 Product of the Day launches on Product Hunt and was ranked the #65 most influential user on the platform in 2025.
Those achievements point to influence, but Zhang often returns to a simpler lesson. Creative AI depends on listening. Builders need to understand how people work, where they get stuck, and what kind of expression they are trying to protect. That might mean watching how experienced creators build full workflows. It might also mean learning from a 17-year-old in Colombia making strong AI videos on a phone without access to a laptop.
For Zhang, both users matter.
“People do not all create from the same place,” she says. “A great product has to respect that. You have to listen closely enough to understand the workflow, the constraint, and the ambition.”
That belief is guiding her toward a broader version of creative AI. The tools of the last few years made generation more accessible. The next challenge is making that generation more aligned with the person using it. Zhang’s goal is a version of AI that supports the person at the center of the work, learns from the way that person communicates, and helps turn intent into stronger creative expression.
Realism was one milestone. Authenticity is another. Helena Zhang’s work sits in the space between them, where AI begins to move past producing content and starts learning how to support a human voice.
