📊 Key Data
  • 200ms latency: Loqua's system achieves near-instantaneous response time for natural voice input.
  • Multimodal processing: Combines audio, screen context, and app rules to resolve ambiguities in dictation.
  • Target users: Aims to assist developers, writers, project managers, and professionals with accessibility needs.
🎯 Expert Consensus

Experts would likely conclude that Loqua's 'voice plus vision' approach represents a significant innovation in dictation technology, though its long-term success hinges on real-world performance and clear privacy policies.

5 days ago
Voice Typing Gets Eyes: Loqua Launches AI That Sees Your Screen

Voice Typing Gets Eyes: Loqua Launches AI That Sees Your Screen

WILMINGTON, Del. – July 15, 2026 – For decades, the holy grail of dictation software was perfect transcription. Today, with near-human accuracy now table stakes, a new startup is arguing that the industry has been solving the wrong problem. Loqua, which officially launched its voice typing tool for Mac and Windows, claims the future isn't a better ear, but a pair of eyes.

The company is tackling what it calls the “destination problem.” A perfectly accurate transcript is useless if it’s contextually wrong. The phrase “add a guard before fetch profile” means something entirely different in a code editor versus a Slack message. Traditional dictation can’t tell the difference. Loqua’s solution is a multimodal system that combines voice with vision, analyzing not just what you say, but where you’re saying it.

“The truth the benchmarks hide is this: a transcript can be perfectly accurate and still be completely wrong,” the company stated in its launch announcement. This marks a potential paradigm shift, moving the contest from pure transcription accuracy to a more nuanced field of contextual intelligence.

Beyond Hearing: The Technology of ‘Voice That Sees’

Loqua’s core innovation is its multimodal architecture, which processes three signals simultaneously the moment a user speaks. The first is the traditional audio path, which transcribes the words. The second, and most crucial, is a context path that acts as the system’s “eyes.” It reads a small slice of the screen: the active application, the specific input field, any selected text, and the syntax of the surrounding words.

The third signal is an app path, which understands the destination’s rules—whether it requires Markdown, strict code syntax, or plain prose. By integrating these three local signals, the tool aims to resolve ambiguities that have long plagued dictation. Homophones like “cache the auth client” versus “cash the auth client” can be distinguished by the file type you’re working in. Spoken function names like “fetch profile” are correctly formatted as fetchProfile because the system sees the identifier on the screen.

To make this practical for daily use, the company claims its system runs primarily on-device, leveraging native hardware acceleration to achieve an end-to-end latency near 200 milliseconds. This speed is critical for making voice input feel natural. Early users seem to agree. “The words show up before I’ve even finished speaking,” one testimonial noted. “The zero-latency feel makes me forget the tool is even there.” While the company’s documentation mentions a “hybrid voice typing stack” that uses optional cloud processing, the user-facing experience is designed to feel instantaneous, avoiding the lag common with purely cloud-based services.

A New Tool for the Desktop Professional

The team behind the tool, comprised of engineers and AI researchers, is explicitly targeting desktop professionals who spend their day navigating a complex web of applications. For software developers, the promise is the ability to dictate code, comments, and commit messages with correct formatting. One user, a developer, praised its uncanny accuracy with technical jargon: “A screen full of framework names, library names, acronyms—Loqua nailed every single one. I’ve stopped going back to proofread.”

For writers and project managers, the benefit lies in seamlessly switching between drafting documents in Notion, sending quick updates on Slack, and composing formal emails. The system automatically adapts its output, turning a spoken command into a formatted task in a project tracker or a casual sentence in a chat app. This ability to “think out loud” productively is what the company believes will finally bridge the gap between solved transcription and underused dictation.

The tool is also positioned as a powerful accessibility aid. For professionals with repetitive strain injuries (RSI) or other physical limitations that make typing difficult, a context-aware voice tool can be transformative. “It didn’t just improve my workflow — it made work possible again,” reported one user with RSI, highlighting a critical use case beyond pure productivity enhancement.

The Privacy Paradox of a Seeing AI

Granting an application the ability to “see” your screen, even a small part of it, inevitably raises significant privacy questions. Loqua appears to have anticipated this, building its privacy story directly into its architecture. The company’s marketing emphasizes that its context-awareness is narrowly focused. “Loqua understands your current writing context in real time,” the press release states, adding that it “doesn’t OCR remote content, summarize windows you aren’t typing in, or keep a visual history.”

On its website, the company makes a bold claim: “Your words are processed ephemerally and never stored.” However, a review of its official privacy policy, last updated May 26, 2026, introduces a layer of ambiguity. The policy states that the company “may collect and process content you provide, including: Voice recordings; Audio input; Speech data; [and] Text input.”

This apparent contradiction between the promise of ephemeral processing and the stated collection of voice and text data is a critical point of friction. It highlights a central challenge for the entire category of context-aware AI: earning user trust. While the functionality is compelling, users, especially professionals handling sensitive corporate or client data, will require absolute clarity on how their information is handled. Reconciling its marketing promises with its legal policies will be a crucial next step for the startup.

Navigating a Crowded Field

Loqua enters a market that is rapidly heating up. The concept of “context-aware” voice input is no longer novel. Competitors like Voiskey, which launched just a week prior, also promise to tailor output for code, Slack, and Gmail using “Expression Intelligence.” Other tools like Contextli, Wispr Flow, and Willow are all vying for the same space, touting similar features of high accuracy, low latency, and intelligent formatting.

In this increasingly crowded field, Loqua is betting that its “voice plus vision” approach provides a deeper, more robust form of context than competitors who may rely more heavily on linguistic analysis alone. The company’s future roadmap, which hints at exploring reinforcement learning to handle rare terms and using audio cues like laughter as structural input, suggests a long-term vision focused on creating a truly natural and intuitive interface between human thought and the machine.

Ultimately, the success of Loqua will depend not just on the cleverness of its technology, but on its real-world performance in the messy, multifaceted workflows of modern professionals and its ability to provide unequivocal transparency about its data practices.

Topics & Related

Sector:
AI & Machine Learning
Software & SaaS
Event:
Product Launch
Theme:
Artificial Intelligence

📝 This article is still being updated

Are you a relevant expert who could contribute your opinion or insights to this article? We'd love to hear from you. We will give you full credit for your contribution.

Contribute Your Expertise →
UAID: 43162