Use Cases · 9 min read · July 30, 2026
10 Voice Commands That Let You Run Your Entire SaaS Stack Without Touching Your Phone
Every time you unlock your phone to check Slack, your brain pays a tax you never see on your calendar. UC Irvine researcher Gloria Mark found that knowledge workers need an average of 23 minutes and 15 seconds to fully refocus after a single interruption — and Harvard Business Review estimates that workers toggle between apps over 1,200 times per day [1][2]. Aurex was built to end that cycle: speak one sentence, execute across your entire SaaS stack, and never break the thread of thought you're actually paid to maintain.
- Speed over everything: Aurex targets sub-second voice-to-voice response — under 1.2 s on the first turn, under 1 s on follow-ups — so the interaction feels like talking to a fast colleague, not waiting on a loading spinner.
- One voice, twelve services: Gmail, Slack, Linear, Asana, Notion, Google Calendar, Drive, Mercury, Typefully, and your internal Kiloforge servers are all reachable from a single spoken command.
- Read and write: Aurex doesn't just surface information — it sends replies, creates tickets, updates statuses, and moves money (with a spoken confirmation gate on high-stakes actions).
- Truly hands-free: Push-to-talk and wake-word modes both supported; barge-in lets you cut the assistant off mid-sentence when you already have what you need.
- No new integrations: Every connection reuses your existing MCP server credentials — no re-authorizing the same Gmail account for the fourth time in a new tool.
- Context switching eliminated: Instead of pulling up five apps on a commute, you fire off five spoken commands and stay in the moment.
| Dimension | Traditional approach | Aurex |
|---|---|---|
| Check Slack DMs | Unlock → app → scroll | One spoken command |
| Create a Linear ticket | Unlock → app → form → submit | One spoken command |
| Draft & send email | Unlock → Gmail → compose → send | One spoken command |
| Update a Notion page | Unlock → app → find page → edit | One spoken command |
| Confirm a payment | Unlock → Mercury → navigate → confirm | Spoken command + confirmation gate |
| Average latency | 45–90 s per task [2] | < 1.2 s perceived voice-to-voice |
TL;DR: Aurex turns your entire SaaS stack into a voice-operated command line — ten commands below show exactly how, across the services you already use every day.
The Context-Switching Tax You're Paying Right Now
Before we get to the commands, it's worth understanding what the alternative actually costs. The numbers are worse than most people guess.
23 Minutes Per Interruption
Gloria Mark, professor at UC Irvine, spent years observing knowledge workers in their natural environment and found that the average worker is interrupted every 3 minutes and 5 seconds [3]. Each interruption carries a recovery cost — 23 minutes and 15 seconds of refocus time — that accumulates quietly across the day [1]. Context switching has been shown to reduce cognitive efficiency by up to 40% [3].
The mechanism researchers call "attention residue" makes this worse: when you switch from deep work to check a Slack notification, part of your cognitive bandwidth stays behind on the original task even after you've moved on [1]. You're not executing Task B at full capacity; you're executing it with a degraded processor.
"Harvard Business Review estimates knowledge workers toggle between applications and websites 1,200 times per day, costing roughly four hours of productive time per worker per day." — Context Switching Kills Creative Productivity, MTM Video [1]
Why Phones Make It Worse for Operators
For operators and founders who live across a dozen SaaS tools, the phone intensifies the problem. Every check-in is a context switch: unlock, navigate, read, navigate back, re-lock, re-orient. Microsoft's own research adds a sharper edge — the typical knowledge worker spends less than three minutes on a digital screen before switching to something else [1]. The screen itself has become a switching machine.
The voice-native model inverts this. You don't switch to a new context; you extend your current one by speaking. The command goes out, the response comes back, and you're still where you were mentally — walking to a meeting, in a cab, on a treadmill.
The 10 Commands: Read Operations
Read commands are where Aurex earns its daily-driver status. They replace the most repetitive app-opening habit in an operator's day.
1. "What email came in from investors this morning?"
This routes to the Gmail MCP server, which supports real read/write access — not just search [5]. Aurex scans your inbox, filters by sender domain or label, and reads back a summary. You get the gist in 15 seconds without unlocking your phone. On a first turn this resolves in under 1.2 s of voice-to-voice latency; by the second turn of the same session, it's under 600 ms.
What it replaces: Open Gmail → tap Search → type investor name → read thread.
2. "Any new Slack DMs I haven't seen?"
MCP servers for Slack support channel reads and direct message queries, making it straightforward to surface unread DMs or mentions [6]. Aurex reads the sender name and first sentence of each unread DM so you can decide what warrants a real response. No unlock required.
What it replaces: Unlock → Slack → DMs → scroll.
3. "What's on my calendar this afternoon?"
Via the Google Calendar MCP connection, Aurex queries your events by time range and reads them back with title, time, and location [6]. A typical response is two sentences. Ask it naturally: "Do I have anything between 2 and 5?" or "What time is my next call?"
What it replaces: Unlock → Calendar → scroll to today → identify time window.
4. "What shipped in Linear this week?"
MCP servers for project management tools like Linear let an AI agent query sprint progress, filter by status, and pull context about ongoing work [5]. Ask for shipped issues, open blockers, or issues assigned to a specific teammate. Aurex reads back a short list — title and status — with no manual dashboard navigation.
What it replaces: Open Linear → navigate to board → filter by week → scan.
5. "Give me a summary of the KombuVault Notion doc"
The Notion MCP server provides genuine read/write access to your databases and pages [5]. For a read, Aurex fetches the page content and summarizes it on the fly. Long-form documents get condensed to the key points you actually need while walking. The voice recognition market hit $18.39 billion in 2025 and is projected to reach $61.71 billion by 2031 — the infrastructure making this kind of real-time document query practical at speed is now industrial-grade [4].
What it replaces: Open Notion → search for doc → read → summarize mentally.
The 10 Commands: Write Operations
Write commands are where Aurex crosses from interesting demo to genuine leverage. The confirmation gate on high-stakes actions keeps the power from becoming a liability.
6. "Reply to Sarah's last email — tell her I'll call at 3"
The Gmail MCP server supports composing and sending messages [5]. Aurex reads back the proposed reply before sending: "Replying to Sarah Chen: 'Hi Sarah, I'll give you a call at 3pm. Talk then.' Shall I send it?" Say "yes" and it's gone. This is the value read-back + confirmation gate pattern: any command containing exact values (names, times, amounts) gets confirmed before execution.
What it replaces: Open Gmail → find thread → reply → type → proofread → send.
7. "Create a Linear ticket: 'Raghav to review the paywall — P1'"
MCP servers for Linear let an AI agent create tickets and update statuses [5]. Aurex parses the title, infers the project from context or asks if ambiguous, sets the priority, and confirms: "Creating: 'Raghav to review the paywall' — Priority 1, in the Paywall project. Confirm?" One spoken "yes" and the ticket exists.
What it replaces: Open Linear → New Issue → title → priority → assignee → project → save.
8. "Mark the KombuVault PR as merged in Linear"
Status update commands are read-then-write: Aurex finds the issue by name, confirms it found the right one, and updates the status. Multi-step actions are narrated in real time — "Pulling up KombuVault PR… found it, marking merged now" — so you're never waiting in silence on a tool call [6]. This narration pattern is a deliberate UX decision: the voice loop stays active even during MCP round-trips.
What it replaces: Open Linear → search issue → click → update status → save.
9. "Post to Slack #engineering: deploy is live"
Slack MCP integration supports channel writes — sending a message to a specific channel by name [6]. Aurex confirms the channel and message text before posting. No channel lookup, no emoji debate. One confirmation, done.
What it replaces: Open Slack → find channel → compose → send.
10. "Send $500 to contractor — Alex Rivera — from Mercury, note: design invoice"
Mercury is connected via MCP, but payment commands carry the strictest confirmation gate in Aurex's safety model. Aurex reads back the full action: "Sending $500 to Alex Rivera from your Mercury operating account, memo: design invoice. Confirm?" Only a clearly spoken "yes" executes. Mis-heard values — a wrong amount, wrong recipient — are caught before they matter.
What it replaces: Open Mercury → Payments → recipient → amount → memo → review → submit.
How Aurex Keeps Latency Below One Second
The ten commands above are only useful if they're fast. An assistant that takes three seconds to respond isn't replacing app-opening — it's adding a new layer of friction. Aurex's architecture makes a single design bet to avoid that outcome.
Speech-to-Speech, Not a Pipeline
Most voice tools are actually three tools stitched together: a speech-to-text model, a language model, and a text-to-speech synthesizer. Each handoff adds latency. Aurex uses the OpenAI Realtime API (gpt-realtime-2.1-mini) as a single speech-to-speech brain — reasoning, tool-calling, and voice synthesis all happen inside one model loop. This eliminates the extra transcription and synthesis hops that roughly double the latency of a stitched pipeline. Read the detailed comparison at /blog/openai-realtime-api-vs-stt-llm-tts-pipeline-comparison.
The numbers: < 1.2 s glass-to-glass on the first turn, < 600 ms on follow-up turns within the same session. By 2026, voice AI has crossed from experimental to essential — the infrastructure is now fast enough to feel like a real conversation, not a voice-controlled menu [4].
MCP Servers Attach Directly to the Voice Loop
The OpenAI Realtime API supports remote MCP servers natively. This means your Gmail, Slack, Linear, Notion, and the internal Kiloforge MCP servers attach directly to the model session — there is no separate orchestration service, no extra routing hop, no custom integration layer to maintain [5][6]. Tool selection and execution happen inside the same reasoning loop that's already generating the voice response.
For a deeper look at how MCP servers work as a universal toolchain connector, see /blog/ultimate-guide-mcp-servers-voice-interface-toolchain.
The Token Broker: Security Without Compromise
A single lightweight backend component — a Cloudflare Worker or Vercel function — mints short-lived session tokens and holds the OpenAI API key server-side. Your credentials never touch the device. MCP server credentials are also held server-side. The phone only ever sees an ephemeral token that expires when the session ends. High-stakes write commands (payments, mass sends, deletes) get an additional spoken confirmation gate so mis-heard values can't trigger irreversible actions. For a detailed breakdown of push-to-talk vs. wake-word activation patterns and their security tradeoffs, see /blog/push-to-talk-vs-wake-word-activation-hands-free-mobile-2026.
| Command type | Latency target | Confirmation gate |
|---|---|---|
| Simple read (calendar, Slack DMs) | < 600 ms (subsequent turns) | None |
| Complex read (Linear sprint, Notion doc) | < 1.2 s | None |
| Low-stakes write (ticket, Slack message) | < 1.2 s | Read-back + "yes" |
| High-stakes write (payment, delete, mass send) | < 1.5 s (includes confirmation turn) | Explicit spoken confirmation required |
"87.5% of builders are actively building voice agents, not just researching them — according to the 2026 Voice Agent Report. That's not curiosity. That's commitment." — AssemblyAI, Voice AI in 2026 [4]
Who This Is Built For
Aurex is not a general-purpose chatbot, a customer support agent, or a voice-controlled smart speaker. It is a command-and-control layer for a single operator who already lives inside a dozen SaaS tools and is frequently mobile — traveling between meetings, walking between offices, commuting.
The profile that gets the most out of it:
- You have Gmail, Slack, Linear (or Asana), Notion, Google Calendar, and Mercury already in your daily workflow.
- You regularly find yourself pulling your phone out mid-commute to check "just one thing" — and that check turns into five minutes.
- You've felt the friction of a voice tool that needs two seconds to respond: that half-beat of silence is enough to break the habit.
- You want writes, not just reads: you want to fire off a reply, create a ticket, and send a payment without opening a single app.
The voice recognition market is undergoing a structural shift — $18.39 billion in 2025 on its way to $61.71 billion by 2031 — and the operator-layer use case is one of the clearest applications driving that growth [4]. Voice is faster than typing for many people, and for mobile, on-the-go operators, it's not even close.
If you're ready to run your stack with your voice, Aurex is available now on iOS — connect your first MCP server in under five minutes and fire off your first command before you reach your next meeting.
Frequently asked questions
How many SaaS tools can Aurex connect to at once?▾
Aurex connects to all your existing MCP-compatible services simultaneously — Gmail, Slack, Linear, Asana, Notion, Google Calendar, Drive, Mercury, Typefully, and internal Kiloforge servers are all attached to the same voice session. There's no limit enforced per session, though keeping MCP tool schemas lean ensures tool-selection reasoning stays fast.
How fast does Aurex respond to a voice command?▾
Aurex targets under 1.2 seconds of perceived voice-to-voice latency on the first turn of a session, and under 600 milliseconds on follow-up turns. This is achieved by using the OpenAI Realtime API as a single speech-to-speech model rather than a stitched STT→LLM→TTS pipeline, which roughly doubles latency due to extra transcription and synthesis hops.
Is Aurex safe to use for financial commands like sending money through Mercury?▾
Yes — high-stakes write commands (payments, deletes, mass sends) are protected by a mandatory spoken confirmation gate. Aurex reads back the exact action and values before executing, and only proceeds after you clearly say 'yes.' Your API credentials and MCP server credentials are held server-side and never stored on your device.
Does Aurex work without an internet connection?▾
No. Aurex requires an active internet connection because it routes voice audio to the OpenAI Realtime API session over WebRTC. Offline mode is not included in v1. For the best latency, a strong Wi-Fi or 5G connection is recommended.
What's the difference between push-to-talk and wake-word activation in Aurex?▾
Push-to-talk requires holding a button while speaking, giving you precise control over when the microphone is active — ideal for noisy environments or privacy-sensitive situations. Wake-word activation listens continuously for a trigger phrase, enabling fully hands-free operation. Both modes support barge-in, which lets you interrupt the assistant mid-response the moment you have the information you need.
Does Aurex support team or multi-user access?▾
Not in v1. Aurex is designed as a single-operator command layer — one person, all their tools, maximum speed. Multi-user and team features are explicitly out of scope for the current release.
Sources
- Context Switching Kills Creative Productivity: The Real Cost
- The Hidden Cost of Context Switching — BasicOps
- Context Switching: The Hidden Productivity Killer — 10000Hours Blog
- Voice AI in 2026: Inside the companies and investments shaping the future of speech — AssemblyAI
- How to Connect Claude Code to Notion, Gmail, and Other Apps Using MCP Servers — MindStudio
- MCP integration for AI: How easy it is to connect your tools with language models — novalutions
- Best Voice AI Productivity Tools 2026 — Speechify
- Best AI voice assistants in 2026: 15 tools tested and compared — Guideflow Blog
Keep reading
Ready to see it for yourself?
Back to home →