Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Join a Google Meet call, transcribe live captions, optionally speak in realtime, and do the followup work afterwards. Use when the user asks the agent to sit in on a meeting, take notes, summarize, respond in-call, or action items from it.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 42% | 0% |
The user says any of:
| Mode | What the bot does | |---|---| | transcribe (default) | Joins, enables captions, scrapes a transcript. Listen-only. | | realtime | Same as transcribe PLUS speaks into the meeting via OpenAI Realtime. The agent calls meet_say(text) and the bot's voice comes out of the call. |
Pick realtime only when the user actually wants the agent to speak. It costs real money (OpenAI Realtime is pay-per-audio-minute) and requires a virtual audio device set up on the machine running the bot.
| Location | When | |---|---| | Local (default) | Gateway machine runs the Playwright bot directly. | | Remote node (node="<name>") | Bot runs on a different machine that has a signed-in Chrome and (for realtime) a configured audio bridge. Useful when the gateway runs on a headless Linux box but the user's real signed-in Chrome lives on their Mac. |
Easiest path — run the built-in installer:
bashhermes plugins enable google_meet hermes meet install # pip deps + Chromium (transcribe only) hermes meet install --realtime # + pulseaudio-utils / brew blackhole+ffmpeg hermes meet auth # optional; skips guest-lobby wait hermes meet setup # preflight checks
hermes meet install --realtime prompts before running sudo apt-get (Linux) or brew install (macOS). Pass --yes to skip the prompt. It will NOT touch your macOS default-input setting — you have to select BlackHole 2ch in System Settings yourself before starting a realtime meeting.
Or do it manually:
bashpip install playwright websockets && python -m playwright install chromium # For realtime mode, additionally: # Linux: sudo apt install pulseaudio-utils # macOS: brew install blackhole-2ch ffmpeg # → System Settings → Sound → Input → BlackHole 2ch # Then set OPENAI_API_KEY or HERMES_MEET_REALTIME_KEY in ~/.hermes/.env
For a remote node:
bash# on the user's Mac (where Chrome is signed in): pip install playwright websockets && python -m playwright install chromium hermes plugins enable google_meet hermes meet node run --display-name my-mac # persistent server # copy the printed token # on the gateway: hermes meet node approve my-mac ws://<mac-ip>:18789 <token> hermes meet node ping my-mac # confirm reachable
Run hermes meet setup to preflight local prereqs.
meet_join(url=..., mode=..., node=...). Returns immediately.meet_status() for liveness, meet_transcript(last=20) for recent captions. Don't re-read the whole transcript every turn.meet_say(text="...") queues text for TTS. The speech lags by ~2s. Don't spam it.meet_leave() when done, or set duration="30m" on meet_join for auto-leave.meet_transcript() in full, summarize, and use regular tools to send the recap, file issues, schedule followups.| Tool | Parameters | Use | |---|---|---| | meet_join | url, mode?, guest_name?, duration?, headed?, node? | Start bot | | meet_status | node? | Liveness + progress | | meet_transcript | last?, node? | Read captions | | meet_leave | node? | Close bot | | meet_say | text, node? | Speak in realtime meeting |
node? on all tools: pass a registered node name (or "auto" for the sole node) to operate a remote bot instead of a local one. Omit for local.
hermes meet auth avoids this.HERMES_MEET_LOBBY_TIMEOUT env), the bot leaves and meet_status reports leaveReason: "lobby_timeout".meet_join leaves the first.meet_status().error.meet_say requires mode='realtime' on the originating meet_join. Calling it against a transcribe-mode meeting returns a clear error.response.cancel to OpenAI Realtime. Captions take ~500ms to show up, so the bot will talk over the first second or so of a human interruption.meet_status() returns (subset shown, there are more):
| Key | Meaning | |---|---| | inCall | Past the lobby. False while waiting for admission. | | lobbyWaiting | Clicked "Ask to join", waiting on host. | | joinAttemptedAt / joinedAt | Timestamps for lobby-click and actual admission. | | captioning | Caption observer is installed. | | transcriptLines / lastCaptionAt | Transcript progress. | | realtime / realtimeReady | Realtime mode provisioned / WS connected. | | realtimeDevice | Audio device name the bot is feeding (e.g. hermes_meet_src). | | audioBytesOut / lastAudioOutAt | How much PCM the OpenAI session has produced. | | lastBargeInAt | Timestamp of the most recent response.cancel sent. | | leaveReason | duration_expired, lobby_timeout, denied, page_closed, or null. | | error | Last error (soft — bot may still be running). |
Local:
$HERMES_HOME/workspace/meetings/<meeting-id>/transcript.txtRemote node: transcript lives on the node host's disk. Use meet_transcript(node=...) to read it over RPC.
https://meet.google.com/... URLs pass.$HERMES_HOME/workspace/meetings/node_token.json) and must be copied to the gateway via hermes meet node approve.meet_say text is rate-limited by the OpenAI Realtime session; spam-protection is the bot's problem, not yours, but still — don't queue hundreds of lines.Other measured skills in the registry, with their headline benchmark lift.