# pchat chat & voice API — device (ESP32) client guide The photonichat server gives a device a **talking market assistant**: hold a button, speak, release, and it answers out loud — with live prices, news, an English-tutor mode, and a private conversation history per device. This page is the contract for building that client; it complements the quote/clock API in [`device-api-agent-guide.md`](./device-api-agent-guide.md) (served at `/docs/device-guide`). This page is served at **`/docs/chat-api`**. All example bodies below were captured from the live server. --- ## 1. Basics | | | |---|---| | Base URL | `https://pchat.photonicat.com` (or `https://pchat-api.photonicat.com`, same API) | | TLS | Cloudflare certificates — on ESP-IDF use `esp_crt_bundle_attach` | | Auth | none | | Identity | header **`X-Client-Id`** on every `/chat/*` call except `/chat/models` and `/chat/tts` | | Bodies | JSON (UTF-8), except audio uploads/downloads | | Errors | `{"detail": "…"}` with a 4xx/5xx status (422 = malformed request) | **`X-Client-Id` is the device's identity, not a secret.** Each id has its own private chat history and recordings (another id gets `404` for them). Use the device serial number, e.g. `PRS-7ZZZZ-00001-K` — stable across reboots and reflashes, up to 64 characters of `A-Z a-z 0-9 _ -`. Optional `X-Client-Name` is stored with new chats for reference. **Say where the device is.** Optional **`X-Timezone`** (IANA name, e.g. `Asia/Shanghai`) and **`X-City`** (e.g. `Shanghai` or `上海`) on message / talk calls give the assistant the local date and time and the place "today's weather" means. Without them the server guesses from the connection's IP, which is wrong behind a VPN (the owner's traffic exits in San Jose). For a device (`X-Client-Id` = its `PRS-…` serial) the home city of the shared device settings (`/device/config`, the Prism's Settings → Home city with *Automatic location* off) is used before the IP, and a city the user told the assistant ("我在上海") is remembered and used before either (§4.6); `CHAT_HOME_CITY` on the server is the last resort before the IP. ### Limits | What | Limit | When exceeded | |---|---|---| | POST/PATCH/DELETE | 60 / min per client id | `429` + `Retry-After` | | GET `/chat/tts` | 60 / min per IP (shared by devices behind one NAT) | `429` + `Retry-After` | | other GETs | 240 / min per IP | `429` + `Retry-After` | | replies generating at once | 1 per chat, 2 per client id, 3 per IP | `409` (same chat) / `429` | | message text | 8000 characters | `413` | | audio upload | 8 MB (keep recordings ≤ 60 s) | `413` | | TTS text | 600 characters per request | `422` | Rate-limit `429`s look like `{"error":"rate_limited","tier":"w","retry_after":37}` — wait `Retry-After` seconds. ### Timing (so you can set timeouts) | Step | Typical | Worst case | |---|---|---| | `POST /chat/transcribe` | 1.5 s (GPU Whisper) | 5–8 s (GPU busy / fallback) | | `POST …/messages` | 2–5 s | **~70 s** when the language model is cold-loaded | | `GET /chat/tts` | ~1 s per sentence | ~5 s | Long waits are safe: a non-streaming reply sends a space every 15 s while it waits (see §4.3), so a **30 s per-read socket timeout** works everywhere. Don't set a short *total* timeout on the message call. --- ## 2. A push-to-talk turn (the whole device flow) ``` boot: sid = nvs_get("chat_sid") — or POST /chat/sessions and store its id (optional) GET /chat/models — see which models are online on hold: record 16 kHz, 16-bit, mono PCM into PSRAM (32 KB per second, cap 60 s) on release: 1. POST /chat/transcribe?lang=auto body = WAV → {"text", "audio_id", …} text == "" → "didn't catch that", stop 2. POST /chat/sessions/{sid}/messages {"content": text, "voice": true, "audio_id": id, "stream": false} → {"message": {"content": reply}, …} 404 → the chat is gone: POST /chat/sessions, store the new sid, retry once 3. for each sentence of reply: GET /chat/tts?text=&format=wav&rate=16000 → WAV → I2S (fetch sentence n+1 while n plays) ``` `"voice": true` makes the assistant answer in a few short spoken sentences with numbers as digits (no tables or markdown), so it reads well through TTS. The reply comes back **in the language of the question** (Chinese in → Chinese out); a bare follow-up like "BTC?" keeps the previous question's language. --- ## 3. Speech ### 3.1 `POST /chat/transcribe` — speech to text Body: the recording as-is. **Send a WAV: 16 kHz, 16-bit little-endian, mono** (`Content-Type: audio/wav`). WebM/Opus, Ogg, MP4/AAC and MP3 also work. | Query | Values | | |---|---|---| | `lang` | `auto` (default), `zh`, `en` | fixing the language avoids misdetection and adds a market-vocabulary hint (茅台, 上证指数, Nasdaq…) | | `model` | `large-v3-turbo` (default), `sensevoice`, `auto`, `base`, `tiny` | ids from `GET /chat/models` → `stt_models` | The default runs Whisper large-v3-turbo on the owner's GPU speech server; if that server can't be reached, it falls back to Whisper on this server's CPU transparently (`"fallback": true`). ```jsonc POST /chat/transcribe?lang=zh (Content-Type: audio/wav, 68 KB, 2 s) → 200 { "text": "黄金现在多少钱,", "language": "zh", "seconds": 2.02, "model": "large-v3-turbo", "engine": "gpu", // "local" when the CPU fallback was used "ms": 1623, "audio_id": 4 // absent when text is "" (nothing is stored then) } ``` - Clips shorter than ~0.4 s, or silence, return `"text": ""`. - So does noise that Whisper turned into words anyway — a segment "ending" past the end of the clip, an unsure foreign language ("Продолжение следует...", "Não sei bem"), or a subtitle credit from its training data ("请不吝点赞 订阅 转发 打赏支持明镜与点点栏目", "谢谢观看"). Those add `"why": "silence" | "language" | "phantom"` and `"discarded": ""`; show "didn't catch that". - The recording is saved on the server (private to your `X-Client-Id`). Pass `audio_id` with the message to attach it; unattached recordings are deleted after an hour. Deleting the message or chat deletes the recording. - `429 {"detail":"voice is busy — try again in a moment"}` when the CPU fallback is queued up; `400` for an empty or undecodable body. ### 3.2 `GET /chat/tts` — text to speech | Query | Values | | |---|---|---| | `text` | ≤ 600 characters (URL-encoded UTF-8) | send one or a few sentences | | `lang` | `zh`, `en`, or empty = detect (any CJK → zh) | picks the voice | | `format` | `mp3` (default, 24 kHz mono 48 kbps) or **`wav`** | | | `rate` | 8000–48000, default 16000 | WAV sample rate | Voices: Microsoft Edge neural voices — `zh-CN-XiaoxiaoNeural` for Chinese, `en-US-AvaMultilingualNeural` for English. No `X-Client-Id` needed. `format=wav` returns **16-bit little-endian mono PCM with a canonical 44-byte header** (`RIFF`/`fmt ` (16)/`data`) — skip 44 bytes and write the rest to I2S at `rate`: ``` GET /chat/tts?text=黄金现在是4157美元。&format=wav&rate=16000 → 200 content-type: audio/wav content-length: 111404 (3.48 s of audio) 00000000: 5249 4646 24b3 0100 5741 5645 666d 7420 RIFF$...WAVEfmt 00000010: 1000 0000 0100 0100 803e 0000 007d 0000 .........>...}.. 00000020: 0200 1000 6461 7461 00b3 0100 0000 0000 ....data........ ``` Split replies into sentences at `。!?!?;` or a `.` followed by a space/newline, merging short ones up to ~200 characters; request the next sentence while the current one plays. Responses are cached for a day (`Cache-Control: private, max-age=86400`) and the server also caches recent sentences, so replays are cheap. `502 {"detail":"speech service unavailable"}` if synthesis fails — skip speaking (the text is still valid). ### 3.3 `GET /chat/audio/{audio_id}` — play back your own recording Returns the upload exactly as sent (e.g. your WAV). `404` for other client ids. `GET /chat/audio/{audio_id}/info` analyses it — handy while tuning the mic: ```jsonc { "container": "wav", "codec": "pcm_s16le", "sample_rate": 16000, "channels": 1, "bits": 16, "bit_rate": 256000, "lossless": true, "duration": 2.02, "bytes": 64684, "peak_dbfs": -6.0, "rms_dbfs": -21.4, "clipped_pct": 0.0, "dc_offset": 0.0002, "waveform": [0.01, 0.2, …], // 200 points, 0-1 "wav": {"riff_size": 64676, "audio_format": 1, "channels": 1, "sample_rate": 16000, "byte_rate": 32000, "block_align": 2, "bits": 16, "data_offset": 44, "data_size": 64640, "chunks": [{"id": "fmt ", "offset": 12, "size": 16}, {"id": "data", "offset": 36, "size": 64640}]}, "warnings": [], // e.g. "very quiet (peak below −30 dBFS) — raise the mic gain", // "clipping on 2.1% of samples", "DC offset +0.05", "RIFF size … ≠ file length" "transcript": "黄金现在多少钱", "language": "zh" } ``` The same panel is on the web: ⓘ next to ▶ on a recording (Chat page, and Settings → Chat history for admins). Link your device on the Chat page (📟 Devices → "+ link", its serial) to see its conversations there, read-only. --- ## 4. Chat ### 4.1 `GET /chat/models` ```jsonc { "default": "huihui_ai/Qwen3.8-abliterated:latest", "models": ["huihui_ai/Qwen3.8-abliterated:latest", "qwen3.8:27b", "gemma4:31b"], "online": true, // false: the model server is down "stt_models": [ {"id": "auto", "name": "Auto", "gpu": true, "label": "now: turbo · GPU"}, {"id": "large-v3-turbo", "name": "turbo", "gpu": true, "label": "GPU · best"}, {"id": "sensevoice", "name": "SenseVoice", "gpu": true, "label": "GPU busy · CPU"}, {"id": "base", "name": "base", "gpu": false, "label": "local CPU"}, {"id": "tiny", "name": "tiny", "gpu": false, "label": "local · fastest"} ], "stt_default": "large-v3-turbo" } ``` ### 4.1b `GET /chat/health` — is the GPU up? No `X-Client-Id` needed. Cached ~10 s. What the Chat page's badges show; a device can light a status dot from `llm.level`: ```jsonc { "ok": true, // false = the LLM is unreachable "llm": {"level": "ok", "state": "warm", "label": "GPU · ready", "detail": "Qwen3.8-abliterated in VRAM (18.4 GB), kept loaded 7 h 57 min more", …}, "stt": {"level": "ok", "state": "gpu", "label": "GPU · turbo", "detail": "Whisper large-v3-turbo on the GPU (~1.5 s a question) · 5.2 GB VRAM free"}, "tts": {"level": "ok", "state": "ok", "label": "Edge · ok", "detail": "Edge neural voices: last spoke 2 min ago"} } ``` `level` is `ok` (green), `warn` (amber: works, but slower) or `down` (red). `llm.state`: `warm` (loaded, replies start at once), `partial` (partly on the CPU, slower), `cold` (not loaded — the next reply first loads it, ~16-27 s; say "waking up…"), `offline`. `stt.state`: `gpu`, `gpu-cold` (loads with the first question), `cpu` (the GPU was full: ~7 s a question), `cpu-cold`, `fallback` (speech server down: this server's CPU), `local`. `tts.state`: `ok`, `idle` (not used since the restart), `failing`. ### 4.2 Sessions (conversations) | Call | Body | Returns | |---|---|---| | `POST /chat/sessions` | `{}` or `{"model": ""}` | the new session | | `GET /chat/sessions` | — | `{"sessions": [session + "generating"]}`, newest activity first (max 300; older ones are pruned) | | `GET /chat/sessions/{id}` | — | session + `"messages"`, `"generating"`, `"partial"` (text so far while a reply is being written) | | `PATCH /chat/sessions/{id}` | `{"title": "…"}` | the session | | `DELETE /chat/sessions/{id}` | — | `{"deleted": id}` (messages + recordings too) | ```jsonc POST /chat/sessions → {"id": 11, "title": "", "model": "huihui_ai/Qwen3.8-abliterated:latest", "created_at": 1790652475.38, "updated_at": 1790652475.38} ``` A session's title is set from its first message. A chat holds up to 400 messages (then `409 "this chat is too long — start a new one"`); the model sees the last ~40 messages. ### 4.3 `POST /chat/sessions/{id}/messages` — ask | Field | Default | | |---|---|---| | `content` | — | the question (≤ 8000 chars) | | `voice` | `false` | spoken question → short, speakable answer | | `audio_id` | `null` | from `/chat/transcribe`, attaches the recording | | `mode` | `"auto"` | `auto`, `market`, `english` (see §5) | | `level` | `"auto"` | English tutor level: `auto`, `kids`, `middle`, `high`, `adult` (see §5) | | `model` | the chat's model | a `models` id | | `stream` | `true` | **`false` = one JSON object when done — use this on the device** | | `tz` | `""` | the client's IANA timezone (same as the `X-Timezone` header) | | `city` | `""` | where the user is (same as `X-City`) | | `stt_ms`, `stt_model` | — | voice: `ms` / `model` from `/chat/transcribe`, kept with the question | **`"stream": false`** — the response is `200 application/json`, possibly preceded by spaces (heartbeats every 15 s while the model works; any JSON parser skips them). Success: ```jsonc { "session": {"id": 11, "title": "黄金现在多少钱,", "model": "huihui_ai/Qwen3.8-abliterated:latest", …}, "user": {"id": 21, "role": "user", "content": "黄金现在多少钱,", "meta": "{\"voice\": true, \"audio\": 4}", …}, "mode": "market", "level": null, // tutor level when mode is "english" "tools": [{"name": "get_quotes", "args": {"symbols": ["GOLD"]}}], // live lookups it made "message": { "id": 22, "role": "assistant", "content": "黄金现在报每盎司4138.7美元,过去24小时下跌了1.363%。", // ← speak this "model": "huihui_ai/Qwen3.8-abliterated:latest", "meta": "{\"tokens\": 56, \"load_ms\": 64419, \"tps\": 129.8, \"ms\": 66231, …}", // JSON string, informational (timings below) "created_at": 1790652544.32 } } ``` Failure *after* the request was accepted comes back as `200` with an `error` key: ```jsonc {"error": "the model server is unreachable — try again in a moment", "dropped_user": true} ``` `dropped_user: true` means the question wasn't kept — ask again. Failures *before* generation starts use status codes: `400` empty content or missing `X-Client-Id`, `404` unknown session (or someone else's), `409` a reply is already being written for this chat, `413` too long, `429` too many replies in progress / rate limit. If the connection drops mid-wait, the reply is still generated and saved: poll `GET /chat/sessions/{id}` until `"generating": false` and take the last message. A reply's `meta` says where its time went (ms): `queue_ms` (waiting for a model slot), `load_ms` (model loaded into the GPU), `prompt_ms` / `prompt_tokens` (reading the conversation), `tool_ms` (live lookups), `first_ms` (until the first words), `ms` (total). Settings → Chat history shows the breakdown. A reply that looked things up is saved (and `done` carries it) as its final answer only: a line the model wrote before the lookups ("我来看看…") shows while streaming but isn't kept or spoken by `/chat/device/speech`. **`"stream": true`** (default) — Server-Sent Events (`text/event-stream`), one JSON object per `data:` line, events separated by a blank line; lines starting with `:` are keep-alive comments: ``` data: {"type": "start", "session": {…}, "user": {…}, "mode": "market", "level": null} data: {"type": "queued"} (waiting for a free slot) data: {"type": "tool", "name": "get_quotes", "args": {"symbols": ["BTC"]}} data: {"type": "delta", "content": "BTC"} data: {"type": "delta", "content": " 现在"} data: {"type": "reset", "content": "…"} (rare: replace ALL text so far with this — the model's hidden reasoning leaked and was removed) … data: {"type": "done", "message": {…same shape as above…}} or data: {"type": "error", "error": "…", "dropped_user": true, "content": ""} ``` Stream only if you want to start speaking before the reply is finished; otherwise `stream: false` is simpler (its `message.content` is always the final, cleaned text). ### 4.4 `POST /chat/sessions/{id}/stop` Stops the reply being written (keeps what was written, marked `"stopped": true` in its meta). → `{"stopped": true}` (or `false` if nothing was running). ### 4.5 `POST /chat/sessions/{id}/messages/{message_id}/feedback` 👍 / 👎 on a reply: `{"rating": 1}` or `{"rating": -1, "note": "wrong stock"}` (note optional, ≤ 500 chars); `{"rating": 0}` takes it back. → the message, its meta now holding `"feedback": {"rating", "at", "note"}`. Only for replies (`404` for a question or someone else's chat). Settings → Chat history can list just the chats with a 👎, to find answers that need fixing. ### 4.6 Memory: `GET /chat/memory`, `DELETE /chat/memory/{key}`, `DELETE /chat/memory` The assistant remembers what a user tells it about themselves, per `X-Client-Id` and across chats: the city they live in, their name, their English level / grade, a preference ("prefers short answers"). It saves them with its `remember` tool (shown as a `tool` event), the tutor also remembers a level the learner states, and every later chat starts with them. A remembered city is where "today's weather" is. At most 12 facts; the oldest go first. ```jsonc GET /chat/memory → {"memory": [{"key": "city", "value": "Shanghai", "updated_at": 1790690000.0}, {"key": "level", "value": "kids", …}]} DELETE /chat/memory/city → {"forgot": "city"} DELETE /chat/memory → {"forgot": 2} ``` --- ## 5. What the assistant does - **Live data.** It looks things up before answering — `tools` shows what it used: `get_quotes` (Hyperliquid / Sina / eToro: stocks, crypto, FX, commodities, indices incl. 上证指数 / 创业板指 / 恒生指数), `search_markets`, `get_movers` (today's top gainers / losers / most traded: A-shares, US stocks, crypto), `get_news` (per instrument, or general; `lang: zh` = Chinese sources, for China news) and `get_weather` (now + 3 days; where the user is unless a city is named — see `X-Timezone` / `X-City`). Prices carry their currency, % change and basis (24h or 1d). A mis-heard stock name is matched by sound (加力创 → 嘉立创) and the answer checks it ("你是说嘉立创吧?"). - **Modes.** `auto` routes each message: "teach me English", "…用英语怎么说", "correct my sentence" go to an **English tutor** (corrections, phrasing, pronunciation tips, practice); a real market question goes back to the market assistant, and so does asking it to look something up ("帮我查查…天气"); anything else stays in the chat's current mode. Force one with `"mode": "english"` or `"market"`. The reply's meta has `"mode"`. The model only sees the earlier turns of the mode it is answering in. - **Learner levels** (tutor only). The tutor teaches differently for **`kids`** (kindergarten/primary: tiny words, games, repeat-after-me, praise, strictly child-safe), **`middle`** (初中: textbook grammar in Chinese, 中考 practice), **`high`** (高中: 高考 vocabulary, clauses, writing) and **`adult`** (daily/work/travel English, IELTS/TOEFL/四六级). With `"level": "auto"` it's taken from what the learner says ("我是三年级小学生", "我上初二", "准备高考", "I'm 8", "IELTS"…) and then kept for the chat; if still unknown, the tutor adapts to how they write and asks their grade once. A device for a child can simply send `"mode": "english", "level": "kids"` on every message. - **Languages.** Replies follow the question's language (Simplified Chinese or English); the tutor explains in Chinese and keeps English examples in English. --- ## 6. ESP-IDF snippets (v5.x, `esp_http_client` + `cJSON`) ```c #include #include #include "esp_http_client.h" #include "esp_crt_bundle.h" #include "cJSON.h" #define PCHAT "https://pchat.photonicat.com" static esp_http_client_handle_t pchat_client(const char *url, esp_http_client_method_t m) { esp_http_client_config_t cfg = { .url = url, .method = m, .timeout_ms = 30000, /* per read/write; replies send heartbeats */ .crt_bundle_attach = esp_crt_bundle_attach, .buffer_size = 4096, .buffer_size_tx = 1024, }; esp_http_client_handle_t c = esp_http_client_init(&cfg); esp_http_client_set_header(c, "X-Client-Id", device_serial()); /* e.g. "PRS-7ZZZZ-00001-K" */ esp_http_client_set_header(c, "X-Timezone", "Asia/Shanghai"); /* optional: local time + weather place */ return c; } ``` **Upload a recording** (16 kHz mono int16 PCM in PSRAM) → text + audio_id: ```c typedef struct __attribute__((packed)) { char riff[4]; uint32_t riff_size; char wave[4]; char fmt[4]; uint32_t fmt_size; uint16_t audio_fmt, channels; uint32_t rate, byte_rate; uint16_t block_align, bits; char data[4]; uint32_t data_size; } wav_hdr_t; /* 44 bytes, little-endian like the ESP32 */ static void wav_hdr(wav_hdr_t *h, uint32_t pcm_bytes, uint32_t rate) { memcpy(h->riff, "RIFF", 4); h->riff_size = 36 + pcm_bytes; memcpy(h->wave, "WAVE", 4); memcpy(h->fmt, "fmt ", 4); h->fmt_size = 16; h->audio_fmt = 1; h->channels = 1; h->rate = rate; h->byte_rate = rate * 2; h->block_align = 2; h->bits = 16; memcpy(h->data, "data", 4); h->data_size = pcm_bytes; } /* returns 0 and fills text/audio_id; text[0] == 0 means "nothing heard" */ int pchat_transcribe(const int16_t *pcm, size_t samples, const char *lang, char *text, size_t text_len, int *audio_id) { char url[128]; snprintf(url, sizeof url, PCHAT "/chat/transcribe?lang=%s", lang ? lang : "auto"); esp_http_client_handle_t c = pchat_client(url, HTTP_METHOD_POST); esp_http_client_set_header(c, "Content-Type", "audio/wav"); wav_hdr_t h; wav_hdr(&h, samples * 2, 16000); int rc = -1; if (esp_http_client_open(c, sizeof h + samples * 2) == ESP_OK) { esp_http_client_write(c, (const char *)&h, sizeof h); const char *p = (const char *)pcm; for (size_t left = samples * 2; left; ) { int n = esp_http_client_write(c, p, left > 4096 ? 4096 : left); if (n <= 0) goto done; p += n; left -= n; } esp_http_client_fetch_headers(c); char body[1536]; int n = esp_http_client_read_response(c, body, sizeof body - 1); body[n > 0 ? n : 0] = 0; if (esp_http_client_get_status_code(c) == 200) { cJSON *j = cJSON_Parse(body); const cJSON *t = cJSON_GetObjectItem(j, "text"), *a = cJSON_GetObjectItem(j, "audio_id"); snprintf(text, text_len, "%s", cJSON_IsString(t) ? t->valuestring : ""); *audio_id = cJSON_IsNumber(a) ? a->valueint : 0; cJSON_Delete(j); rc = 0; } } done: esp_http_client_cleanup(c); return rc; } ``` **Ask** (non-streaming) → reply text: ```c /* returns HTTP status (200 ok; check reply[0]), 404 = create a new session */ int pchat_ask(int sid, const char *text, int audio_id, char *reply, size_t reply_len) { char url[96]; snprintf(url, sizeof url, PCHAT "/chat/sessions/%d/messages", sid); cJSON *req = cJSON_CreateObject(); cJSON_AddStringToObject(req, "content", text); cJSON_AddBoolToObject(req, "voice", true); cJSON_AddBoolToObject(req, "stream", false); if (audio_id) cJSON_AddNumberToObject(req, "audio_id", audio_id); char *body = cJSON_PrintUnformatted(req); esp_http_client_handle_t c = pchat_client(url, HTTP_METHOD_POST); esp_http_client_set_header(c, "Content-Type", "application/json"); int status = -1; reply[0] = 0; if (esp_http_client_open(c, strlen(body)) == ESP_OK) { esp_http_client_write(c, body, strlen(body)); esp_http_client_fetch_headers(c); status = esp_http_client_get_status_code(c); size_t cap = 16 * 1024; /* meta carries tool results */ char *buf = heap_caps_malloc(cap, MALLOC_CAP_SPIRAM); int n = esp_http_client_read_response(c, buf, cap - 1); buf[n > 0 ? n : 0] = 0; cJSON *j = cJSON_Parse(buf); /* skips the heartbeat spaces */ const cJSON *msg = cJSON_GetObjectItem(j, "message"); const cJSON *content = cJSON_GetObjectItem(msg, "content"); if (cJSON_IsString(content)) snprintf(reply, reply_len, "%s", content->valuestring); /* else: {"error": ...} (status 200) or {"detail": ...} (4xx) */ cJSON_Delete(j); free(buf); } esp_http_client_cleanup(c); cJSON_Delete(req); cJSON_free(body); return status; } ``` **Speak** one sentence (WAV → I2S): ```c static void url_encode(const char *in, char *out, size_t cap) /* UTF-8 safe */ { static const char hex[] = "0123456789ABCDEF"; size_t o = 0; for (const unsigned char *p = (const unsigned char *)in; *p && o + 4 < cap; p++) { if (isalnum(*p) || strchr("-_.~", *p)) out[o++] = *p; else { out[o++] = '%'; out[o++] = hex[*p >> 4]; out[o++] = hex[*p & 15]; } } out[o] = 0; } int pchat_say(const char *sentence, i2s_chan_handle_t tx) /* I2S set to 16 kHz, 16-bit mono */ { char *url = malloc(3 * 1024); int o = snprintf(url, 3 * 1024, PCHAT "/chat/tts?format=wav&rate=16000&text="); url_encode(sentence, url + o, 3 * 1024 - o); esp_http_client_handle_t c = pchat_client(url, HTTP_METHOD_GET); int rc = -1; if (esp_http_client_open(c, 0) == ESP_OK && esp_http_client_fetch_headers(c) >= 0 && esp_http_client_get_status_code(c) == 200) { uint8_t buf[2048]; int skip = 44, n; /* canonical WAV header */ while ((n = esp_http_client_read(c, (char *)buf, sizeof buf)) > 0) { int off = skip > n ? n : skip; skip -= off; size_t wrote; if (n > off) i2s_channel_write(tx, buf + off, n - off, &wrote, portMAX_DELAY); } rc = 0; } esp_http_client_cleanup(c); free(url); return rc; } ``` Tips: - **Keep the session id in NVS**; create a new one on `404`, or when the user asks for a fresh conversation. Many short chats are fine (300 per device are kept). - **Record at 16 kHz mono 16-bit.** An INMP441 gives 32-bit I2S frames: take the top 16 bits (`(int16_t)(s >> 16)`, add gain if quiet). 10 s = 320 KB → PSRAM. Ignore presses shorter than ~0.5 s. - **Language button?** `lang=zh` / `lang=en` makes recognition more reliable than `auto` for short phrases. - **Be polite with retries:** on `429` wait `Retry-After`; on network errors back off 2 → 4 → 8 s. One device should never have more than one reply in flight. - **Speak while fetching:** start `GET /chat/tts` for sentence *n+1* in another task while sentence *n* plays, so there are no gaps. --- ## 7. curl cheat sheet ```bash B=https://pchat.photonicat.com; H="X-Client-Id: PRS-7ZZZZ-00001-K" curl -s $B/chat/models SID=$(curl -s -X POST $B/chat/sessions -H "$H" -H 'Content-Type: application/json' -d '{}' | jq .id) curl -s -X POST "$B/chat/transcribe?lang=zh" -H "$H" -H 'Content-Type: audio/wav' --data-binary @q.wav curl -s -X POST $B/chat/sessions/$SID/messages -H "$H" -H 'Content-Type: application/json' \ -d '{"content":"黄金现在多少钱?","voice":true,"stream":false}' | jq -r .message.content curl -sN -X POST $B/chat/sessions/$SID/messages -H "$H" -H 'Content-Type: application/json' \ -d '{"content":"BTC?"}' # SSE curl -s -G $B/chat/tts --data-urlencode 'text=Gold is at 4,157 dollars.' \ -d format=wav -d rate=16000 -o say.wav curl -s $B/chat/sessions/$SID -H "$H" | jq '.messages[-1].content' ``` --- ## 8. One-call push-to-talk, one conversation per device (`/chat/device/*`) What the Prism firmware uses (`chat_device.py`). The server keeps the device's conversation, keyed on its hardware serial (`X-Client-Id`): the **current conversation** is the device's most recently active chat, so the device stores no session id and carries on after a reboot. After **30 minutes without a turn** the next `talk` starts a new conversation by itself (stale turns otherwise steered new answers; `CHAT_DEVICE_IDLE_S` on the server). A browser whose client id is the serial sees the same chats. | Call | | | |---|---|---| | `POST /chat/device/talk?lang=auto` | body = the recording: Ogg Opus (`audio/ogg`), or WAV | SSE: `heard` first, then the §4.3 events of the reply | | `GET /chat/device/speech/{message_id}?rate=16000` | | the whole reply as one streamed WAV | | `GET /chat/device/history?limit=12` | | `{"session", "generating", "messages": [{id, role, content}]}`, oldest first | | `POST /chat/device/new` | | start a fresh conversation (the current one if it is still empty) | The Prism sends Ogg Opus: 48 kHz stereo (its two microphones), 64 kbps, wideband, 20 ms packets, encoded while recording (~8 KB a second; 16 kHz mono WAV is 32 KB). Any format `/chat/transcribe` takes works. `talk` transcribes (as `/chat/transcribe`), asks in the current conversation with `"voice": true` and streams: ``` data: {"type": "heard", "text": "黄金现在多少钱", "language": "zh", "seconds": 2.0, "audio_id": 7} data: {"type": "start", "session": {…}, "user": {…}, "mode": "market", "level": null} data: {"type": "tool", "name": "get_quotes", "args": {"symbols": ["GOLD"]}} data: {"type": "delta", "content": "黄金"} (also "reset": replace the text so far) … data: {"type": "done", "message": {"id": 42, "content": "…", …}} ``` `"text": ""` = nothing heard, and the stream ends there (with `"why"` when the recognizer only made something up out of noise, see §3.1). A reply that can't start (one already in flight, too long) is `heard` + `error`. Errors before the stream (bad audio, speech busy, rate limit) are plain HTTP errors with `{"detail"}`. `speech` reads a reply of this device's aloud: markdown, links and emoji removed, split into sentences, synthesized a few ahead of the one being sent. The response is a 44-byte WAV header whose two sizes are `0xFFFFFFFF` (length unknown up front), then 16-bit mono PCM at `rate` until the stream ends: play it as it arrives. It counts as an expensive call (60 / min per IP, like `/chat/tts`). Edge takes ~3-5 s a sentence, so `talk` already starts synthesizing the reply (first sentence alone, then the next pieces as they are complete) while it streams: ask for the speech as soon as `done` arrives and the first audio is usually ready.