音声と会話ランタイム
接続時の Orbz は無音です。@neongate-ai/orbz@1.0.3 に挨拶、人格、会話は組み込まれていません。talk は凍結された空のオブジェクト、DEFAULT_TALK_FLOW は凍結された空配列です。ホストが一文用の speech または会話用の型付き talkFlow と音声エンジンを設定し、ユーザーの明示的な操作後にだけ開始します。
speech は talkFlow より優先されます。どちらもなければ startTalking() は何もしません。空でないフローの開始で実行時コンテキストがリセットされ、ask ステップはホストの receive() 呼び出し時に指定の capture キーへ入力を保存します。後続ステップはその値を埋め込めます。Cookie、ローカルストレージ、IndexedDB、バックエンドへの永続化はしません。
組み込みの会話データ
import { DEFAULT_TALK_FLOW, talk } from "@neongate-ai/orbz";
console.log(talk); // {}
console.log(DEFAULT_TALK_FLOW); // []フローを続ける
下記の独自フローが開始して ask に到達した後、ユーザーが送信した入力を receive() に渡します。fullName は例の capture: "fullName" 宣言で作るキーであり、組み込みデータではありません。
import "@neongate-ai/orbz/browser";
import type { OrbzElement } from "@neongate-ai/orbz";
const orb = document.querySelector<OrbzElement>("orb-z");
await orb?.receive("Jonatas");
console.log(orb?.talkContext.fullName);ブラウザーの英語音声
導入済みパッケージの WebSpeechAdapter の既定値は pt-BR です。例文が英語なので、この例では en-US を明示します。他の言語ではホストのテキストとアダプターの言語を両方設定します。非同期の音声一覧を待ち、指定言語に一致する音声を優先しますが、利用可能な音声は訪問者の環境によります。
import '@neongate-ai/orbz/browser'
import {
WebSpeechAdapter,
type OrbzElement
} from "@neongate-ai/orbz";
const orb = document.createElement("orb-z") as OrbzElement;
orb.speech = "Hello. This is an explicit speech example.";
orb.voiceEngine = new WebSpeechAdapter({
language: "en-US",
preferredVoices: ["Google US English", "Microsoft Aria Online"]
});
document.body.append(orb);
const startVoiceButton = document.querySelector<HTMLButtonElement>("[data-start-voice]");
startVoiceButton?.addEventListener("click", async () => {
await orb.startTalking();
});実際にインストールされている音声は、訪問者のブラウザーと OS に依存します。WebSpeechAdapter は選択を改善しますが、システム音声を OpenAI 音声に変えることはできません。
OpenAI 品質の音声
1.0.3 の OpenAISpeechAdapter は gpt-4o-mini-tts、marin、MP3、ブラジルポルトガル語の読み上げ指示が既定です。この英語例はホストの文に合わせて instructions を上書きします。例のエンドポイントはアプリが実装し、保護する必要があります。
import '@neongate-ai/orbz/browser'
import {
OpenAISpeechAdapter,
type OrbzElement
} from "@neongate-ai/orbz";
const orb = document.createElement("orb-z") as OrbzElement;
orb.speech = "Hello. This is an explicit speech example.";
orb.voiceEngine = new OpenAISpeechAdapter({
endpoint: "/api/orbz/speech",
instructions: "Speak in natural American English. Do not change the supplied text."
});
document.body.append(orb);
const startVoiceButton = document.querySelector<HTMLButtonElement>("[data-start-voice]");
startVoiceButton?.addEventListener("click", async () => {
await orb.startTalking();
});エンドポイントは、連携を実装するアプリケーションが管理します。input、instructions、model、response_format、voice を含む OpenAI 互換の JSON 本文を受け取り、生成された音声を返します。OpenAI API キーはそのサーバー側エンドポイントで保持し、ブラウザーコードや npm パッケージには決して含めないでください。
生成音声を使うアプリケーションは、その声が AI によって生成されていることを明確に伝える必要があります。
明示的な開始とブラウザーポリシー
音声を開始などの明確なラベルを付けたネイティブの <button> を表示し、クリックハンドラーから startTalking() を呼び出します。オーブの接続、音声エンジンの割り当て、ページへの移動で音声が始まることはありません。
ブラウザーが要求された音声を NotAllowedError で拒否した場合、Orbz は元のエラーを含む orbz-talk-error を発火し、次のポインター、キーボード、タッチ操作の後に要求されたフローを再試行します。リセット操作では、音声の制限を解除するためだけに要素を再マウントしないでください。
独自のフローを指定する
import "@neongate-ai/orbz/browser";
import { WebSpeechAdapter, type OrbzElement, type OrbzTalkStep } from "@neongate-ai/orbz";
const flow = [
{ id: "welcome", kind: "say", needsAuth: false, text: "Hello." },
{ id: "name", kind: "ask", needsAuth: false, text: "What is your name?", capture: "fullName" },
{ id: "help", kind: "say", needsAuth: false, text: "How can I help, {{fullName}}?" },
{ id: "answer", kind: "respond", needsAuth: false, strategy: "openai", fallback: "I cannot answer that right now." }
] as const satisfies readonly OrbzTalkStep[];
const orb = document.createElement("orb-z") as OrbzElement;
orb.talkFlow = flow;
orb.voiceEngine = new WebSpeechAdapter({ language: "en-US" });
document.body.append(orb);
const startVoiceButton = document.querySelector<HTMLButtonElement>("[data-start-voice]");
startVoiceButton?.addEventListener("click", async () => {
await orb.startTalking();
});明示的に開始する実行で利用できるよう、startTalking() を呼び出す前に voiceEngine、talkFlow、intelligence を設定します。
任意のインテリジェンス連携
import type {
OrbzElement,
OrbzIntelligencePort
} from "@neongate-ai/orbz";
const intelligence: OrbzIntelligencePort = {
async respond(input, context) {
return productAgent.respond({ context, input });
}
};
orb.intelligence = intelligence;イベントと視覚状態
音声の再生中、Orbz は一時的に speaking の視覚状態を使い、その後で元の状態に戻します。
| イベント | 詳細 |
|---|---|
orbz-speaking-change | { speaking: boolean } |
orbz-talk-error | { error: unknown } |
上記のテキスト読み上げフローはマイクを収録しません。入力 UI、権限、文字起こし、製品ロジック、receive() はホストが管理します。Orbz には別の Realtime 会話 API もありますが、この startTalking() の例では起動しません。