添加链接
link之家
链接快照平台
  • 输入网页链接,自动生成快照
  • 标签化管理网页链接

本指南介绍如何将文本转语音头像与实时合成配合使用。 输入文本后,虚拟形象视频几乎会立即生成。

语音直播 API 中可以使用文本转语音头像来创建更个性化的语音对话。 有关详细信息,请参阅 Voice Live API 概述 和 Voice Live 示例代码中的虚拟形象 。

Prerequisites

Azure subscription: 免费创建一个订阅 。 Foundry 资源: 在一个受支持的区域中 创建 Microsoft Foundry 资源 。 有关区域可用性的详细信息,请参阅 文本转语音头像区域 。

或者,如果使用 Speech Studio:

Speech resource: 在 Azure 门户中创建语音资源 。 选择标准 S0 定价层来访问虚拟形象 。 语音资源密钥和区域:部署后,选择“转到资源” 以查看和管理密钥。

若要使用实时虚拟形象合成,请为首选平台和编程语言安装语音 SDK。 查看 安装语音 SDK 。

实时虚拟形象使用 WebRTC 将视频从服务器流式传输到客户端。 确保网络允许 WebRTC 流量。 如果你有防火墙,请添加规则以允许出站流量到 WebRTC 使用的 TURN(中继)服务器。 如果你使用默认的通信服务 TURN 服务器,请允许到 relay.communication.microsoft.com 的 UDP 端口 3478 和 TCP 端口 443 流量。 以下是你需要添加入站流量的防火墙规则:

来源 IP地址 目标 FQDN 目标 IP

选择文本转语音语言和语音

语音服务支持多种 语言和语音 。

要匹配输入文本并使用指定的声音,可以在 SpeechSynthesisLanguage 对象中设置 SpeechSynthesisVoiceName 或 SpeechConfig 属性:

const speechConfig = SpeechSDK.SpeechConfig.fromSubscription("YourSpeechKey", "YourSpeechRegion");
// Set either the `SpeechSynthesisVoiceName` or `SpeechSynthesisLanguage`.
speechConfig.speechSynthesisLanguage = "en-US";
speechConfig.speechSynthesisVoiceName = "en-US-Ava:DragonHDLatestNeural";   

所有神经语音都是多语言的,并且能够流利地使用自己的语言和英语。 例如,如果选择“es-ES-ElviraNeural”并输入英语文本,则虚拟形象会用西班牙口音说英语。

如果语音不支持输入语言,则语音服务不会创建音频。 如要查看完整列表,请参阅语言和语音支持。

默认语音选择:

  • 如果未设置 SpeechSynthesisVoiceName 或 SpeechSynthesisLanguage,则会使用 en-US 的默认语音。
  • 如果仅设置 SpeechSynthesisLanguage,则使用该区域设置中的默认语音。
  • 如果同时设置两者,则 SpeechSynthesisVoiceName 优先。
  • 如果使用 SSML 设置语音,则忽略这两个属性。
  • 选择虚拟形象角色和风格

    请参阅 受支持的标准虚拟形象。

    设置虚拟形象角色和风格:

    const avatarConfig = new SpeechSDK.AvatarConfig(
        "lisa", // Set avatar character here.
        "casual-sitting", // Set avatar style here.
    

    设置照片头像:

    const avatarConfig = new SpeechSDK.AvatarConfig(
        "anika", // Set photo avatar character here.
    avatarConfig.photoAvatarBaseModel = "vasa-1"; // Set photo avatar base model here.
    

    设置与实时虚拟形象的连接

    实时虚拟形象使用 WebRTC 协议流式传输视频。 使用 WebRTC 对等连接设置与虚拟形象服务的连接。

    首先,创建 WebRTC 对等连接对象。 WebRTC 是对等的,依赖于 ICE 服务器进行网络中继。 语音服务提供获取 ICE 服务器信息的 REST API。 建议从语音服务提取 ICE 服务器详细信息,但可以使用自己的详细信息。

    提取 ICE 信息的示例请求:

    GET /tts/cognitiveservices/avatar/relay/token/v1 HTTP/1.1
    Host: YourResourceName.cognitiveservices.azure.com
    Ocp-Apim-Subscription-Key: YOUR_RESOURCE_KEY
    

    使用 ICE 服务器 URL、用户名和上一响应中的凭据创建 WebRTC 对等连接:

    // Create WebRTC peer connection
    peerConnection = new RTCPeerConnection({
        iceServers: [{
            urls: [ "Your ICE server URL" ],
            username: "Your ICE server username",
            credential: "Your ICE server credential"
    

    ICE 服务器 URL 可以以 turn 开头(例如 turn:relay.communication.microsoft.com:3478),也可以以 stun 开头(例如 stun:relay.communication.microsoft.com:3478)。 urls 中仅包含 turn URL。

    然后,在对等连接的 ontrack 回调中设置视频和音频播放器元素。 此回调运行两次 - 一次用于视频,一次用于音频。 在回调中创建两个玩家元素:

    // Fetch WebRTC video/audio streams and mount them to HTML video/audio player elements
    peerConnection.ontrack = function (event) {
        if (event.track.kind === 'video') {
            const videoElement = document.createElement(event.track.kind)
            videoElement.id = 'videoPlayer'
            videoElement.srcObject = event.streams[0]
            videoElement.autoplay = true
        if (event.track.kind === 'audio') {
            const audioElement = document.createElement(event.track.kind)
            audioElement.id = 'audioPlayer'
            audioElement.srcObject = event.streams[0]
            audioElement.autoplay = true
    // Offer to receive one video track, and one audio track
    peerConnection.addTransceiver('video', { direction: 'sendrecv' })
    peerConnection.addTransceiver('audio', { direction: 'sendrecv' })
    

    然后,使用语音 SDK 创建虚拟形象合成器,并使用对等连接连接到虚拟形象服务:

    // Create avatar synthesizer
    var avatarSynthesizer = new SpeechSDK.AvatarSynthesizer(speechConfig, avatarConfig)
    // Start avatar and establish WebRTC connection
    avatarSynthesizer.startAvatarAsync(peerConnection).then(
        (r) => { console.log("Avatar started.") }
    ).catch(
        (error) => { console.log("Avatar failed to start. Error: " + error) }
    

    实时 API 在 5 分钟空闲或 30 分钟连接后断开连接。 为了让虚拟形象运行更长时间,请启用自动重新连接。 请参阅此 JavaScript 示例代码(搜索“自动重新连接”)。

    从文本输入合成会说话的虚拟形象视频

    设置完成后,虚拟形象视频会在浏览器中播放。 虚拟形象微微闪烁并移动,但在发送文本输入之前不会说话。

    将文本发送到虚拟形象合成器,使虚拟形象能够说话:

    var spokenText = "I'm excited to try text to speech avatar."
    avatarSynthesizer.speakTextAsync(spokenText).then(
        (result) => {
            if (result.reason === SpeechSDK.ResultReason.SynthesizingAudioCompleted) {
                console.log("Speech and avatar synthesized to video stream.")
            } else {
                console.log("Unable to speak. Result ID: " + result.resultId)
                if (result.reason === SpeechSDK.ResultReason.Canceled) {
                    let cancellationDetails = SpeechSDK.CancellationDetails.fromResult(result)
                    console.log(cancellationDetails.reason)
                    if (cancellationDetails.reason === SpeechSDK.CancellationReason.Error) {
                        console.log(cancellationDetails.errorDetails)
    }).catch((error) => {
        console.log(error)
        avatarSynthesizer.close()
    

    关闭实时虚拟形象连接

    为了避免额外的费用,请在完成后关闭连接:

  • 关闭浏览器会释放 WebRTC 对等连接,并在几秒钟后关闭虚拟形象连接。

  • 如果虚拟形象空闲 5 分钟,连接将自动关闭。

  • 可以手动关闭虚拟形象连接:

    avatarSynthesizer.close()
    

    设置背景色

    使用 backgroundColor 的 AvatarConfig 属性设置虚拟形象视频的背景颜色:

    const avatarConfig = new SpeechSDK.AvatarConfig(
        "lisa", // Set avatar character here.
        "casual-sitting", // Set avatar style here.
    avatarConfig.backgroundColor = '#00FF00FF' // Set background color to green
    

    颜色字符串的格式应为 #RRGGBBAA。 忽略 alpha 通道 (AA) - 实时虚拟形象不支持透明背景。

    设置背景图像

    使用 backgroundImage 的 AvatarConfig 属性设置背景图像。 将图像上传到公共 URL 并将其分配给 backgroundImage:

    const avatarConfig = new SpeechSDK.AvatarConfig(
        "lisa", // Set avatar character here.
        "casual-sitting", // Set avatar style here.
    avatarConfig.backgroundImage = "https://www.example.com/1920-1080-image.jpg" // A public accessiable URL of the image.
    

    设置背景视频

    API 不直接支持背景视频,但你可以在客户端自定义背景:

  • 将背景色设置为绿色(以方便着色)。
  • 创建大小与虚拟形象视频相同的画布元素。
  • 对于每一帧,将绿色像素设置为透明并将帧绘制到画布上。
  • 隐藏原始视频。
  • 这样,画布上就会出现一个透明的虚拟形象。 请参阅 JavaScript 示例代码。

    然后,可以将任意动态内容(如视频)放置在画布后面。

    虚拟形象视频默认为 16:9。 若要裁剪为不同的纵横比,请使用左上角和右下角的坐标指定矩形区域:

    const videoFormat = new SpeechSDK.AvatarVideoFormat()
    const topLeftCoordinate = new SpeechSDK.Coordinate(640, 0) // coordinate of top-left vertex, with X=640, Y=0
    const bottomRightCoordinate = new SpeechSDK.Coordinate(1320, 1080) // coordinate of bottom-right vertex, with X=1320, Y=1080
    videoFormat.setCropRange(topLeftCoordinate, bottomRightCoordinate)
    const avatarConfig = new SpeechSDK.AvatarConfig(
        "lisa", // Set avatar character here.
        "casual-sitting", // Set avatar style here.
        videoFormat, // Set video format here.
    

    有关完整示例,请参阅我们的 code 示例并搜索 crop。

    设置视频分辨率

    4K 分辨率头像的默认输出分辨率为 3840x2160。 但是,你可以调整输出分辨率,同时保持原始纵横比以满足你的要求。 例如,在流式处理期间将输出分辨率设置为 1920x1080 可以减少网络带宽消耗。

    const videoFormat = new SpeechSDK.AvatarVideoFormat();
    videoFormat.width = 1920;
    videoFormat.height = 1080;
    

    有关更多详细信息,请参阅 sample 代码。

    设置照片虚拟形象的虚拟形象场景

    对于照片头像,可以通过调整缩放、位置和旋转参数来配置场景。 这样就可以自定义虚拟形象在视频输出中的显示方式。

    使用以下参数创建对象 AvatarSceneConfig :

    缩放:介于 0 和 1 之间的值,其中 1.0 表示 100% 缩放(默认值)。
  • positionX:水平位置偏移量,范围从 -1 到 1,其中 0 居中。 positionY:垂直位置偏移量,范围从 -1 到 1,其中 0 居中。 rotationX:绕 X 轴以弧度旋转 rotationY:以弧度绕 Y 轴旋转。 rotationZ:以弧度绕 Z 轴旋转。 振幅:虚拟形象运动的振幅,范围为 0 到 1,其中小于 1 的值减少运动振幅,1.0(默认值)表示全振幅。

    创建虚拟形象配置时设置初始场景配置:

    const avatarConfig = new SpeechSDK.AvatarConfig(
        "anika", // Set photo avatar character here.
    avatarConfig.photoAvatarBaseModel = "vasa-1";
    avatarConfig.scene = new SpeechSDK.AvatarSceneConfig(
        1.0,  // zoom: 100%
        0.0,  // positionX: centered
        0.0,  // positionY: centered
        0.0,  // rotationX: no rotation
        0.0,  // rotationY: no rotation
        0.0,  // rotationZ: no rotation
        1.0   // amplitude: full movement
    

    您还可以在化身会话开始后使用 updateSceneAsync 方法动态更新场景。

    // Convert percentage and degree values to normalized values
    const zoom = 0.85; // 85% zoom
    const positionX = 0.1; // 10% offset to the right
    const positionY = -0.05; // 5% offset upward
    const rotationX = 10 * Math.PI / 180; // 10 degrees in radians
    const rotationY = 0;
    const rotationZ = 0;
    const amplitude = 0.8; // 80% movement amplitude
    const sceneConfig = new SpeechSDK.AvatarSceneConfig(
        zoom,
        positionX,
        positionY,
        rotationX,
        rotationY,
        rotationZ,
        amplitude
    

    场景配置仅适用于照片头像。 视频头像不支持此功能。

    在 Speech SDK 的 GitHub 存储库中查找文本转语音头像代码示例。 这些示例演示如何在 Web 和移动应用中使用实时虚拟形象:

    服务器端 + 客户端 Python (服务器) + JavaScript (客户端) C# (服务器) + JavaScript (客户端) Node.js (服务器) + JavaScript (客户端) JavaScript Android Python Node.js 什么是文本转语音虚拟形象 安装语音 SDK