30行Python代碼打造一款簡單的人工語音對話

Posted on 2021-05-21 by WalkonNet

@Author：Runsen

1876年，亞歷山大·格雷厄姆·貝爾（Alexander Graham Bell）發明瞭一種電報機，可以通過電線傳輸音頻。托馬斯·愛迪生（Thomas Edison）於1877年發明瞭留聲機，這是第一臺記錄聲音並播放聲音的機器。

最早的語音識別軟件之一是由Bells Labs在1952年編寫的，隻能識別數字。1985年，IBM發佈瞭使用“隱馬爾可夫模型”的軟件，該軟件可識別1000多個單詞。

幾年前，一個replace("?","")代碼價值一個億

如今，在Python中Tensorflow，Keras，Librosa，Kaldi和語音轉文本API等多種工具使語音計算變得更加容易。

今天，我使用gtts和speech_recognition，教大傢如何通過三十行代碼，打造一款簡單的人工語音對話。思路就是將語音變成文本，然後文本變成語音。

gtts

gtts是將文字轉化為語音，但是需要在VPN下使用。這個因為要接谷歌服務器。

具體gtts的官方文檔：

下面，讓我們看一段簡單的的代碼

from gtts import gTTS

def speak(audioString):
    print(audioString)
    tts = gTTS(text=audioString, lang='en')
    tts.save("audio.mp3")
    os.system("audio.mp3")
    
speak("Hi Runsen, what can I do for you?")

執行上面的代碼，就可以生成一個mp3文件，播放就可以聽到瞭Hi Runsen, what can I do for you?。這個MP3會自動彈出來的。

speech_recognition

speech_recognition用於執行語音識別的庫，支持在線和離線的多個引擎和API。

speech_recognition具體官方文檔

安裝speech_recognition可以會出現錯誤，對此解決的方法是通過該網址安裝對應的whl包

在官方文檔中提供瞭具體的識別來自麥克風的語音輸入的代碼

下面就是 speech_recognition 用麥克風記錄下你的話，這裡我使用的是
recognize_google，speech_recognition 提供瞭很多的類似的接口。

import time
import speech_recognition as sr

# 錄下來你講的話
def recordAudio():
    # 用麥克風記錄下你的話
    print("開始麥克風記錄下你的話")
    r = sr.Recognizer()
    with sr.Microphone() as source:
        audio = r.listen(source)
    data = ""
    try:
        data = r.recognize_google(audio)
        print("You said: " + data)
    except sr.UnknownValueError:
        print("Google Speech Recognition could not understand audio")
    except sr.RequestError as e:
        print("Could not request results from Google Speech Recognition service; {0}".format(e))
    return data

if __name__ == '__main__':
    time.sleep(2)
    while True:
        data = recordAudio()
        print(data)

下面是我亂說的英語

對話

上面，我們實現瞭用麥克風記錄下你的話，並且得到瞭對應的文本，那麼下一步就是字符串的文本操作瞭，比如說how are you，那回答"I am fine”，然後將"I am fine”通過gtts是將文字轉化為語音

# @Author：Runsen
# -*- coding: UTF-8 -*-
import speech_recognition as sr
from time import ctime
import time
import os
from gtts import gTTS


# 講出來AI的話
def speak(audioString):
    print(audioString)
    tts = gTTS(text=audioString, lang='en')
    tts.save("audio.mp3")
    os.system("audio.mp3")


# 錄下來你講的話
def recordAudio():
    # 用麥克風記錄下你的話
    r = sr.Recognizer()
    with sr.Microphone() as source:
        audio = r.listen(source)

    data = ""
    try:
        data = r.recognize_google(audio)
        print("You said: " + data)
    except sr.UnknownValueError:
        print("Google Speech Recognition could not understand audio")
    except sr.RequestError as e:
        print("Could not request results from Google Speech Recognition service; {0}".format(e))

    return data


# 自帶的對話技能（邏輯代碼：rules）
def jarvis():
    while True:
        data = recordAudio()
        print(data)
        if "how are you" in data:
            speak("I am fine")
        if "time" in data:
            speak(ctime())
        if "where is" in data:
            data = data.split(" ")
            location = data[2]
            speak("Hold on Runsen, I will show you where " + location + " is.")
            # 打開谷歌地址
            os.system("open -a Safari https://www.google.com/maps/place/" + location + "/&amp;")

        if "bye" in data:
            speak("bye bye")
            break


if __name__ == '__main__':
    # 初始化
    time.sleep(2)
    speak("Hi Runsen, what can I do for you?")

    # 跑起
    jarvis()

當我說how are you？會彈出I am fine的mp3

當我說where is Chiana？會彈出Hold on Runsen, I will show you where China is.的MP3

同樣也會彈出China的谷歌地圖

本項目對應的Github

以上就是30行Python代碼打造一款簡單的人工語音對話的詳細內容，更多關於Python人工語音對話的資料請關註WalkonNet其它相關文章！

30行Python代碼打造一款簡單的人工語音對話

gtts

speech_recognition

對話

推薦閱讀：

發佈留言取消回覆

近期文章

gtts

speech_recognition

對話

推薦閱讀：

發佈留言 取消回覆

近期文章

標籤

發佈留言取消回覆