Simple, hackable offline speech to text - using the VOSK-API.

Last update: Jan 07, 2023

Related tags

Overview

Nerd Dictation

Offline Speech to Text for Desktop Linux.

This is a utility that provides simple access speech to text for using in Linux without being tied to a desktop environment.

Simple: This is a single file Python script with minimal dependencies.
Hackable: User configuration lets you manipulate text using Python string operations.
Zero Overhead: As this relies on manual activation there are no background processes.

Dictation is accessed manually with begin/end commands.

This uses the excellent vosk-api.

Usage

It is suggested to bind begin/end/cancel to shortcut keys.

nerd-dictation begin

nerd-dictation end

For details on how this can be used, see: nerd-dictation --help and nerd-dictation begin --help.

Features

Specific features include:

Numbers as Digits

Optional conversion from numbers to digits.

So Three million five hundred and sixty second becomes 3,000,562nd.

A series of numbers (such as reciting a phone number) is also supported.

So Two four six eight becomes 2,468.

Time Out

Optionally end speech to text early when no speech is detected for a given number of seconds. (without an explicit call to end which is otherwise required).

Output Type

Output can simulate keystroke events (default) or simply print to the standard output.

User Configuration Script

User configuration is just a Python script which can be used to manipulate text using Python's full feature set.

See nerd-dictation begin --help for details on how to access these options.

Dependencies

Python 3.
The VOSK-API.
parec command (for recording from pulse-audio).
xdotool command to simulate keyboard input.

Install

pip3 install vosk
git clone https://github.com/ideasman42/nerd-dictation.git
cd nerd-dictation
wget https://alphacephei.com/kaldi/models/vosk-model-small-en-us-0.15.zip
unzip vosk-model-small-en-us-0.15.zip
mv vosk-model-small-en-us-0.15 model

To test dictation:

./nerd-dictation begin --vosk-model-dir=./model &
# Start speaking.
./nerd-dictation end

Reminder that it's up to you to bind begin/end/cancel to actions you can easily access (typically key shortcuts).
To avoid having to pass the --vosk-model-dir argument, copy the model to the default path:
```
mkdir -p ~/.config/nerd-dictation
mv ./model ~/.config/nerd-dictation
```

Hint

Once this is working properly you may wish to download one of the larger language models for more accurate dictation. They are available here.

Configuration

This is an example of a trivial configuration file which simply makes the input text uppercase.

# ~/.config/nerd-dictation/nerd-dictation.py
def nerd_dictation_process(text):
    return text.upper()

A more comprehensive configuration is included in the examples/ directory.

Hints

The processing function can be used to implement your own actions using keywords of your choice. Simply return a blank string if you have implemented your own text handling.
Context sensitive actions can be implemented using command line utilities to access the active window.

Paths

Local Configuration

~/.config/nerd-dictation/nerd-dictation.py

Language Model

~/.config/nerd-dictation/model

Note that --vosk-model-dir=PATH can be used to override the default.

Details

Typing in results will never press enter/return.
Pulse audio is used for recording.
Recording and speech to text a performed in parallel.

Examples

Store the result of speech to text as a variable in the shell:

SPEECH="$(nerd-dictation begin --timeout=1.0 --output=STDOUT)"

Limitations

Text from VOSK is all lower-case, while the user configuration can be used to set the case of common words like I this isn't very convenient (see the example configuration for details).
For some users the delay in start up may be noticeable on systems with slower hard disks especially when running for the 1st time (a cold start).

This is a limitation with the choice not to use a service that runs in the background. Recording begins before any the speech-to-text components are loaded to mitigate this problem.

Further Work

And a general solution to capitalize words (proper nouns for example).
Preview output while dictating.
Wayland support (this should be quite simple to support and mainly relies on a replacement for xdotool).
Add a setup.py for easy installation on uses systems.
Possibly other speech to text engines (only if they provide some significant benefits).
Possibly support Windows & macOS.

Simple, hackable offline speech to text - using the VOSK-API.

Related tags

Overview

Nerd Dictation

Usage

Features

Dependencies

Install

Configuration

Hints

Paths

Details

Examples

Limitations

Further Work

Owner

Campbell Barton

Analysis of voices based on the Mel-frequency band

Speech Algorithms Collections

Use python MIDI to write some simple music

Sequencer: Deep LSTM for Image Classification

Code to work with wave files!

Tradutor de um arquivo MIDI para ser usado em um simulador RISC-V(RARS)

𝙰 𝙼𝚞𝚜𝚒𝚌 𝙱𝚘𝚝 𝙲𝚛𝚎𝚊𝚝𝚎𝚍 𝙱𝚢 𝚃𝚎𝚊𝚖𝙳𝚕𝚝 💖

Make an audio file (really) long-winded

The project aims to develop a personal-assistant for Windows & Linux-based systems

A python wrapper for REAPER

Learn chords with your MIDI keyboard !

Tune in is a Collaborative Music Playing Systems where multiple guests can join a room and enjoy the song being played

📺Headless全自动B站直播录播、切片、上传一体工具

Bot duniya Music Player

This bot can stream audio or video files and urls in telegram voice chats

commonfate 📦commonfate 📦 - Common Fate Model and Transform.

?️ Open Source Audio Matching and Mastering

Audio Retrieval with Natural Language Queries: A Benchmark Study

Inner ear models for Python

Generating a structured library of .wav samples with Python.