This repository implements a brute-force spellchecker utilizing the Damerau-Levenshtein edit distance.

Last update: Dec 11, 2021

Overview

About spellchecker.py

Implementing a highly-accurate, brute-force, and dynamically programmed spellchecking program that utilizes the Damerau-Levenshtein string metric for measuring edit distance between two sequences of characters.

How to Write Your Own Test Cases

In the lib folder, you will see two different text files called 'candidate_words.txt' and 'incorrect_words.txt':

The candidate_words.txt text file can contain an unlimited amount of CORRECTLY spelled words, with each word written on a new line.
The incorrect_words.txt text file can contain an unlimited amount of INCORRECTLY spelled words, with each word written on a new line. However, each incorrectly spelled word in this list MUST have its correctly spelled counterpart contained somewhere in the 'candidate_words.txt' text file. It doesn't matter where, since the 'candidate_words.txt' file will be randomly shuffled anyway.

In the test folder, you will see a text file called target_words.txt:

The 'target_words.txt' file will contain the CORRECT spelling of each word contained in the 'incorrect_words.txt' text file, with each being on a new line in the same exact order that you inserted their incorrectly spelled counterparts in the 'incorrect_words.txt' text file. It is important that both the incorrectly and correctly spelled words are in the same order to be able to calculate the accuracy of the spell checker.

To view an example on how to create your own test cases, take a look at the files provided in either folder.

How to Run the Program

Enter the folder's directory using your terminal. Then, simply run python3 spellchecker.py

The only thing you will need to modify are the files in the lib and test folders if you want to try the program with your own test cases. The program does not need to be touched, unless you'd like to modify the global variable 'THRESHOLD', which is used as the threshold to find an incorrectly spelled word's closest approximation.
The incorrectly spelled words in 'incorrect_words.txt' will be run through the program to find its closest lexical match from the candidate_words.txt text file using the Damerau-Levenshtein algorithm.
The spellchecked words will then be, in order, cross checked against its intended counterparts in target_words.txt to calculate the overall accuracy of the spellchecking algorithm.

The results of the program will then be printed to your terminal.

Dependencies

Ensure that you have difflib installed for python3.

Final Words

Feel free to use or modify this program for your intended purposes!

This repository implements a brute-force spellchecker utilizing the Damerau-Levenshtein edit distance.

Related tags

Overview

About spellchecker.py

How to Write Your Own Test Cases

How to Run the Program

Dependencies

Final Words

Owner

Raihan Ahmed

glow-speak is a fast, local, neural text to speech system that uses eSpeak-ng as a text/phoneme front-end.

Official code of our work, Unified Pre-training for Program Understanding and Generation [NAACL 2021].

a CTF web challenge about making screenshots

Indobenchmark are collections of Natural Language Understanding (IndoNLU) and Natural Language Generation (IndoNLG)

Question answering app is used to answer for a user given question from user given text.

Bpe algorithm can finetune tokenizer - Bpe algorithm can finetune tokenizer

Code for text augmentation method leveraging large-scale language models

End-to-end text to speech system using gruut and onnx. There are 40 voices available across 8 languages.

Code from the paper "High-Performance Brain-to-Text Communication via Handwriting"

Transformer-based Text Auto-encoder (T-TA) using TensorFlow 2.

I can help you convert your images to pdf file.

A Pytorch implementation of "Splitter: Learning Node Representations that Capture Multiple Social Contexts" (WWW 2019).

The code for the Subformer, from the EMNLP 2021 Findings paper: "Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers", by Machel Reid, Edison Marrese-Taylor, and Yutaka Matsuo

Code for the paper TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks

[KBS] Aspect-based sentiment analysis via affective knowledge enhanced graph convolutional networks

This is a modification of the OpenAI-CLIP repository of moein-shariatnia

🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.

This repository is home to the Optimus data transformation plugins for various data processing needs.

PyTorch original implementation of Cross-lingual Language Model Pretraining.

This repo contains simple to use, pretrained/training-less models for speaker diarization.