GitHub - gildasch/stopwords: Removes most frequent words (stop words) from a text content. Based on a Curated list of language statistics.

stopwords is a go package that removes stop words from a text content. If instructed to do so, it will remove HTML tags and parse HTML entities. The objective is to prepare a text in view to be used by natural processing algos or text comparison algorithms such as SimHash.

It uses a curated list of the most frequent words used in these languages:

arabic
bulgarian
czech
danish
english
finnish
french
german
hungarian
italian
japanese
latvian
norwegian
persian
polish
portuguese
romanian
russian
slovak
spanish
swedish
thai
turkish

If the function is used with an unsupported language, it doesn't fail, but will apply english filter to the content.

How to use this package

You can find an example here https:github.com/bbalet/gorelated where stopwords package is used in conjunction with SimHash algorithm in order to find a list of related content for a static website generator:

import (
      "github.com/bbalet/stopwords"
)

//Example with 2 strings containing P html tags
//"la", "un", etc. are (stop) words without lexical value in French
string1 := []byte("<p>la fin d'un bel après-midi d'été</p>")
string2 := []byte("<p>cet été, nous avons eu un bel après-midi</p>")

//Return a string where HTML tags and French stop words has been removed
cleanContent := stopwords.CleanContent(string1, "fr", true)

//Get two (Sim) hash representing the content of each string
hash1 := stopwords.Simhash(string1, "fr", true)
hash2 := stopwords.Simhash(string2, "fr", true)

//Hamming distance between the two strings (diffference between contents)
distance := stopwords.CompareSimhash(hash1, hash2)

//Clean the content of string1 and string2, compute the Levenshtein Distance
stopwords.LevenshteinDistance(string1, string2, "fr", true)

Where fr is the ISO 639-1 code for French (it accepts a BCP 47 tag as well). https:en.wikipedia.org/wiki/List_of_ISO_639-1_codes

Credits

Most of the lists were built by IR Multilingual Resources at UniNE http:members.unine.ch/jacques.savoy/clef/index.html

License

stopwords is released under the BSD license.

Name		Name	Last commit message	Last commit date
Latest commit History 12 Commits
.gitignore		.gitignore
.travis.yml		.travis.yml
LICENSE		LICENSE
README.md		README.md
benchmark_test.go		benchmark_test.go
levenshtein.go		levenshtein.go
levenshtein_test.go		levenshtein_test.go
simhash.go		simhash.go
simhash_test.go		simhash_test.go
stopwords.go		stopwords.go
stopwords_ar.go		stopwords_ar.go
stopwords_bg.go		stopwords_bg.go
stopwords_cs.go		stopwords_cs.go
stopwords_da.go		stopwords_da.go
stopwords_de.go		stopwords_de.go
stopwords_el.go		stopwords_el.go
stopwords_en.go		stopwords_en.go
stopwords_es.go		stopwords_es.go
stopwords_fa.go		stopwords_fa.go
stopwords_fi.go		stopwords_fi.go
stopwords_fr.go		stopwords_fr.go
stopwords_hu.go		stopwords_hu.go
stopwords_it.go		stopwords_it.go
stopwords_ja.go		stopwords_ja.go
stopwords_lv.go		stopwords_lv.go
stopwords_nl.go		stopwords_nl.go
stopwords_no.go		stopwords_no.go
stopwords_pl.go		stopwords_pl.go
stopwords_pt.go		stopwords_pt.go
stopwords_ro.go		stopwords_ro.go
stopwords_ru.go		stopwords_ru.go
stopwords_sk.go		stopwords_sk.go
stopwords_sv.go		stopwords_sv.go
stopwords_test.go		stopwords_test.go
stopwords_th.go		stopwords_th.go
stopwords_tr.go		stopwords_tr.go

License

gildasch/stopwords

Folders and files

Latest commit

History

Repository files navigation

How to use this package

Credits

License

About

Resources

License

Stars

Watchers

Forks

Languages