- Add a
keywordsexecutable for TF-IDF keyword extraction.keywords fitbuilds a vocabulary from files or standard input,keywords extractreads a file,keywords infoprints model statistics, and a bare text argument transforms that text. The command printsterm:scorepairs and maps stems back to whole words, sokeywords "Ruby is elegant"printselegant:0.61 ruby:0.61. Options set the model path, the top-N terms, quiet mode,--min-df,--max-df, and--ngram. Usage errors exit 2 and other errors exit 1. The gem now installs two executables,classifierandkeywords. - Treat each line as a separate document during
keywords fit. The document count controls the inverse document frequency, so a file with many lines contributes many documents. - Add
Classifier::Streaming::MultiIO. It reads several IO objects or file paths as one sequential stream. It opens and closes each path one at a time, so a corpus larger than the file descriptor limit still fits. - Add
String#stem_to_word_hash. It maps each stemmed root to the most frequent original word in the text. - Add the
min_dfandmax_dfreaders toClassifier::TFIDF. - Accept a
MultiIOinTFIDF#fit_from_streamandStreaming::LineReader. - Fix
Marshalsupport inClassifier::Bayes. The dump left outmin_word_length, so the restored classifier raisedArgumentError: comparison of Integer with nil failedon its firstclassify. A dump written by an older version still loads, and takes the configured default. - Fix
LSI#highest_relative_content, which returned an Enumerator rather than the documented array of documents. - Fix
LSI#highest_ranked_stems, which repeated one stem when the document vector held equal weights. It looked up each weight by value, so tied weights all resolved to the same index. - Add a
docs/reference for the command line tools, each classifier, persistence, streaming, and configuration.
- Add a
--searchflag to the command line tool to filter local models. - Add a model detail view to the command line tool.
- Read model information in advance when the tool lists local models.
- Accept keyword arguments in
train_from_stream.
- Make
min_word_lengthconfigurable, so a caller can keep or drop short words during tokenization. - Add a Claude Code plugin with a skill and slash commands.
- Force UTF-8 encoding on the HTTP response body, so a remote model loads under any locale.
- Force UTF-8 encoding when the locale is not UTF-8. Model data and user input no longer raise an encoding error.
- Add a
classifierexecutable with a model registry. The tool trains, classifies, and manages saved models from the shell. - Fix the examples and the broken links in the README.
- Add Dependabot for automated dependency updates.
- Breaking: set
required_ruby_versionto>= 3.1. Older Ruby versions are no longer supported. - Add a k-Nearest Neighbors classifier.
- Add a Logistic Regression classifier.
- Add a TF-IDF vectorizer.
- Add streaming training and incremental SVD, so a corpus larger than memory can train a model.
- Add a hash-style API to add items to LSI.
- Add keyword arguments to
Bayes#trainandBayes#untrain. - Accept an array of categories in the classifier constructor.
- Fix the sentence and paragraph splits in
Summary. - Add property-based tests for the probabilistic invariants.
- Replace the optional GSL dependency with a bundled C extension for LSI. The extension has no external dependency and falls back to pure Ruby.
- Add pluggable persistence backends through a storage API.
- Add
saveandloadmethods for classifier persistence. - Add thread safety to the Bayes and LSI classifiers.
- Expose the LSI tuning parameters, with validation and an introspection API.
- Cache the expensive computations in the Bayes classifier.
- Breaking: replace the fixed 0.1 constant in the Bayes classifier with
add-one (Laplace) smoothing, where
P(word|category) = (count + 1) / (total + vocabulary_size). The smoothing now scales with the vocabulary size and applies to seen and unseen words alike. Classification scores change as a result. - Fix an LSI dimension mismatch in the pure Ruby SVD.
- Fix the numerical stability of the SVD implementation.
- Replace the separate RBS files with inline annotations.
- Add a GitHub Actions workflow that publishes the gem on a version tag.
- Add an LSI benchmark that compares GSL against pure Ruby.
- Add RuboCop and SimpleCov.
- Improve the scaling of the LSI content node.
- Require
setand use the explicit::Setnamespace. - Refactor
prepare_category_name.
- Fix the word count when
remove_categoryruns. - Add the
mutex_mdependency and update thefast-stemmerversion.
- Add
remove_categoryto the Bayes classifier.
- Add
classify_with_confidenceto the LSI classifier. - Require
mathnonly for Ruby 2.5 and later, and addcmathfor Ruby 2.7 and later. - Silence the warnings about an uninitialized
$GSL, and correct the rb-gsl URL hint. - Package the test files in the gem.
- Add a Gemfile.lock.
- Use Minitest for the test suite.
- Add the
mathndependency, which Ruby 2.5.0 removed. - Make the gem installable through Bundler and a git remote.
- Fix the gemspec and the unit tests.
- Use a prior in the Bayes classifier.
- Change the skip word list from an array to a set.
- Reduce the number of regular expression matches during tokenization.
- Stem only the words that the tokenizer keeps.
- Default the word hash values to 0.
- Add a gemspec, a README, and Travis CI.
- Use fast-stemmer for the Porter stemmer.
- Check
$GSLbefore a call toMatrix.diag, so the code usesGSL::Matrix.
- Fix the reported issue #1.
Versions 1.0 through 1.3.1 reached RubyGems on 2009-07-25. The git history before 2010 is too sparse to attribute each change to one of these versions.