We will explore two techniques to compress our word embeddings:
- PCA dimensionality reduction
- Product quantization
You can find the demo on this page
Run fetch_vectors.sh to download pre-trained vectors from fastText trained on the common crawl dataset. The file contains 50k word vectors, each of 300 dimensions.
foo@bar:~$ ./fetch_vectors.sh
Downloading fastText word vectors for 50k most common words...
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 35.5M 100 35.5M 0 0 1272k 0 0:00:28 0:00:28 --:--:-- 1278k
Archive: data/crawl-300d-50K.vec.zip
inflating: data/crawl-300d-50K.vec
Run compressor.py to load the words vectors into gensim and calculate the accuracy using the word analogies test.
foo@bar:~$ python main.py -i data/crawl-300d-50K.vec
Loading vectors in gensim...
Calculating accuracy...
Accuracy: 87.078803%Add the -c parameter to compress the embeddings and re-calculate the accuracy. The compressed embeddings, codebook of centroids and vocabulary are saved in the model directory as JSON files.
foo@bar:~$ python compressor.py -c -i data/crawl-300d-50K.vec
Loading vectors in gensim...
Reduce dimensions using PCA...
Size reduction: 50.000000%
Compress embeddings using product quantization...
Size reduction: 93.000000%
Calculating accuracy...
Accuracy: 81.957310%
Saving generated/vocab.json
Saving generated/codes.json
Saving generated/centroids.jsonCompress the files using lz compression, bundle everything into a single JSON file
# Run this from the project root
yarn build-embeddings