What is and why make a “Bag of Words” model?
Like everyone, I’ve become interested in the epistemic nature of Large Language Models (LLM).
So, how to proceed? When you focus directly at LLMs they are complicated and difficult to grok.
I’ve opted for an evolutionary approach. I think of this as stepping stones to the LLM revolution. What significant artefacts came before? And then how, who, what, why and when did improvements or leaps occur from those earlier tries to use computers to understand and imitate language, semantic language?
There is excellent information out there on the internet. Step by step I can uncover the path. As well as history, this involves concepts, code (python code) and maths.
First cab off the rank is “Bag of Words” (BOW) or properly speaking the Bag of Words model.
Bag of Words means take all the words from a document and count them. The order of the words is ignored. This turns out to be useful, more about that below.
I’ve imitated and adapted a BOW model for a Dr Seuss book:
Let’s suppose our documents have a small vocabulary. For instance, Dr. Seuss’ book Green Eggs and Ham has only fifty unique words. In alphabetical order, they are: a, am, and, anywhere, are, be, boat, box, car, could, dark, do, eat, eggs, fox, goat, good, green, ham, here, house, I, if, in, let, like, may, me, mouse, not, on, or, rain, Sam, say, see, so, thank, that, the, them, there, they, train, tree, try, will, with, would and you.
If we treat each page of the book as a single document, we can embed each of them as a 50-dimensional vector. Consider the page that reads:
I would not like them here or there.
I would not like them anywhere.
I do not like green eggs and ham.
I do not like them, Sam-I-am.
Summary of the process of making a BOW model:
- preprocessing: tokenize the sentences, removing punctuation and unnecessary spaces
- tokenize the words
- count the word frequencies
- filter out stop words if required (not done for the Dr Seuss case since not really necessary)
- build the BOW model: a binary matrix where each row corresponds to a sentence and each column represents one of the top N frequent words
- visualise the BOW model, in this case I made a heatmap. Other options are a frequency graph and a word cloud.
One of my current goals is to improve my python programming. So, I’ll document that as well. Usually I still feel like a python beginner, but am slowly, very slowly improving. My current goal is to become more fluent with list and dictionary comprehensions, lambda and a few others (RE, heapq, sorting, zip). I have learnt some new things about tokenization, REs, numpy, matplotlib and seaborn, the latter two for visualisation. But, I do need to learn more about seaborn since I’ve yet to figure out how to rotate the xticklabels and yticklabels.
Still not happy with my python coding skills but as I said, slowly improving.
CODEHere’s my python code, with comments for the BOW model. This particular code has been adapted from a couple of helpful sources, which are acknowledged in the doc string. I’ve left out the ordered frequency graph and Word Cloud. They are useful features but the BOW model is the main game.
# -*- coding: utf-8 -*- """ Created on Thu Oct 8 09:43:12 2026 @author: billk Dr Seuss story https://builtin.com/machine-learning/bag-of-words modified to create BOW visual with the help of https://www.geeksforgeeks.org/nlp/bag-of-words-bow-model-in-nlp/ """ import re # regular expressions import nltk # natural language tool kit import numpy as np import seaborn as sns import matplotlib.pyplot as plt # Assumes that 'doc' is a list of strings and 'vocab' is some iterable of vocab # words (e.g., a list or set) def get_bag_of_words(doc, vocab): # Create initial dictionary which maps each vocabulary word to a count of 0 word_count_dict = dict.fromkeys(vocab, 0) # For each word in the doc, increment its count for word in doc: word_count_dict[word] += 1 # Now, initialize a vector to a list of zeros bag = [0] * len(vocab) # For every vocab word, set its index equal to its count for i, word in enumerate(vocab): bag[i] = word_count_dict[word] return bag # Define the vocabulary of the whole document vocab = ['a', 'am', 'and', 'anywhere', 'are', 'be', 'boat', 'box', 'car',\ 'could', 'dark', 'do', 'eat', 'eggs', 'fox', 'goat', 'good', 'green',\ 'ham', 'here', 'house', 'i', 'if', 'in', 'let', 'like', 'may', 'me',\ 'mouse', 'not', 'on', 'or', 'rain', 'sam', 'say', 'see', 'so', 'thank',\ 'that', 'the', 'them', 'there', 'they', 'train', 'tree', 'try', 'will',\ 'with', 'would', 'you'] # Define one page of the document doc = ("I would not like them here or there.\n" "I would not like them anywhere.\n" "I do not like green eggs and ham.\n" "I do not like them, Sam-I-am.") # convert the doc string to list of sentences dataset = nltk.sent_tokenize(doc) #print(dataset) check it's working for i in range(len(dataset)): dataset[i] = dataset[i].lower() # eliminate capitals dataset[i] = re.sub(r'\W', ' ', dataset[i]) # eliminate non letters dataset[i] = re.sub(r'\s+', ' ', dataset[i]) # eliminate extra spaces # run next two lines to check it's working # for i, sentence in enumerate(dataset): # print(f"Sentence {i+1}: {sentence}") # initialise and make a word2count dictionary word2count = {} for data in dataset: words = nltk.word_tokenize(data) # tokenize each sentence # count the words for word in words: if word not in word2count: word2count[word] = 1 # add dictionary item else: word2count[word] += 1 # count the words print(word2count) # print the dictionary {'word' : num} # create BOW matrix graph BOW = [] for data in dataset: vector = [] for word in word2count: if word in nltk.word_tokenize(data): vector.append(1) else: vector.append(0) BOW.append(vector) # BOW is a list of lists BOW = np.asarray(BOW) # convert to numpy array # make a heatmap with seaborn plt.figure(figsize=(10, 6)) sns.heatmap(BOW, cmap='RdYlGn', cbar=False, annot=True, fmt="d", xticklabels=word2count, yticklabels=[f"Sentence {i+1}" for i in range(len(dataset))]) plt.title('Bag of Words Matrix') plt.xlabel('Frequent Words', rotation=0, ha="right") plt.ylabel('Sentences',rotation=0, ha="right") plt.tight_layout() plt.show()
Dr Seuss Bag of Words heat map (click to see more clearly):
USEFULNESSThis article summarises the surprising (given its simplicity) uses of BOW:
Although BoW is a very simple technique that turns documents into vectors it can be used by some machine learning algorithms for following use cases:
1. Plagiarism detection
If two documents produce the same or very similar BoW vectors, they likely contain the same words in similar proportions.
2. Spam detection
Emails containing words like “free”, “win”, “lottery” in high frequency can be flagged.Simple ML algorithms trained on BoW vectors still perform well for this use case.
3. Document classification
For example, news categories. Sports articles will have high counts for words like “team”, “win”, “match”. Politics articles will have high counts for words like “election”, “policy”, “minister”.
Issues arising from this research, to be elaborated on in future blogs:
- Other more efficient pathways: scikit-learn, there is always an improved way, TF-IDF to cull out trivial words
- Language classification algorithms, such as the naive Bayes classifier; how do they work? See the third reference below.
https://builtin.com/machine-learning/bag-of-words
Provides the code, which I adapted, for the Dr Seuss Green Eggs and Ham
https://www.geeksforgeeks.org/nlp/bag-of-words-bow-model-in-nlp/
Provides NLTK code for Bag of Words plus a few visualisation techniques (Frequency table and Word Cloud)
naive Bayes classifiers link
https://scikit-learn.org/stable/modules/naive_bayes.html
In spite of their apparently over-simplified assumptions, naive Bayes classifiers have worked quite well in many real-world situations, famously document classification and spam filtering. They require a small amount of training data to estimate the necessary parametersNLTK book
https://www.nltk.org/book/
Comprehensive account (concepts and python code) of Natural Language Processing (NLP) up until roughly 2010. Source of the NLTK python library.
