Saturday, October 10, 2026

Bag of Words model

What is and why make a “Bag of Words” model?

Like everyone, I’ve become interested in the epistemic nature of Large Language Models (LLM).

So, how to proceed? When you focus directly at LLMs they are complicated and difficult to grok.

I’ve opted for an evolutionary approach. I think of this as stepping stones to the LLM revolution. What significant artefacts came before? And then how, who, what, why and when did improvements or leaps occur from those earlier tries to use computers to understand and imitate language, semantic language?

There is excellent information out there on the internet. Step by step I can uncover the path. As well as history, this involves concepts, code (python code) and maths.

First cab off the rank is “Bag of Words” (BOW) or properly speaking the Bag of Words model.

Bag of Words means take all the words from a document and count them. The order of the words is ignored. This turns out to be useful, more about that below.

I’ve imitated and adapted a BOW model for a Dr Seuss book:

Let’s suppose our documents have a small vocabulary. For instance, Dr. Seuss’ book Green Eggs and Ham has only fifty unique words. In alphabetical order, they are: a, am, and, anywhere, are, be, boat, box, car, could, dark, do, eat, eggs, fox, goat, good, green, ham, here, house, I, if, in, let, like, may, me, mouse, not, on, or, rain, Sam, say, see, so, thank, that, the, them, there, they, train, tree, try, will, with, would and you.

If we treat each page of the book as a single document, we can embed each of them as a 50-dimensional vector. Consider the page that reads:

I would not like them here or there.

I would not like them anywhere.

I do not like green eggs and ham.

I do not like them, Sam-I-am.

Summary of the process of making a BOW model:

  • preprocessing: tokenize the sentences, removing punctuation and unnecessary spaces
  • tokenize the words
  • count the word frequencies
  • filter out stop words if required (not done for the Dr Seuss case since not really necessary)
  • build the BOW model: a binary matrix where each row corresponds to a sentence and each column represents one of the top N frequent words
  • visualise the BOW model, in this case I made a heatmap. Other options are a frequency graph and a word cloud.
PYTHON KNOWLEDGE

One of my current goals is to improve my python programming. So, I’ll document that as well. Usually I still feel like a python beginner, but am slowly, very slowly improving. My current goal is to become more fluent with list and dictionary comprehensions, lambda and a few others (RE, heapq, sorting, zip). I have learnt some new things about tokenization, REs, numpy, matplotlib and seaborn, the latter two for visualisation. But, I do need to learn more about seaborn since I’ve yet to figure out how to rotate the xticklabels and yticklabels.

Still not happy with my python coding skills but as I said, slowly improving.

CODE

Here’s my python code, with comments for the BOW model. This particular code has been adapted from a couple of helpful sources, which are acknowledged in the doc string. I’ve left out the ordered frequency graph and Word Cloud. They are useful features but the BOW model is the main game.

# -*- coding: utf-8 -*-
"""
Created on Thu Oct  8 09:43:12 2026

@author: billk
Dr Seuss story
https://builtin.com/machine-learning/bag-of-words
modified to create BOW visual with the help of
https://www.geeksforgeeks.org/nlp/bag-of-words-bow-model-in-nlp/

"""
import re   # regular expressions
import nltk # natural language tool kit
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt


# Assumes that 'doc' is a list of strings and 'vocab' is some iterable of vocab
# words (e.g., a list or set)
def get_bag_of_words(doc, vocab):
   # Create initial dictionary which maps each vocabulary word to a count of 0
   word_count_dict = dict.fromkeys(vocab, 0)
   # For each word in the doc, increment its count
   for word in doc:
       word_count_dict[word] += 1
   # Now, initialize a vector to a list of zeros
   bag = [0] * len(vocab)
   # For every vocab word, set its index equal to its count
   for i, word in enumerate(vocab):
       bag[i] = word_count_dict[word]
   return bag

# Define the vocabulary of the whole document
vocab = ['a', 'am', 'and', 'anywhere', 'are', 'be', 'boat', 'box', 'car',\
        'could', 'dark', 'do', 'eat', 'eggs', 'fox', 'goat', 'good', 'green',\
        'ham', 'here', 'house', 'i', 'if', 'in', 'let', 'like', 'may', 'me',\
        'mouse', 'not', 'on', 'or', 'rain', 'sam', 'say', 'see', 'so', 'thank',\
        'that', 'the', 'them', 'there', 'they', 'train', 'tree', 'try', 'will',\
        'with', 'would', 'you']
# Define one page of the document
doc = ("I would not like them here or there.\n"
      "I would not like them anywhere.\n"
      "I do not like green eggs and ham.\n"
      "I do not like them, Sam-I-am.")

# convert the doc string to list of sentences
dataset = nltk.sent_tokenize(doc) 
#print(dataset) check it's working

for i in range(len(dataset)):
    dataset[i] = dataset[i].lower()  # eliminate capitals
    dataset[i] = re.sub(r'\W', ' ', dataset[i]) # eliminate non letters
    dataset[i] = re.sub(r'\s+', ' ', dataset[i]) # eliminate extra spaces

# run next two lines to check it's working
# for i, sentence in enumerate(dataset):
#    print(f"Sentence {i+1}: {sentence}") 

# initialise and make a word2count dictionary
word2count = {} 

for data in dataset:
    words = nltk.word_tokenize(data) # tokenize each sentence
    # count the words
    for word in words:
        if word not in word2count:
            word2count[word] = 1 # add dictionary item
        else:
            word2count[word] += 1 # count the words

print(word2count) # print the dictionary {'word' : num}

# create BOW matrix graph
BOW = []

for data in dataset:
    vector = []
    for word in word2count:
        if word in nltk.word_tokenize(data):
            vector.append(1)
        else:
            vector.append(0)
    BOW.append(vector)

# BOW is a list of lists
BOW = np.asarray(BOW) # convert to numpy array

# make a heatmap with seaborn
plt.figure(figsize=(10, 6))
sns.heatmap(BOW, cmap='RdYlGn', cbar=False, annot=True, fmt="d", 
            xticklabels=word2count, yticklabels=[f"Sentence {i+1}" 
            for i in range(len(dataset))])

plt.title('Bag of Words Matrix')
plt.xlabel('Frequent Words', rotation=0, ha="right") 
plt.ylabel('Sentences',rotation=0, ha="right")
plt.tight_layout()
plt.show()

Dr Seuss Bag of Words heat map (click to see more clearly):

USEFULNESS

This article summarises the surprising (given its simplicity) uses of BOW:

Although BoW is a very simple technique that turns documents into vectors it can be used by some machine learning algorithms for following use cases:

1. Plagiarism detection

If two documents produce the same or very similar BoW vectors, they likely contain the same words in similar proportions.

2. Spam detection

Emails containing words like “free”, “win”, “lottery” in high frequency can be flagged.Simple ML algorithms trained on BoW vectors still perform well for this use case.

3. Document classification

For example, news categories. Sports articles will have high counts for words like “team”, “win”, “match”. Politics articles will have high counts for words like “election”, “policy”, “minister”.

Issues arising from this research, to be elaborated on in future blogs:

  • Other more efficient pathways: scikit-learn, there is always an improved way, TF-IDF to cull out trivial words
  • Language classification algorithms, such as the naive Bayes classifier; how do they work? See the third reference below.
REFERENCE
https://builtin.com/machine-learning/bag-of-words
Provides the code, which I adapted, for the Dr Seuss Green Eggs and Ham

https://www.geeksforgeeks.org/nlp/bag-of-words-bow-model-in-nlp/
Provides NLTK code for Bag of Words plus a few visualisation techniques (Frequency table and Word Cloud)

naive Bayes classifiers link
https://scikit-learn.org/stable/modules/naive_bayes.html
In spite of their apparently over-simplified assumptions, naive Bayes classifiers have worked quite well in many real-world situations, famously document classification and spam filtering. They require a small amount of training data to estimate the necessary parameters
NLTK book
https://www.nltk.org/book/
Comprehensive account (concepts and python code) of Natural Language Processing (NLP) up until roughly 2010. Source of the NLTK python library.

Sunday, July 26, 2026

Game of Life: colour, seeds, references

A Gosper gun which is formed from 2 queen bee shuttles back to back. The debris from the bee hive sparks creates a glider every cycle of 30 generations.

The Conway life appspot site has an extensive pattern library, with hundreds of interesting seeds, which are run with changing colours.

I couldn’t figure out how to do the changing colours so I asked on Stack Exchange and within hours had received a solution. It turns out that the changing colours at the Conway appspot site are based on the age of the cells:

  • green: new cells
  • yellow: one generation old
  • orange: two generations old
  • red: three generations old

The code below shows how and includes the Gosper gun seed.

I found an article by Paul Rendell about how to make a Turing Machine in the Cellular Automaton Conway’s Game of Life!

I also found a lexicon which explains all Game of Life terminology as well as many patterns with grids.

I spent some time making some of the patterns used to make the Turing Machine but eventually gave up because it was too much for me. By the way you can watch videos of the GOL Turing Machine here.

Golly is an open source, cross-platform application for exploring Conway's Game of Life and many other types of cellular automata: John von Neumann's 29-state CA, Wolfram's 1D rules, WireWorld, Generations, Paterson's Worms, Larger than Life, etc.

CODE, just with comments about the colour update (for full comments see previous post)
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
import numpy as np
import matplotlib.pyplot as plt
from matplotlib.animation import FuncAnimation
from matplotlib.colors import ListedColormap

N = 80 # array size
universe = np.zeros((N, N), dtype = int)

s= N//2 - 20 

# Gosper glider gun
universe[s, s+24] = 1
universe[s+1,[s+22,s+24]] = 1
universe[s+2,[s+12,s+13,s+20,s+21,s+34,s+35]] = 1
universe[s+3,[s+11,s+15,s+20,s+21,s+34,s+35]] = 1
universe[s+4,[s+0,s+1,s+10,s+16,s+20,s+21]] = 1
universe[s+5,[s+0, s+1, s+10,s+14,s+16,s+17,s+22,s+24]] = 1
universe[s+6,[s+10,s+16,s+24]] = 1
universe[s+7,[s+11,s+15]] = 1
universe[s+8,[s+12,s+13]] = 1

generations = 0
# initalise age
age = np.zeros(universe.shape)
colors = ['white', 'green', 'yellow', 'orange', 'red']
discrete_cmap = ListedColormap(colors)

def animate(frame, universe, img):
    global generations, age
    new_u = np.zeros((N, N), dtype = int)

    universe[0,:] = universe[-1,:] = universe[:,0] = universe[:,-1] = 0
    new_u[1:-1,1:-1] = (universe[ :-2, :-2] + universe[ :-2,1:-1] + universe[ :-2,2:] +
                        universe[1:-1, :-2]                + universe[1:-1,2:] +
                        universe[2:  , :-2] + universe[2:  ,1:-1] + universe[2:  ,2:])

    birth = (new_u==3)[1:-1,1:-1] & (universe[1:-1,1:-1]==0)
    survive = ((new_u==2) | (new_u==3))[1:-1,1:-1] & (universe[1:-1,1:-1]==1)
    universe[:] = np.zeros((N, N), dtype = int)
    universe[1:-1,1:-1][birth | survive] = 1
    population = np.count_nonzero(universe)
    
    age += universe # monitor age of cells
    age *= universe # remove dead cells
    age[age>4] = 4 # cap the age at 4
    img.set_data(age)
    ax.set_xlabel('generations = {}, population = {}'.format(generations, population))
    generations += 1


fig = plt.figure(figsize=(5, 5))
ax = plt.axes()
ax.set_yticklabels([])
ax.set_xticklabels([])
img = ax.imshow(universe, vmin=0, vmax=4, cmap=discrete_cmap)
ani = FuncAnimation(fig, animate, fargs=(universe, img,),
                    frames=220, interval=200)
ani.save("gosperGliderGun_colour.gif") # save the animations as a gif
plt.show()

Sunday, July 19, 2026

coding Game of Life

Code is at the bottom, with extensive comments

Notes on how to code Game of Life (GOL) including an animation, using numpy and matplotlib.

I learn through doing inspirational projects and Conway’s Game of Life cellular automata is certainly that. More about that in a later blog.

I’m also learning NumPy, as a step in learning the maths behind neural nets, so initially I’m following Nicolas P. Rougier’s online book “From Python to Numpy” (2017), Ch 4.2, where he has both a python and numpy version of Game of Life. The numpy version uses vectorization, rather than for loops.

The GOL rules are:

  • Survival: if an existing alive cell has 2 or 3 neighbours then it survives.
  • Birth: if an empty cell has 3 neighbours then it becomes alive.
  • All other neighbourhood counts lead to dead cells

Each cell has 8 neighbouring cells. By slicing the universe grid for each neighbouring cell and adding the values we obtain a neighbourhood count (new_u) for each cell. Then the rules can be implemented using boolean masks which reference both the new_u neighbourhood count grids and the universe grids.

Rougier doesn’t do the animation in his book. He stops after showing how to compute neighbours and does one iteration. So I found a version on the web by finxter using python for loops and FuncAnimation

The code works but the finxter explanation is cursory.

It is essential to make a new copy of the universe. Not doing this held my up for a while. Part of the problem was that I didn’t fully understand matplotlib’s FuncAnimation so I researched that for a bit.

Here is a good explanation:
The FuncAnimation class allows us to create an animation by passing a function that iteratively modifies the data of a plot. This is achieved by using the setter methods on various Artist (examples: Line2D, PathCollection, etc.). A usual FuncAnimation object takes a Figure that we want to animate and a function func that modifies the data plotted on the figure. It uses the frames parameter to determine the length of the animation. The interval parameter is used to determine time in milliseconds between drawing of two frames.
- source

Which setter method, plotting method and Artist do I need?

  • data set method: set_data()
  • plotting method: Axes.imshow()
  • Artist: image.AxesImage
FuncAnimation looks like this:

class matplotlib.animation.FuncAnimation(fig, func, frames=None, init_func=None, fargs=None, save_count=None, *, cache_frame_data=True, **kwargs)
  • makes an animation by repeatedly calling a function func
  • You must store the created Animation in a variable (called ani in my code) that lives as long as the animation should run. Otherwise, the Animation object will be garbage-collected and the animation stops.
  • func callable: The function to call at each frame. The first argument will be the next value in frames.
  • frames: Source of data to pass func and each frame of the animation
  • fargs: Additional arguments to pass to each call to func.
  • interval: delay between frames in milliseconds
  • repeat: bool, default True

Here is how to generate the animation gif which I used for this blog:
ani.save("GOL_random.gif")

Then I worked out how to count the population of each frame and the number of generations of the game and added that to the xlabel inside the animation function

Finally, I noticed that one Conway site had an extensive pattern library with hundreds of interesting seeds which can be run WITH DYNAMICALLY CHANGING COLOURS. I looked around at how to achieve this but so far haven’t worked it out.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
"""
Game of Life
"""
import numpy as np
import matplotlib.pyplot as plt
from matplotlib.animation import FuncAnimation

def create_universe(N=50, p=0.5):
    return np.random.choice([0, 1], size=(N, N), p=[1-p, p])
# generates random 1 and 0s with p probability in a N*N 2D array
N = 50 # array size

universe = create_universe(N=N, p=0.5)
# universe determines image

generations = 0 # measure the number of generations

# animation function
def animate(frame, universe, img):
    global generations
    new_u = np.zeros((N, N), dtype = int)
    # new_u initialises as same size as universe, array of all 0s 
    # then calculates neighbour count
    
    # set edges of universe to zero
    # to avoid edge calculation complications
    universe[0,:] = universe[-1,:] = universe[:,0] = universe[:,-1] = 0
    # universe is 2D array with 1s and 0s
    
    # count the neighbours of the universe cells (by grid displacement)
    # new_u is neighbour counts
    new_u[1:-1,1:-1] = (universe[ :-2, :-2] + universe[ :-2,1:-1] + universe[ :-2,2:] +
                     universe[1:-1, :-2]                + universe[1:-1,2:] +
                     universe[2:  , :-2] + universe[2:  ,1:-1] + universe[2:  ,2:])

    # boolean masks to implement GOL rules
    birth = (new_u==3)[1:-1,1:-1] & (universe[1:-1,1:-1]==0) 
    # 3 neighbours for an empty cell -> birth
    survive = ((new_u==2) | (new_u==3))[1:-1,1:-1] & (universe[1:-1,1:-1]==1)
    # 2 or 3 neighbours in occupied cell -> survive
    
    # reset universe to zeros, make a new copy, then update 1s & 0s
    universe[:] = np.zeros((N, N), dtype = int) # new copy essential
    universe[1:-1,1:-1][birth | survive] = 1
    # count the population
    population = np.count_nonzero(universe)
    # update image with the data set method
    img.set_data(universe)
    # display generations and population
    ax.set_xlabel('generations = {}, population = {}'.format(generations, population))
    generations += 1 # update generations

fig = plt.figure(figsize=(5, 5)) # figure size
ax = plt.axes()
ax.set_yticklabels([]) # turn off axis ticks (but keep x axis label above)
ax.set_xticklabels([])
img = ax.imshow(universe, interpolation='nearest')
# imshow is the plotting method for AxesImage artist
ani = FuncAnimation(fig, animate, fargs=(universe, img,),
                    frames=200, interval=200)
# ani lives while animation runs, otherwise garbage collected
# fargs: arguments to pass to each call of the function 
# interval in milliseconds
ani.save("GOL_random.gif") # save the animations as a gif
plt.show()

Thursday, January 01, 2026

Books 2026

Books, videos (and long articles) I am reading / watching in 2026:

Ananthaswamy, Anil. Why Machines Learn: The Elegant Math Behind Modern AI (2024)
Banks, Iain M. Excession (1996)
Backman, Fredrik. A Man Called Ove (2014)
Bennett, Max. A Brief History of Intelligence: Why the Evolution of the Brain Holds the Key to the Future of AI (2023)
Bhattacharya, Ananyo. The Man from the Future: The Visionary Life of John van Neumann (2021)
Bird Steven, Klein Ewan, Loper Edward. Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit
Boldizar, Alexander. The Man who saw Seconds (2024)
Brooks, Rodney. Predictions Scorecard, 2026 January 01
Brooks, Rodney. Why Today’s Humanoids Won’t Learn Dexterity (September 2025)
Brooks, Rodney. Intelligence without Reason (1991)
Brooks, Rodney. Elephants Don't Play Chess (1990)
Curry, Judith. Climate Uncertainty and Risk (2023)
Dennett, Daniel. Intuition Pumps and other Tools for Thinking (2013)
Dennett, Daniel. From Bacteria to Bach and Back: The Evolution of Minds (2017)
Disher, Garry. The Divine Wind (1998)
Eisenberg, David. SVG Essentials. (2002)
Epstein, Alex. Fossil Future (2022)
Goldstein, Rebecca. The Mattering Instinct (2016)
Haverbeke, Marijn. Eloquent JavaScript (Fourth Edition). (2024)
Holson, Benjie. Benjie's Humanoid Olympic Games (September 2025)
Lawhon, Ariel. The Frozen River (2024)
Larson, Erik. The Myth Of AI: Why Computers can't think the way we do (2021)
Larson, Erik. Benderland (2026)
Lee, Tim & Trott, Sean. A jargon-free explanation of how AI large language models work (2023)
Mearsheimer, John J. The Darkness Ahead: Where The Ukraine War Is Headed (June, 2023)
Mesbahi, Farzed. Abundance or Collapse (2026)
McKinney, Wes. Python for Data Analysis, 3E (2022-23)
North, Claire. Slow Gods (2025)
Pfarrer, Chuck and Brewer, Alan. Indications and Warnings | Day 1418 | Putin's Dubious Milestone (Jan 14, 2026)
Raeini, Mohammad Ghaseminejad. The evolution of language models: From N-Grams to LLMs, and beyond (2025)
Ramalho, Luciano. Fluent Python: Clear, Concise and Effective Progamming, second edition (2022)
Robinson, Kim Stanley. Blue Mars (2009)
Rowson, Jonathan. Understanding the Grünfeld (1998)
Rougier, Nicolas P. From Python to Numpy (2017)
Schoch, Richard. The Secrets of Happiness (2006)
Seth, Anil. Being You: A New Science of Consciousness (2021)
Stewart, James. Calculus Fourth Edition (1999)
Stanley, Kenneth; Lehman, Joel. Why Greatness cannot be Planned: The Myth of the Objective (2015)
Akarsh Kumar, Jeff Clune, Joel Lehman, Kenneth O. Stanley. Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis (2025)
Tchaikovsky, Adrian. Children of Time (2015)
Tchaikovsky, Adrian. Children of Ruin (2019)
Tchaikovsky, Adrian. Children of Memory (2023)
Trask, Andrew. Grokking Deep Learning (2019)
Whitehead, Colson. Underground Railroad (2016)
Yann LeCun, Yoshua Bengio, Geoffrey Hinton. Deep Learning (2015)
Zahavy,Tom. LLMs can't Jump (2026)
CODING
Plotly Open Source Graphing Library for Python
Di Méo, Grace. Creating Interactive Visualizations with Plotly
kaizenkode, Python Programming Projects for Beginners (2026)
seaborn
pandas
matplotlib
NumPy: the absolute basics for beginners
Python NumPy For Your Grandma by Ben Gorman
NumPy tutorials
Machine Learning Crash Course
GAME OF LIFE
LifeWiki
The Game of Life
Conway's Game of Life
A Universal Turing Machine in Conway’s Game of Life
Golly

Previous: Books and articles 2025

Friday, November 07, 2025

innovative school curriculum reform

SUBJECTS for CURRICULUM INNOVATION

As a teacher I have developed some innovative pathways to introduce students to the three digital revolutions: computation, communication, fabrication.

1) Turtle art, including the Turtle Art tile project
Turtle art is software which produces beautiful patterns with a simplified coding interface.

This can be a standalone or lead into the Turtle Art tile project which involves some significant transformations from Turtle Art design to Tinkercad to 3D prints and then to a painted clay product.

2) Scratch multimedia story telling
Scratch multimedia coding is for telling stories in an interesting and entertaining way. I have developed several iterations of an 18 lesson course I have taught to Year 7s. Any story can be told.

3) Fabulous Fabrication
Students cycle through the learning of some basics, design, making, collaboration, reflection, debugging and presentation using these materials: microbits, neopixel strips, servos, cardboard, building materials and tools, 3D printers. When I trialled this approach recently with year 8s the sort of things they decided to build with moving parts were a complete exo skeleton, a submarine made from geodesic domes, a sword and scythe weapon set, a mini computer, a dancing cactus and a couple of others.

4) AI pathways
This is the latest big thing. I have found and developed a variety of resources suitable for secondary students. Dale Lane, an IBM developer, has integrated Machine Learning with a Scratch User Interface. Ken Kahn has developed materials using Large Language Models to make web apps, games and more. This can also serve as an introduction to web site design with HTML, CSS and JavaScript. Google’s Teachable Machine is another resource accessible to school students.

5) Chess
I have worked as a chess coach. Chess can promote a variety of sought after skills such as complex decision making, logical, rational thinking, time management, mental toughness, resilience, will power, determination and persistence.

6) Drones
I am familiar with the Tello drone and have also built a drone (the Air:bit) from a kit which uses microbits developed by WonderKit Technology, a company in Norway. There are a variety of options here in a rapidly developing field.

7) Python coding
PyGame as an engaging way for students to learn python. Python can also offer a pathway to data science through Jupyter, numpy, pandas and matplotlib.

8) Website design with HTML, CSS and JavaScript

9) App Inventor provides block code (easier to learn) for making apps for any phone

Monday, September 08, 2025

Machine learning for kids (Dale Lane site)

Web site: https://machinelearningforkids.co.uk/

What, why and who

Machine Learning for Kids (MLK) is the name of an AI project based education program developed by Dale Lane. He does continually update it since its inception in 2017. In this fascinating blog post Dale writes about how things happened and what made him do them.

Machine Learning (ML) is one type of AI. I’ll explain at the end how it fits into the bigger AI picture.

Dale has integrated his ML tools as Scratch extensions to provide an easy to use interface and kid friendly experience. Find this Scratch fork at https://machinelearningforkids.co.uk/scratch/. This is a great idea to help kids learn AI by making (actually training not making the whole thing, which is beyond the scope of kids) an AI in a familiar User Interface. This fits well with the constructionist learning concept that we understand a thing by (in this case partially) making that thing.

Project templates page of Dale's Scratch fork

On his About page Dale provides a What? and Why? he did it.

What?

“It provides an easy-to-use guided environment for training machine learning models to: recognise text, numbers, images, sounds; predict numbers; generate text.”
Why?
“Machine learning is all around us. We all use machine learning systems every day - such as spam filters, recommendation engines, language translation services, chatbots and digital assistants, search engines, and fraud detection systems."

It will soon be normal for machine learning systems to drive our cars, and help doctors to diagnose and treat our illnesses.

It's important that kids are aware of how our world works. The best way to understand the capabilities and implications is to be able to build with this technology for themselves (my emphasis).”

On the same About page find two videos:

Education in the age of AI (Artificial Intelligence) | Dale Lane | TEDxWinchester, 16 min (2023)
This one is an overview of the What? When? and How? explaining the importance of teaching AI in schools.

Machine Learning for Kids (2019), 22 min
This one takes you through the Smart classroom project where the program learns to recognise a variety of commands that turn a fan or lamp on or off. The program is first trained by the students and learns to recognise other commands which have not been in the training set.

Who?

Dale is an IBM developer which gives him ready access to watsonx, a further development of the famous watson which won Jeopardy in 2011, a significant milestone in the history of AI: https://en.wikipedia.org/wiki/IBM_Watson. Some of his projects, eg. I Spy, utilise the watson API. Dale writes a blog about a variety of topics. Search his blog with "Scratch" or "Machine learning" or "mlforkids-tech" for a more detailed account of the topics discussed here.

The Projects and Worksheets

Dale has developed a comprehensive series of worksheets which take you through step by step about how to complete a project. All worksheets are released under the Creative Commons Licence :-)

The worksheets are categorised under project types (recognise text, images, numbers, sounds, faces; predict number or generate text), difficulty level and whether they are Scratch or Python projects. At this stage I’ve completed about 20 projects (out of 52). Usually, but not always, they run smoothly and produce fascinating outcomes.

Examples which I have completed so far

For more detail go to the worksheets page:
recognise text: Smart Classroom, Make me Happy, Quiz Show recognise images: Judge a Book, I Spy, Shy Panda, CAPTCHA
recognise numbers: Pac-Man, Noughts and Crosses
predict numbers: Catch the Ball
recognise sounds: Alien sounds, Shoebox
recognise faces: Face Finder
generate text: Language models, Story Teller

70 years of AI

Working and reflecting through the projects became part of my own education about AI, how it works and how accessible it is to a curious non expert such as myself. We now live in a world of big data and machine learning pretrained models that do useful things.

The main thing we hear about these days is the hype and commercialisation around Large Language Models (LLMs). However, AI has been around for almost 70 years now and most of it has been in niche fields. This is where Dale’s site is so valuable. It provides an evolutionary, historical perspective. He has developed a bunch of fun games and activities around the media themes listed above (recognise text, images etc). For example, we can train a model to play Pacman by playing Pacman ourselves. As we play, the model tracks the co-ordinates of the player and the ghost and learns to play by itself.

By the way, Dale has also kept up to date with recent developments and as noted above has models that generate text, similar to chatGPT etc. but with a difference. As well as the commercial LLMs out there there are also many Small Language Models (SLMs) that don’t require such powerful (GPU) processors.

Story Teller: Of course I had to try the Small Language Model activity since that is currently in favour. I learnt that there are several Small Language Models out there a range of which are made available at the site, eg. “SmolLM2” (made by Hugging Face), Llama 3.2 (made by Meta) and there are others. I learnt that you can play around with the probability and temperature settings and how that affects the responses to a prompt. Each time you trigger a given prompt you will get a different story. I had a lot of fun with this one putting in characters I knew and suggesting a general story line.
Aside: What will English teachers set for homework from now on? Given that the model generates a different story each time I think the plagiarism checkers have had their day.

Workflow

The workflow can vary. Sometimes you go straight to a Scratch project page. Sometimes you start a new project. Sometimes you load a project template. It’s all laid out in the worksheets for each project.

Most of the projects have this screen in common. This particular one is for the Smart Classroom project, where you train the Scratch app to turn a light or fan on or off with text commands.

Train mode: You add a diverse variety of text instructions to category boxes: lamp_on, lamp_off etc. This can be arduous and I would anticipate some students saying “boring” for this mode.

Learn and Test mode: This screen reminds you of how many items you have entered into each category box and invites you to train a new machine learning model. After you train it you can then do some confidence testing (reported as a %) on new instructions which differ from the training set.

Make mode: This screen transitions you normally to Scratch 3 (or occasionally to Python or App Inventor) where you can load a Project template and test the model in an activity or game.

Models and Templates everywhere

I became interested in all the non commercial free AI models out there. They are everywhere and Dale has done a great job of finding and using them to generate a rich set of activities. Some examples to show the diversity of models used:

  • a model which uses data from wikipedia to develop a Quiz game
  • Use IBM® watsonx™ Assistant to build your own live chatbot
  • use the OpenLibrary API to enable access to information about books

and much more ...

This passage from one of Dale’s blog provides the big picture view of what he has achieved:

“Most of the work I do on Machine Learning for Kids involves adding machine learning models into Scratch. To enable students to create interesting projects, it also helps to make it easier to get external data into Scratch that they can use for training and classifying. A few examples of where I’ve done this in the past include creating Scratch blocks to access weather data, data from Spotify, and data from Wikipedia.”
https://dalelane.co.uk/blog/?p=5244
Pretrained models

Going back to the MLK Scratch site, if you open the Extensions page you find a variety of pretrained models. While for some of the projects students train their own models (aka supervised training, so they get to learn what is involved in training a model) it’s also a good idea to have pretrained models so the students can quickly move onto the more interesting aspects of ML. My abbreviated descriptors of the pre trained models below gives you some idea of the wide scope of possible narrow AI apps that has developed over time. I feel that all the hype surrounding LLMs tends to obscure this.

Some of the pretrained models on the Scratch Extensions page

There are pretrained models for:
Speech to Text: Google Chrome only
Face detection: find the x,y coordinates of your eyes, nose and mouth.
Pose detection: find the x,y coordinates of different parts of your body, like shoulders, elbows, wrists, knees, and ankles.
Hand detection: find the x,y coordinates of different parts of your hand: the tips of each of your fingers, and your wrist.
Toxicity: predict the percentage probability that some provided text contains toxic content such as threatening language, insults, obscenities, or identity-based hate.
Imagenet: will predict the main object shown in a sprite … (based on MobileNet (a ML model designed for mobile devices, so it doesn't need much computing power, more details here)
Question answering: It is a type of machine learning model called BERT which is useful for projects with text ... It has been trained using a set of questions and answers from Wikipedia articles collected by Stanford University called 'SQuAD'.
Pitch estimation: gives you blocks that will return the frequency of a note it recognized, and to convert that into the name or MIDI note

There is even a TensorFlow model for more advanced use. So Dale Lane’s site leads into more advanced aspects of AI. He doesn’t leave that out just because it is focused on kids.

There are more models at Scratch > Extensions including Wikipedia, Weather, Books, Voice tuner and MQTT

And then at the Scratch Project templates tab you will find another multitude under the categories: Text, Images, Numbers, Sounds, Regression

Stories and Learning Outcomes

Under the stories tab Dale uses stories to illustrate what students will learn from his program. For example:

Machine learning hasn’t replaced the need to code
Students see that machine learning adds new tools to their existing toolbox. They see that what they've been learning about coding is still important and valuable, and that machine learning expands the types of things they're able to build. This is followed by a story which illustrates this principle.
Crowd sourcing and gamification can help to generate training data

The boring side of AI is having to type in all the training data yourself. To overcome this problem Dale has developed some projects where the whole class does the training data.

There are many other story lessons listed here. You could read as “learning outcomes” from completing the projects.

Problems

Entering data can be arduous (boring), this varies from project to project depending on the type of data you need to enter.

One of the projects (Quiz Show) required too much memory for my computer

Some of the projects (eg. Shoebox) were flakey in that I would say one number clearly and two numbers were entered into the dialog. But the general issue of often requiring more data / training is not so much a problem but an opportunity to explain to students that ML works better with more training.

Practicalities

There are guidelines for teachers on this page.

I’d have to see how that pans out in practice:
  • setting up student IDs quickly
  • site reliability
  • ability to keep track of student progress
ML as one type of AI
ML is a subset of AI. I have previously blogged about this: an AI taxonomy

This diagram, found on the web, is missing the two robotic forms of AI and the Neuro-Symbolic hybrids

ML is learning by example. In some cases lots of examples. It improves its performance with training over time. With machine learning, we use algorithms that have the ability to learn.

You can have forms of AI that don’t learn over time. Symbolic AI, Traditional robotics and Behaviour based robotics could all fit this category. They are programmed, they do some human like stuff but don’t change or improve over time without human reprogramming. They are still important but sidelined at the moment due to the LLM (Large Language Models) hype.

To repeat, Machine Learning is a subfield of AI that uses algorithms trained on data to produce adaptable models that can perform a variety of complex tasks. Deep learning is a subset of machine learning that uses several layers within neural networks to do some of the most complex ML tasks

So we have here with Dale Lane's site an introduction to one type of AI, the type that has received the most attention lately.

Sunday, June 08, 2025

machine learning explained

Below is a summary of an article by Rodney Brooks, Machine Learning Explained (2017).

This inspiried me to take on the development of my own version of hexapawn (referenced below) in python. After reading the Rodney Brooks article I did search around and found a hexapawn version written in Scratch by puttering. The rules are outlined there. The machine improves their play as it plays. I played 30 games with it. For the first 10 I won 6 and the machine 4. For the next 10 I won 4 and the machine 6. For the final 10 I won 2 and the machine 8. I hasten to add that in this game black has the advantage and with perfect play should win 0 to 10.

Anyway Machine Learning is all the rage now so I'll summarise large section of the Brooks article. Of course, you should read the whole thing:

Machine Learning

  • is what has enabled the new assistants in our houses such as the Amazon Echo (Alexa) and Google Home by allowing them to reliably understand as we speak to them.
  • is how Google chooses what advertisements to place, how it saves enormous amounts of electricity at its data centers, and how it labels images so that we can search for them with key words.
  • is how DeepMind (a Google company) was able to build a program called Alpha Go which beat the world Go champion.
  • is how Amazon knows what recommendations to make to you whenever you are at its web site.
  • is how PayPal detects fraudulent transactions.
  • is how Facebook is able to translate between languages. And the list goes on!

Machine Learning is not magic.

Every successful application of ML is hard won by researchers or engineers carefully analyzing the problem that is at hand. They select one or many different ML algorithms, and custom design how to connect them together and to the data. In some cases there is an extensive period of training on very large sets of data before the algorithm can be run on the problem that is being solved. In that case there may be months of work to do in collecting the right sort of data from which ML will actually learn. In other cases the learning algorithm will be integrated in to the application and will learn while doing the task that is desired–it might require some training wheels in the early stages, and they too must be designed. In any case there is always a big design project about how, when the ultimate system is operational, the data that comes in will be organized, processed and mapped before it reaches the ML component of the system.

Alan Turing was assisted by Donald Michie in developing the code breaking Colossus computer Bletchley Park during WW2 which helped shorten the war against fascism

After the war Arthur Samuel developed a machine that could play draughts from 1952-56, the first AI in the USA.

Samuel wondered whether the improvements he was making to the program by hand could be made by the machine itself.

What Samuel had realized, demonstrated, and exploited, was that digital computers were by 1959 fast enough to take over some of the fine tuning that a person might do for a program, as he had been doing since the first version of his program in 1952, and ultimately eliminate the need for much of that effort by human programmers by letting the computer do some Machine Learning on appropriate parts of the problem. This is exactly what has lead, almost 60 years later to the great influence that ML is now having on the world.

Samuel explored two (machine) learning techniques:

  1. Memoization or memoisation is an optimization technique used primarily to speed up computer programs by storing the results of expensive function calls and returning the cached result when the same inputs occur again
  2. The other learning technique that he investigated involved adjusting numerical weights on how much the program should believe each of over thirty measures of how good or bad a particular board position was for the program or its human opponent. This is closer to how ML works today.

Donald Michie himself built a machine that could learn to play the game of tic-tac-toe (Noughts and Crosses in British English) from 304 matchboxes, small rectangular boxes which were the containers for matches, and which had an outer cover and a sliding inner box to hold the matches. He put a label on one end of each of these sliding boxes, and carefully filled them with precise numbers of colored beads. With the help of a human operator, mindlessly following some simple rules, he had a machine that could not only play tic-tac-toe but could learn to get better at it. He called his machine MENACE, for Matchbox Educable Noughts And Crosses Engine, and published⁠5 a report on it in 1961. A matchbox computer is still a computer!

In 1962 Martin Gardner⁠ reported on it in his regular Mathematical Games column in Scientific American, but illustrated it with a slightly simpler version to play hexapawn, three chess pawns against three chess pawns on a three by three chessboard. … Gardner suggested that people try building a matchbox computer to play simplified checkers with two pieces for each player on a four by four board.

Rodney Brooks recounts how he eventually came to realise that this was a wonderful way for explaining how machine learning works.

Today people generally recognize three different classes of Machine Learning, supervised, unsupervised, and reinforcement learning, all three very actively researched, and all being used for real applications. Donald Michie’s MENACE introduced the idea of ML reinforcement learning, and he explicitly refers to reinforcement as a key concept in how it works

Footnote:
I explain in this article, an AI taxonomy, the difference between Machine Learning (ML) and AI. You can have forms of AI that don’t learn over time. Symbolic AI, Traditional robotics and Behaviour based robotics could all fit this category. They are programmed, they do some human like stuff but don’t change or improve over time without human reprogramming. They are still important but sidelined at the moment due to the LLM (Large Language Models) hype.

Wednesday, June 04, 2025

my crystal neopixel lamp

The Concept:

This started from the idea of making a lamp with flashing, coloured LEDs inside. The inspiration here was seeing such a lamp made by Robin at Hackerspace for his grand-daughter. I learnt then about the ESP32 board which enables control of the flashing LEDs, aka Neopixels, from your phone.

Previously, I had made a Sierpinkski sieve a Sierpinski pyramid lamp with transparent PLA filament and lit it up with a couple of Circuit Playgrounds in the base. Up until then that was my favourite make! This background partially determined my direction for this project.

So, I found a crystal lamp on printables, here, which had been partially hollowed out and printed it with transparent filament.

This pic is from Printables: I plan to improve it with changing coloured neopixels that go to the top of the central pillar (see video near the bottom of this article)

Then began a complex process of research and buying of materials. I had to determine what to buy and the voltage and current requirements. Adafruit has a comprehensive uberguide about NeoPixels. Some helpful pointers from the adafruit uberguide were:

  • NeoPixels are usually described as “5 Volt devices” ... (but) ...Lower voltages are always acceptable, with the caveat that the LEDs may be slightly dimmer. There’s a limit below which the LED will fail to light, or will start to show the wrong color.
  • NeoPixels don’t care what end they receive power from. Though data moves in only one direction, electricity can go either way. You can connect power at the head, the tail, in the middle, or ideally distribute it to several points.

I also consulted with Robin at Hackerspace about which strips and ESP32 to purchase.

My plan became something like this:
  • purchase a cuttable RGB+IC 144LEDs/m WS2812B LED Strip. I then cut 3 strips, cutting at the copper dots, each 14 LEDs long and arranged them around a triangular prism so there will be lots of flashing lights from all directions.
  • design and 3D print a hollow triangular prism support for the LED strips after they had been cut
  • design and print a base with a cavity to hold the batteries and the ESP32 board
  • purchase a SuperMini ESP32-S3, 23mmx18mm, keeping it small so as to keep the cavity small
  • power the whole thing with 3xAAA batteries in a cylindrical holder (55mm long x 22mm diameter), amounting to 4.5 volts

In retrospect, this is or was a reasonable plan. Of course there are alternatives, eg. rather than the bulky battery holder run a cable to a 5v power supply. I was muddling through. Doing things for the first time is always hard.

Understanding the circuit and how the LED strip works

Initially I didn’t understand how the Neopixel strips (WS2812) worked. I hadn't grasped the significance of the comment in the adafruit uberguide, that "NeoPixels don’t care what end they receive power from." As it turns out each LED has its own processor. One side is positive, the other side negative and the data flows in one direction marked by an arrow on the strip. My understanding now is that the circuit(s) are made up on the fly as the current flows across through each LED.

Triangular prism designed with SCAD

I wanted to upgrade my 3D design skills so this time I opted to learn SCAD, which has a mathematical approach to design. I found it to be fairly intuitive after looking at a couple of beginner’s tutorials.

SCAD design of triangular prism onto which the LED strips are mounted (I printed this with transparent PLA filament too

Fastening the LED strips, wiring up, soldering and testing

I wanted a transparent hollow triangular prism to fasten the LED strips onto, with small holes at the top and bottom to run wires through to the inside. I designed this with SCAD. It was surprisingly easy and elegant!

I cut three strips, cutting at the copper dots 14 LEDs / strip. I wired up the positives at one one end and the GND and data the other end. The idea here was to avoid clutter at one end.

It took me a while and help from Robin to understand the circuit and data flow

  • the circuit goes through each LED (where is this explained?)
  • the arrows on the strip show the data direction

For stripping the wires I needed the right tool, although the experts can do it by feel. Before soldering I removed one LED from each strip with a heat gun. I did this because the copper dots when cut in half were very small.

I tested the strip using my Sunfounder kit and it was successful!

Design and redesigns of base with SCAD

I had to do a few redesigns along the way after I decided to insert a relatively bulky cylindrical battery holder (55mm long x 22mm diameter) into the base.

SuperMini ESP32-S3

I found helpful information about the SuperMini ESP32-S3 Development Board here.

  • some simple tests to check if the board was working
  • pin layour and which pins were safe to send data through

Then I had to figure out how to flash micropython onto the board. Once again help was there online: Instructions

Coding in micropython

I opted for coding in micropython which I find easier to understand than arduino C++. I bought a Sunfounder ESP32 starters kit, installed Thonny IDE and started working through their online micropython tutorials. They did have a micropython WS2812 tutorial for an 8 LED strip, which gave me a good start for the code.

I thought an effect where the colour grows stronger then fade and then changes to another colour would look good. It took me a while to figure out how to modify the code so it was both elegant and did what I wanted. This involved writing efficient functions and looking up the RGB values for different colours.

The Thonny code tested alright on the LED strips but then I had to figure out how to run it from the ESP32 board. With a Save As... I could download it to the board and then found an article which explained that it had to renamed main.py and then it would run.

There is also the WLED option which provides tremendous variety but you miss out on the joy of coding ;-)

Not the final version, but it does show the neopixels burning brightly!

Trouble shooting / Problem solving: Suitable wire thickness.

You need the right materials. The experts can wing it but mere mortals like me need the right materials.

"For the lack of a nail a kingdom was lost" - Shakespeare
For the lack of 26 AGW wire a crystal neopixel lamp was compromised - Kerr

During my near the end soldering session I suddenly discovered that two out of three LED strips no longer lit up. The problem was that I had used very thin wires and with a little twisting some of them broke. So, I went back to replacing the broken wires with thicker ones (22 AGW) and resoldering. But my problems persisted because now the wires were too thick and it is hard to twist three thick wires on the one side and solder them successfully to one thick wired on the other side. So, what I have planned here is to get some 26 AGW (in between thickness) wires and do the whole thing again!

Also my design lacked an important component given that I'm powering off three AAA batteries. A switch! This will be part of my new model

This experience has taught me to improve my soldering skills (use the third hand, flux and the solder sucker when required) and design skills (the wiring was far from elegant). Such is life. I've learnt a lot.

SUMMING UP: WHAT HAVE I LEARNT
  • Thonny python IDE upgrade and use
  • micropython for loops and functions
  • WS2812 Neopixel strip function (power, data), cutting procedure
  • Supermini ESP32 function and testing
  • OpenSCAD skills (designed and made tri prism and base)
  • Circuit or wiring design
  • Soldering skills (wire to wire, stripping, tinning, third hand, solder sucker, heat shrink)
  • Heat gun, to remove one of the LEDs from each strip
  • Problem solving and receiving help
  • Checking parts carefully before buying (eg. bought momentary switches)
REFERENCE