In this first task we will use cross-correlation to find Waldo in the image below:

[ Image source: https://rare-gallery.com ]
Recall that cross-correlation, which the machine learning world often refers to as convolution is defined as:
for an image matrix
By sliding the waldo-kernel over the image and computing the cross-correlation at each position we can find the position where waldo is located by looking for the maximum value in the resulting matrix.
Navigate to the src/custom_conv.py module.
- Start in
my_conv_directand implement the convolution following the equation above. Both inputs aretorch.Tensors. Test your function with vscode tests ornox -s test. - Go to
src/waldo.py. It already usesmy_conv_directfor convolution. This script finds waldo in the image using your convolution function. Execute it withpython ./src/waldo.pyin your terminal. - If your code passes the pytest but is too slow to find waldo feel free to use
scipy.signal.correlate2dinsrc/waldo.pyinstead of your convolution function.
Navigate to the src/custom_conv.py module.
The function my_conv implements a fast version of the convolution operation above using a flattened kernel. We learned about this fast version in the lecture. Have a look at the slides again and then implement get_indices to make my_conv work. It should return
- A matrix of indices following the flattened convolution rule from the lecture, e.g. for a
$(2\times 2)$ kernel and a$(3\times 3)$ image it should return the index transformation
- The number of rows and columns in the result following
$$o=(i-k)+1$$ where$i$ denotes the input size and$k$ the kernel size.
If you need help, follow the hints below:
- First create a list of starting indices for each row in the output. These are the upper left corners of each kernel application.
- Then create a list of offsets within the kernel. These are the indices that need to be added to each starting index to get the full set of indices for each kernel application.
- Finally use broadcasting to add the two lists together and get the final index matrix.
The tests for this task (
test_conv_fastandtest_get_indices_readme_exampleintests/test_conv.py) are skipped automatically as long asget_indicesreturnsNone. Once you have implemented it, runnox -s testagain. Afterwards switchsrc/waldo.pytomy_convand run the script again. Watch the memory usage: the index matrix has one row per output pixel and one column per kernel pixel.
Open src/mnist.py and implement MNIST digit recognition with CNN in torch
-
Reuse your code from the yesterday's exercise on neural networks:
cross_entropy,sgd_step,zero_grad,get_accand the training loop. - In
cross_entropy,$n$ is the total number of entries of the label tensor, i.e. batch size$\times$ number of classes:$$\mathcal{L} = -\frac{1}{n}\sum_{k=1}^{n} \Big( y_k \log(o_k) + (1-y_k)\log(1-o_k) \Big).$$ -
get_accdoes not need gradients. Usetorch.no_grad()there, otherwise the evaluation on the 10000 test images builds a large autograd graph. - Reuse the
Netfrom the yesterday's exercise, add convolutional layers and pooling.torch.nn.Conv2dandtorch.nn.MaxPool2dwill help you. - Test your functions with
nox -s test.

