Assignment 17: Images as Data and Convolutions
Learning Objectives
- Identify and explain key components of a convolutional neural network (CNN)
- Create and apply filters like a CNN
- Calculate the output size and values resulting from a given filter
Meet Convolutional Neural Networks
Convolutional neural networks were a major step in the world of computer vision (and image generation). In class 17, we did some exploration of why these are cool and how they work. If you missed class, please review these materials. Now, you’ll spend some more time solidifying your understanding.
There are a huge number of resources out there. We suggest you look at two types:
- One that gives a high level overview and a visualization. We suggest the first one, but are providing a few other great options:
- This interactive visual overview of CNNs from a collaboration between Georgia Tech and Oregon State. This one will allow you to explore each of the layers and functions. You can click on each of the parts to see more. There’s a little video at the end that shows how to use the tool.
- This write-up with some helpful visualizations by Ujjwal Karn.
- One of the earlier types of these visualizations focused on handwritten numbers by Adam Harley.
- Training on MNIST in the browser by Karpathy. This one shows the weights and the gradients.
- This lecture by Serena Yeung of Stanford (part of one of the most famous academic AI labs) explaining convolutional neural networks. This lecture provides a little bit of history and does a nice job explaining some key terms and concepts, getting into the specifics and the math. This also gives you a taste of a classic academic lecture on this topic. Here’s the website from their class, which may also be a helpful but not required resource.
As always, you’re welcome to find alternative resources (and share them with everyone if they are awesome)!
For these first two exercises, we’d like you to attempt to recall the answers based on what you learned above (without immediately looking back at these resources). The act of trying to recall things from your memory helps slow the forgetting process (see this article by researcher Dr. Kathleen McDermott if you want evidence of this). Then you can check your answers with the resources and make your answers better.
Based on the materials above, explain the following terms/concepts:
- Convolution (conceptually and as a dot product)
- Filter size (F)
- Stride
- Padding (e.g., zero padding)
- Max pool
- ReLu
- Flatten
Describe the general architecture of a convolutional neural network for image classification. You don’t need to go into a lot of detail here, we just want to draw your attention to the major things that happen and the order that they happen in.
Part A
Given an input feature map of size 32 × 32 with a single channel, a filter size of 5 × 5, a stride of 1, and no padding, calculate the dimensions of the output feature map after a single convolution operation.
Part B
Repeat the above exercise for a filter of size 4 x 4. Why would we not want this filter?
For an RGB image of size 28x28, apply 6 different 7x7x3 filters with zero-padding of 3 and a stride of 1. What is the size of the output (give all dimensions)?
Given a grayscale image of size 64×64, apply a convolutional layer with the following parameters:
Filter size: 5×5
Stride: 2
Padding: 0 (no padding)
Calculate the size of the output feature map after applying this convolution.
Calculate the output from the following filter by hand (calculator fine). There is no padding for the image.
Filter:
\(\begin{bmatrix}
0 & -1 & 0 \\
-1 & 4 & -1 \\
0 & -1 & 0 \\
\end{bmatrix}\)
Image:
\(\begin{bmatrix}
10 & 0 & 10 & 0 \\
10 & 0 & 10 & 0 \\
10 & 10 & 10 & 10 \\
0 & 0 & 10 & 60 \\
\end{bmatrix}\)
In this notebook, you will create your own filters and apply them like they are part of a convolutional neural network. You will need to do a little research on filter types.
https://colab.research.google.com/drive/146PINGkBpFUGo8AhyEk8yNH5JFgOezXv?usp=sharing