Assignment 17: Images as Data and Convolutions

Learning Objectives

Learning Objectives
  • Identify and explain key components of a convolutional neural network (CNN)
  • Create and apply filters like a CNN
  • Calculate the output size and values resulting from a given filter

Meet Convolutional Neural Networks

Convolutional neural networks were a major step in the world of computer vision (and image generation). In class 17, we did some exploration of why these are cool and how they work. If you missed class, please review these materials. Now, you’ll spend some more time solidifying your understanding.

There are a huge number of resources out there. We suggest you look at two types:

  1. One that gives a high level overview and a visualization. We suggest the first one, but are providing a few other great options:
  2. This lecture by Serena Yeung of Stanford (part of one of the most famous academic AI labs) explaining convolutional neural networks. This lecture provides a little bit of history and does a nice job explaining some key terms and concepts, getting into the specifics and the math. This also gives you a taste of a classic academic lecture on this topic. Here’s the website from their class, which may also be a helpful but not required resource.

As always, you’re welcome to find alternative resources (and share them with everyone if they are awesome)!

For these first two exercises, we’d like you to attempt to recall the answers based on what you learned above (without immediately looking back at these resources). The act of trying to recall things from your memory helps slow the forgetting process (see this article by researcher Dr. Kathleen McDermott if you want evidence of this). Then you can check your answers with the resources and make your answers better.

Exercise 1

Based on the materials above, explain the following terms/concepts:

  • Convolution (conceptually and as a dot product)
  • Filter size (F)
  • Stride
  • Padding (e.g., zero padding)
  • Max pool
  • ReLu
  • Flatten
Exercise 2

Describe the general architecture of a convolutional neural network for image classification. You don’t need to go into a lot of detail here, we just want to draw your attention to the major things that happen and the order that they happen in.

Exercise 3

Part A

Given an input feature map of size 32 × 32 with a single channel, a filter size of 5 × 5, a stride of 1, and no padding, calculate the dimensions of the output feature map after a single convolution operation.

Part B

Repeat the above exercise for a filter of size 4 x 4. Why would we not want this filter?

Exercise 4

For an RGB image of size 28x28, apply 6 different 7x7x3 filters with zero-padding of 3 and a stride of 1. What is the size of the output (give all dimensions)?

Exercise 5

Given a grayscale image of size 64×64, apply a convolutional layer with the following parameters:
Filter size: 5×5
Stride: 2
Padding: 0 (no padding)
Calculate the size of the output feature map after applying this convolution.

Exercise 6

Calculate the output from the following filter by hand (calculator fine). There is no padding for the image.

Filter:
\(\begin{bmatrix} 0 & -1 & 0 \\ -1 & 4 & -1 \\ 0 & -1 & 0 \\ \end{bmatrix}\)

Image:
\(\begin{bmatrix} 10 & 0 & 10 & 0 \\ 10 & 0 & 10 & 0 \\ 10 & 10 & 10 & 10 \\ 0 & 0 & 10 & 60 \\ \end{bmatrix}\)

Exercise 7

In this notebook, you will create your own filters and apply them like they are part of a convolutional neural network. You will need to do a little research on filter types.

https://colab.research.google.com/drive/146PINGkBpFUGo8AhyEk8yNH5JFgOezXv?usp=sharing