{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# CS 124 Tutorial: `NumPy`\n",
    "---\n",
    "\n",
    "Based on the `CS 124: Jupyter and Python Tutorial` created by \n",
    "`Krishna Patel (Winter 2020)`, and updated by `Bryan Kim (Winter 2021)`,  \n",
    "`Dilara Soylu (Winter 2022)`, and `Sri Jaladi (Winter 2026)`.\n",
    "\n",
    "Some examples based on the \n",
    "[CS 231n Python Numpy Tutorial (with Jupyter and Colab)](https://cs231n.github.io/python-numpy-tutorial/) by Justin Johnson. "
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='overview'></a>\n",
    "## Overview\n",
    "\n",
    "In this tutorial, we will walk you through some `NumPy` examples as a \n",
    "preparation for our second assignments, `PA 2`.\n",
    "`NumPy` is a very popular `Python` library used for matrix operations and linear\n",
    "algebra.\n",
    "The purpose of this notebook is to give a basic introduction to `NumPy` for \n",
    "students who haven't used it before, and an easy review for those who have.\n",
    "Learning `NumPy` is well worth the effort, as you will be using it constantly if\n",
    "you choose to take further ML/AI courses at Stanford.\n",
    "\n",
    "This notebook is optional and ungraded, so if you have worked with `NumPy` \n",
    "before and feel comfortable working with it, feel free to skim it or skip it."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "## Linear Algebra Review\n",
    "\n",
    "While this review does not cover linear algebra, we *HIGHLY RECOMMEND* you review some basic linear algebra techniques and fundamentals (vectors, dot products, matrix mulitplications, etc.) for the purposes of the class.\n",
    "\n",
    "If you would like to review linear algebra concepts the class expects you to know, you can use the following resources:\n",
    "\n",
    "* [3blue1brown Linear Algebra Lectures](https://www.3blue1brown.com/topics/linear-algebra) \n",
    "* We recommend lectures 1, 2, 3, and the dot product lecture\n",
    "* [Khan Academy Linear Algebra Lectures](https://www.khanacademy.org/math/linear-algebra),\n",
    "* We recommend the \"Vectors\", \"Linear Combinations\", and \"Vector Dot Products\" videos\n",
    "* Ask an LLM to tutor you in linear algebra for computer science, focusing on vectors, matrices, dot products, and cosine similarity."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='contents'></a>\n",
    "## Contents\n",
    "\n",
    "1. [Environment Check](#environment_check)\n",
    "2. [`NumPy` Exercises](#regular_expressions_exercises)\n",
    "   * [Part 1. Basic `NumPy`](#basic_numpy)\n",
    "   * [Part 2. Indexing and Slicing](#indexing_and_slicing)\n",
    "   * [Part 3. Array Math and Functions](#array_math_and_functions)\n",
    "   * [Part 4. Vectorization](#vectorization)\n",
    "   * [Part 5. Broadcasting](#broadcasting)\n",
    "   * [Part 6. Matrix Multiplication](#matrixmultiplication)\n",
    "   * [Part 7. Count Vectorizer](#countvectorizer)\n",
    "3. [Next Steps](#next_steps)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='environment_check'></a>\n",
    "### Environment Check\n",
    "\n",
    "Let's ensure that we are running our notebook in the correct environment."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "import os\n",
    "assert os.environ['CONDA_DEFAULT_ENV'] == \"cs124\"\n",
    "\n",
    "import sys\n",
    "assert sys.version_info.major == 3 and sys.version_info.minor == 10"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "If the above cell causes an error, it means that you are using the wrong \n",
    "environment or `Python` version!\n",
    "If this is the case, please follow the troubleshotting steps shared in the \n",
    "[Jupyter Notebook Tutorial](https://github.com/cs124/pa0-python-jupyter-tutorial/blob/main/jupyter_tutorial.ipynb)."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='numpy_exercises'></a>\n",
    "## `NumPy` Exercises"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='basic_numpy'></a>\n",
    "\n",
    "### Part 1. Basic `NumPy`"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Let's import numpy (aliasing the import as np is traditional)\n",
    "import numpy as np"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "The basic building blocks of `NumPy` are arrays, which are represented with the\n",
    "`np.ndarray` type.\n",
    "Arrays represent multi-dimensional matrices (often also referred to as tensors).\n",
    "\n",
    "We can easily create a 1-D array (a vector) by calling `np.array()` and passing\n",
    "in a `Python` list with the data that should go into the array:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# This is an array (of type np.ndarray)\n",
    "a = np.array([1, 2, 3, 4, 5])\n",
    "\n",
    "print(\"a is an: {}\".format(type(a)))\n",
    "a"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Arrays can have different numbers of dimensions, different shapes, and contain\n",
    "elements of different types.\n",
    "Very frequently you'll want to check these properties (especially shape), \n",
    "which you can do like this:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# a is an array containing integers (in this case, 64-bit integers)\n",
    "print(\"Data type of a: {}\".format(a.dtype))\n",
    "\n",
    "# a is a 1-dimensional array of length 5\n",
    "print(\"Shape of a: {}\".format(a.shape))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Note that b contains floating point numbers, not integers\n",
    "b = np.array([5.0, 1.0])\n",
    "\n",
    "print(\"Data type of b: {}\".format(b.dtype))\n",
    "print(\"Shape of b: {}\".format(b.shape))\n",
    "b"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "As we noted above, you can initialize a 1-D array with a python list.\n",
    "However, most of the time we're interested in higher-dimensional arrays like \n",
    "2-D arrays (matrices).\n",
    "You can initialize them using nested `Python` lists:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# A 2-D array\n",
    "c = np.array([[1, 2],\n",
    "              [3, 4],\n",
    "              [5, 6]])\n",
    "\n",
    "print(\"Shape of c: {}\".format(c.shape))\n",
    "c"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Note that the length of the __outer__ list is the first dimension (in this\n",
    "case of length 3), while the lengths of the __inner__ lists are the second\n",
    "dimension (in this case of length 2).\n",
    "\n",
    "It's easy to get your dimensions/shapes mixed up if you get the order confused.\n",
    "Just remember that the number of (horizontal) rows is always the first\n",
    "dimension, and the number of (vertical) columns is the second dimension."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "In addition to the basic `np.array`, `NumPy` provides a bunch of other convenient\n",
    "methods to create different types of arrays/matrices (to save you from\n",
    "typing out the data by hand, or having to use loops or `Python` list\n",
    "comprehensions). \n",
    "Some useful ones include:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# To create an all-zero array with the given shape (3, 4)\n",
    "zeros = np.zeros((3, 4))\n",
    "print(\"3x4 Array of Zeros\")\n",
    "print(zeros)\n",
    "\n",
    "# To create an all-ones array\n",
    "ones = np.ones((2, 2))\n",
    "print(\"\\n2x2 Array of Ones\")\n",
    "print(ones)\n",
    "\n",
    "# To create an array filled with a single value\n",
    "filled = np.full((3, 3), 5)\n",
    "print(\"\\n3x3 Array of 5s\")\n",
    "print(filled)\n",
    "\n",
    "# To create an identity matrix\n",
    "identity = np.identity(3)\n",
    "print(\"\\n3x3 Identity Matrix\")\n",
    "print(identity)\n",
    "\n",
    "# To create an array filled with random values sampled uniformly\n",
    "# from [0.0, 1.0)\n",
    "random = np.random.random((2, 2))\n",
    "print(\"\\n2x2 Array where each value is sampled from U[0,1)\")\n",
    "print(random)\n",
    "\n",
    "# To create an uninitialized array (junk values). This is useful if you know\n",
    "# you're going to manually fill/overwrite the entire array anyways (it saves\n",
    "# the time NumPy would have spent to set every entry to a particular value)\n",
    "# NOTE: Do NOT confuse this with np.zeros(). \"empty\" does not mean all zeroes,\n",
    "# it just means we don't care what is in it. It could be all zeros, it could\n",
    "# be all ones, it could be anything at all.\n",
    "empty = np.empty((2, 2))\n",
    "print(\"\\n2x2 Array of Junk Values\")\n",
    "print(empty)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Finally, once we have an array that we've initialized with data, we can also\n",
    "reshape it without changing its data using np.reshape!\n",
    "\n",
    "Check out this example:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "a = np.array([[1, 2, 3],\n",
    "              [4, 5, 6]])\n",
    "print(\"a.shape = {}\".format(a.shape))\n",
    "print(a)\n",
    "\n",
    "# Can reshape it from 2x3 to 3x2\n",
    "# Pay careful attention to the new order\n",
    "# of values in the reshaped array\n",
    "a_reshaped = a.reshape((3, 2))\n",
    "print(\"\\na_reshaped.shape = {}\".format(a_reshaped.shape))\n",
    "print(a_reshaped)\n",
    "\n",
    "# Can reshape from 2x3 to 6 (1d vector now)\n",
    "a_reshaped = a.reshape((6,))\n",
    "print(\"\\na_reshaped.shape = {}\".format(a_reshaped.shape))\n",
    "print(a_reshaped)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Note what happens if we try to reshape to a shape with a different number\n",
    "# of elements:\n",
    "a_reshaped = a.reshape((7,))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "<a id='indexing_and_slicing'></a>\n",
    "### Part 2. Indexing and Slicing\n",
    "\n",
    "Arrays can be initialized with lists, as we saw, and in many ways they behave\n",
    "a lot like `Python` lists!\n",
    "You can access specific elements by their index (starting with zero, just like \n",
    "in `Python` lists), and modify them:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "a = np.array([1.0, 2.0, 3.0, 4.0, 5.0])\n",
    "\n",
    "b = np.array([[1, 2, 3],\n",
    "              [4, 5, 6],\n",
    "              [7, 8, 9]])"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# You can access elements in an array like a Python list by indexing:\n",
    "print(\"a[3] = {}\".format(a[3]))\n",
    "\n",
    "# You can index into higher-dimensional arrays the same way.\n",
    "\n",
    "# NOTE: If you had nested python lists instead of a NumPy array, you'd\n",
    "# need to do something like b[0][1] instead of b[0, 1], so it's a little\n",
    "# different, but the idea is the same. The b[0][1] syntax will also work\n",
    "# for NumPy arrays.\n",
    "print(\"b[0, 0] = {}\".format(b[0, 0]))\n",
    "print(\"b[1, 2] = {}\".format(b[1, 2]))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# You can also modify elements just like in a Python list\n",
    "print(\"a before:\\n {}\\n\".format(a))\n",
    "a[2] = 9.0\n",
    "print(\"a after:\\n {}\\n\".format(a))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "`NumPy` also supports more complex forms of indexing (like slicing), which\n",
    "also behaves similarly to `Python` lists:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Use a list and numpy array as examples to show slicing works the same\n",
    "example_list = [1, 2, 3, 4, 5, 6]\n",
    "a = np.array([1, 2, 3, 4, 5, 6])\n",
    "\n",
    "# This gives a slice containing every element starting from position/index 1\n",
    "# (inclusive) up to but EXCLUDING the element at position/index 3\n",
    "print(\"Python List: example_list[1:3] = {}\".format(example_list[1:3]))\n",
    "\n",
    "# NumPy works exactly the same way\n",
    "print(\"Numpy Array: a[1:3] = {}\".format(a[1:3]))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "We can modify slices just like how we modified elements using indexing:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "print(\"a before:\\n {}\\n\".format(a))\n",
    "a[1:3] = [8, 9]\n",
    "print(\"a after:\\n {}\\n\".format(a))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "We can also slice multi-dimensional arrays.\n",
    "Try to figure out what each of the expression below will give before running \n",
    "them, to check your intuition:\n",
    "\n",
    "Pay careful attention to the shapes of the outputs, the number of dimensions, and size of the dimensions"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "b = np.array([[1, 2, 3],\n",
    "              [4, 5, 6],\n",
    "              [7, 8, 9]])\n",
    "\n",
    "print(\"b[:, :] =>\\n {}\\nShape : {}\\n\".format(b[:, :], b[:, :].shape))\n",
    "\n",
    "# Note that 1-D arrays are always treated as row vectors (horizontal)\n",
    "# When we index on any dimension, that dimension is dropped.\n",
    "print(\"b[1, :2] =>\\n {}\\nShape : {}\\n\".format(b[1, :2], b[1, :2].shape))\n",
    "\n",
    "# When we slice on a dimension instead of indexing (1:2 instead of 1), \n",
    "# we DON'T drop a dimension\n",
    "print(\"b[1:2, :2] =>\\n {}\\nShape : {}\\n\".format(b[1:2, :2], b[1:2, :2].shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "It's important to think carefully about the shapes/dimensions of the slices\n",
    "that you extract.\n",
    "Depending on how you slice, your result could have fewer dimensions than the \n",
    "original array, or the same number.\n",
    "\n",
    "Think about what shapes you would expect these slices to be, then run the cell \n",
    "below to double-check:\n",
    "\n",
    "Do the results make sense to you? \n",
    "If not, can you figure out why the shapes came out as they did?"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "<a id='array_math_and_functions'></a>\n",
    "### Part 3. Array Math and Functions\n",
    "\n",
    "We covered creating arrays and reading/writing elements in them, but the main\n",
    "reason we use `NumPy` in the first place is to do math with arrays (linear\n",
    "algebra).\n",
    "For the most part, array/vector math in `NumPy` is extremely straight-forward \n",
    "and intuitive.\n",
    "You can add, subtract, multiply, etc. `NumPy` arrays just like they were \n",
    "`Python` numbers and NumPy will take care of everything for you!\n",
    "\n",
    "For example:"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Note that for all the below element-wise operations, the two\n",
    "arrays being added/subtracted/multiplied etc. must have the same shape!\n",
    "(We will talk about an exception to this rule in the next exercise.)\n",
    "\n",
    "Also note that for many matrix operations, there's two ways of\n",
    "writing it in `NumPy`. \n",
    "Either you can call a dedicated function, or you can use the standard \n",
    "\"+, -, *, /\" operators on NumPy arrays as if they were just numbers"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "a = np.array([[1,2],\n",
    "              [3,4]])\n",
    "\n",
    "b = np.array([[3,3],\n",
    "              [4,4]])\n",
    "\n",
    "print(\"a =>\\n {}\\nShape : {}\\n\".format(a, a.shape))\n",
    "print(\"b =>\\n {}\\nShape : {}\\n\".format(b, b.shape))\n",
    "\n",
    "# To take an element-wise sum\n",
    "print(\"a + b =>\\n {}\\n\".format(a + b))\n",
    "print(\"np.add(a, b) =>\\n {}\\n\".format(np.add(a, b)))\n",
    "\n",
    "# To take an element-wise difference\n",
    "print(\"b - a =>\\n {}\\n\".format(b - a))\n",
    "print(\"np.subtract(b, a) =>\\n {}\\n\".format(np.subtract(b, a)))\n",
    "\n",
    "# To take an element-wise product \n",
    "# IMPORTANT NOTE : THIS IS NOT MATRIX MULTIPLICATION, THIS\n",
    "# IS ELEMENT-WISE MULTIPLICATION. We will see how to do matrix\n",
    "# multiplication in a later part\n",
    "print(\"a * b =>\\n {}\\n\".format(a * b))\n",
    "print(\"np.multiply(a, b) =>\\n {}\\n\".format(np.multiply(a, b)))\n",
    "\n",
    "# To take an element-wise quotient\n",
    "print(\"a / b =>\\n {}\\n\".format(a / b))\n",
    "print(\"np.divide(a, b) =>\\n {}\\n\".format(np.divide(a, b)))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Note what happens when the shapes don't match\n",
    "\n",
    "a = np.array([1, 2, 3])\n",
    "b = np.array([1, 2])\n",
    "\n",
    "a + b"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "`NumPy` also provides a bunch of super-useful functions to compute mathematical\n",
    "functions of NumPy arrays, or do other common operations.\n",
    "Some commonly used ones that you might find useful include:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "a = np.array([1, 2, 3])\n",
    "\n",
    "b = np.array([[1, 2, 3],\n",
    "              [4, 5, 6]])\n",
    "\n",
    "print(\"a =>\\n {}\\nShape : {}\\n\".format(a, a.shape))\n",
    "print(\"b =>\\n {}\\nShape : {}\\n\".format(b, b.shape))\n",
    "\n",
    "# Sum all the elements in an array\n",
    "print(\"np.sum(b) = {}\".format(np.sum(b)))\n",
    "\n",
    "# Sum elements along an axis/dimension\n",
    "print(\"np.sum(b, axis = 1) = {}\".format(np.sum(b, axis = 1)))\n",
    "\n",
    "# Get average of all elements in array\n",
    "print(\"np.mean(b) = {}\".format(np.mean(b)))\n",
    "\n",
    "# Mean elements along an axis/dimension\n",
    "print(\"np.mean(b, axis = 1) = {}\".format(np.mean(b, axis = 1)))\n",
    "\n",
    "# Get maximum element in the WHOLE array\n",
    "print(\"np.max(b) = {}\".format(np.max(b)))\n",
    "\n",
    "# Get maximum element across a single axis/dimension for each\n",
    "print(\"np.max(b, axis=1) = {}\".format(np.max(b, axis=1)))\n",
    "\n",
    "# Take (natural) log element-wise\n",
    "print(\"np.log(a) = {}\".format(np.log(a)))\n",
    "\n",
    "# Take exponential (e^x) element-wise\n",
    "print(\"np.exp(a) = {}\".format(np.exp(a)))\n",
    "\n",
    "# Take square root element-wise\n",
    "print(\"np.sqrt(a) = {}\".format(np.sqrt(a)))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "__NOTE:__ In general, for most `NumPy` functions that operate on an entire \n",
    "array, you can also specify a dimension or dimensions to apply it along\n",
    "\n",
    "__NOTE:__ There are also other packages that build off of `NumPy` or\n",
    "use `NumPy` arrays.\n",
    "We will very briefly encounter a few examples later in the class like `SciPy`\n",
    "and `PyTorch`. \n",
    "For the most part, these packages tend to work very similarly to the built-in \n",
    "`NumPy` functions above."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='vectorization'></a>\n",
    "### Part 4. Vectorization\n",
    "\n",
    "Behind the scenes, `NumPy` uses the concept of vectorization to help speed-up operations \n",
    "that would regularly take much much longer using naive python for loops. For this reason,\n",
    "whenever possible, we prefer to use numpy operations as they are generally faster than\n",
    "naive python functions, computations and loops. Run the next three cells to see this."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "arrays = np.random.random((1000, 1000))\n",
    "other_array = np.random.random((1000))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "%%time\n",
    "\n",
    "#Compute the max inner product\n",
    "max_i = 0\n",
    "max_val = 0\n",
    "for  i in range(1000):\n",
    "    idp = 0\n",
    "    for j in range(1000):\n",
    "        idp += arrays[i][j] * other_array[j]\n",
    "    if idp > max_val:\n",
    "        max_val = idp\n",
    "        max_i = i\n",
    "print(max_i, max_val)\n",
    "        \n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "%%time\n",
    "\n",
    "arrays = np.array(arrays)\n",
    "other_array = np.array(other_array)\n",
    "idps = np.dot(arrays, other_array)\n",
    "print(np.argmax(idps), np.max(idps))\n"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='broadcasting'></a>\n",
    "### Part 5. Broadcasting\n",
    "\n",
    "Another topic that may be useful to you is broadcasting. \n",
    "It's one of the most useful and powerful features of `NumPy`, because it lets \n",
    "you write matrix/array operations in a natural way without having to be too \n",
    "specific about what you want `NumPy` to do.\n",
    "\n",
    "In most cases, `NumPy` can use broadcasting to infer what you wanted to do\n",
    "and do it. \n",
    "Let's look at an example:"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Now recall that earlier, when we tried to add, multiply, subtract, etc. two\n",
    "arrays, they had to have the same shape! \n",
    "Let's see what happens when we do some of these things with a and b:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Create a multi-dimensional (1 x 2) matrix of ones\n",
    "a = np.ones((1,2))\n",
    "\n",
    "# Create a single-element 1-D array\n",
    "b = np.array([4])\n",
    "\n",
    "print(\"a =>\\n {}\\nShape : {}\\n\".format(a, a.shape))\n",
    "print(\"b =>\\n {}\\nShape : {}\\n\".format(b, b.shape))\n",
    "\n",
    "a_plus_b = a + b\n",
    "print(\"a + b =>\\n {}\\nShape : {}\\n\".format(a_plus_b, a_plus_b.shape))\n",
    "\n",
    "a_minus_b = a - b\n",
    "print(\"a - b =>\\n {}\\nShape : {}\\n\".format(a_minus_b, a_minus_b.shape))\n",
    "\n",
    "a_times_b = a * b\n",
    "print(\"a * b =>\\n {}\\nShape : {}\\n\".format(a_times_b, a_times_b.shape))\n",
    "\n",
    "a_divided_b = a / b\n",
    "print(\"a / b =>\\n {}\\nShape : {}\\n\".format(a_divided_b, a_divided_b.shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Wait, what happened?? \n",
    "We just said that the shapes had to match, but `a` and `b` definitely don't \n",
    "have the same shape...\n",
    "\n",
    "This is where broadcasting comes into play. Although `a` and `b` don't have the\n",
    "same shape, they have __compatible__ shapes, so `NumPy` was able to guess\n",
    "what we actually wanted to do and __broadcast__ behind the scenes to make the\n",
    "operation work.\n",
    "\n",
    "What `NumPy` did behind the scenes in this case is see that `b` and `a` don't \n",
    "have the same shape, but also realize that maybe we meant to \"re-use\" the value \n",
    "in `b` for every value in `a`. In other words, it \"broadcast\" `b` from its\n",
    "original shape of (1,) to `a`'s shape of (2,) by duplicating the element\n",
    "in `b`.\n",
    "\n",
    "Once it did that, it could simply do the element-wise operation as usual!\n",
    "\n",
    "This makes our life a lot easier, as we didn't have to explicitly write out\n",
    "that duplication and reshaping ourselves.\n",
    "Of course, that assumes that this behavior is actually what we intended.\n",
    "\n",
    "In this case, it seems to make good sense, as if we tell it to subtract the\n",
    "array `[4]` from `[1, 1]`, it seems reasonable that what we are really asking\n",
    "for is for it to subtract 4 from each 1 in the second array."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Most of the time, `NumPy` assumes that whenever we do an operation on two arrays\n",
    "with different shapes, we would like things to be broadcast if possible to\n",
    "make the operation work.\n",
    "\n",
    "It will then try to expand the smaller array to match the size of the bigger\n",
    "one by copying/repeating the data in the smaller array."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Let's try a slightly more complicated example involving 2-D arrays:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {
    "pycharm": {
     "name": "#%%\n"
    }
   },
   "outputs": [],
   "source": [
    "# Create a 2-D array of ones of shape (4, 3)\n",
    "a = np.ones((4,3))\n",
    "\n",
    "# Create an array of shape (4,)\n",
    "# np.arange(n) creates a vector out of increasing\n",
    "# integers from 0 to n (exclusive)\n",
    "b = np.arange(4)\n",
    "\n",
    "# Reshape to (4, 1)\n",
    "b = b.reshape((4, 1))\n",
    "\n",
    "# Multiply them together\n",
    "ba = b * a\n",
    "\n",
    "print(\"a =>\\n {}\\nShape : {}\\n\".format(a, a.shape))\n",
    "print(\"b =>\\n {}\\nShape : {}\\n\".format(b, b.shape))\n",
    "print(\"a * b =>\\n {}\\nShape : {}\\n\".format(ba, ba.shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "Did it do what you expected?\n",
    "Do you see how broadcasting was applied to create the result?"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Broadcasting also works with scalars. In general, most operations you can do with two scalars, you can do with an array and a scalar and the operation will be broadcast to every element in the array."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "This is true for the obvious operations:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "a = np.arange(5)\n",
    "print(\"a = {}\".format(a))\n",
    "\n",
    "a_plus_5 = a + 5\n",
    "print(\"a + 5 = {}\".format(a_plus_5))\n",
    "\n",
    "a_minus_1 = a - 1\n",
    "print(\"a - 1 = {}\".format(a_minus_1))\n",
    "\n",
    "a_times_10 = a * 10\n",
    "print(\"a * 10 = {}\".format(a_times_10))\n",
    "\n",
    "a_divided_2 = a / 2\n",
    "print(\"a / 2 = {}\".format(a_divided_2))\n",
    "\n",
    "a_squared = a ** 2\n",
    "print(\"a ** 2 = {}\".format(a_squared))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "It's also true in a lot of ways you might not think of right away but that can be very useful! For example:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "a = np.random.random((5, 5))\n",
    "a\n",
    "\n",
    "# Creates a boolean mask where each element in the array\n",
    "# is either true or false based on if the value is > 0.5 or not\n",
    "a_mask = (a > 0.5)\n",
    "\n",
    "print(\"a =>\\n {}\\n Shape : {}\\n\".format(a, a.shape))\n",
    "print(\"a > 0.5 =>\\n {}\\n Shape : {}\\n\".format(a_mask, a_mask.shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Can you think of a way this might be useful?"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "The details and different cases for broadcasting can be tricky, even\n",
    "for experienced `NumPy` users!\n",
    "So you're certainly not expected to be use or need it heavily in this class. \n",
    "However, knowing the basics of how it works may make your life a little easier \n",
    "and your code a little simpler on some of the homeworks.\n",
    "\n",
    "If you ever find yourself in a confusing situation involving broadcasting, we\n",
    "definitely recommend checking out the `NumPy` documentation on\n",
    "[Broadcasting](https://numpy.org/doc/stable/user/basics.broadcasting.html).\n"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='matrixoperations'></a>\n",
    "### Part 6. Matrix Multiplication\n",
    "\n",
    "A final topic that will definitely be useful for the homeworks and is very important\n",
    "in `NumPy` is that of matrix operations. While we covered element-wise multiplication, \n",
    "how can we compute matrix multiplication or dot products of vectors?\n",
    "\n",
    "These next functions, especially np.matmul and np.dot will definitely come up\n",
    "in your homeworks and is valuable to understand moving forward."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Let's discuss matrix multiplication between two numpy arrays `a` and `b`.\n",
    "\n",
    "We can then use either `np.matmul`, `np.dot`, or the @ operator to matrix multiply `a` with `b`\n",
    "\n",
    "Note that `np.matmul` and `np.dot` have subtle differences that you can read more about\n",
    "in their documentation pages. As a general tip, if you ever encounter a `NumPy` function\n",
    "you are not fully understanding, you can search up its documentation. `NumPy` keeps\n",
    "great documentation and has examples to help as well.\n",
    "\n",
    "`np.matmul` documentation: https://numpy.org/doc/stable/reference/generated/numpy.matmul.html\n",
    "\n",
    "`np.dot` documentation: https://numpy.org/doc/stable/reference/generated/numpy.dot.html\n",
    "\n",
    "The primary difference for the purposes of the class is that you can pass a single scalar value \n",
    "into `np.dot` but cannot into `np.matmul`. For example, `np.dot(a,5)` would be valid using\n",
    "our same `a` but `np.matmul(a,5)` would not be valid since `5` is just a scalar value\n",
    "\n",
    "The other differences exist when passing in arrays with more than 2-dimensions but this isn't relevant for this class"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Let's make a have shape (2,5) and b have shape (5,4) for simplicity\n",
    "a = np.ones((2,5))\n",
    "b = np.full((5,4), 2)\n",
    "\n",
    "matmul_result = np.matmul(a,b)\n",
    "dot_result = np.dot(a,b)\n",
    "operator_result = a @ b\n",
    "\n",
    "print(\"a =>\\n {}\\nShape : {}\\n\".format(a, a.shape))\n",
    "print(\"b =>\\n {}\\nShape : {}\\n\".format(b, b.shape))\n",
    "print(\"np.matmul(a, b) =>\\n {}\\nShape : {}\\n\".format(matmul_result, matmul_result.shape))\n",
    "print(\"np.dot(a, b) =>\\n {}\\nShape : {}\\n\".format(dot_result, dot_result.shape))\n",
    "print(\"a @ b =>\\n {}\\nShape : {}\\n\".format(operator_result, operator_result.shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### EXAMPLE 1\n",
    "\n",
    "Assume we have the numpy matrices $x$ and $y$ with shapes $(2,3)$ and $(4,3)$ respectively\n",
    "\n",
    "Without running the next few cells, write down what you think the shapes of the following variables will be. If you think an error will be thrown, write that down instead.\n",
    "\n",
    "* $A = xy$ \n",
    "* $B = xy^T$\n",
    "* $C = x^Ty$\n",
    "* $D = yx^T$"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 10,
   "metadata": {},
   "outputs": [],
   "source": [
    "### YOUR ANSWERS HERE ###\n",
    "answers = \"\"\"\n",
    "A's shape: \n",
    "B's shape: \n",
    "C's shape:\n",
    "D's shape: \n",
    "\"\"\"\n",
    "### YOUR ANSWERS HERE ###"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Declare our matrices x and y\n",
    "x = np.ones((2,3))\n",
    "y = np.ones((4,3))\n",
    "\n",
    "print(\"x's shape:\", x.shape)\n",
    "print(\"y's shape:\", y.shape)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Over the next cells, define our outputs using np.matmul\n",
    "A = np.matmul(x,y)\n",
    "print(\"A's shape:\", A.shape)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "B = np.matmul(x,y.T)\n",
    "print(\"B's shape:\", B.shape)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "C = np.matmul(x.T,y)\n",
    "print(\"C's shape:\", C.shape)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "D = np.matmul(y,x.T)\n",
    "print(\"D's shape:\", D.shape)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### SOLUTION\n",
    "\n",
    "As a general rule, lets assume that `a` has shape (i,j) and `b` has shape (m,n). In order to \n",
    "have a valid matrix multiplication, we must ensure that j == m, or the second \n",
    "dimension of our first array is equal to the first dimension of the second array. \n",
    "The outputted matrix as a result of the multiplication will then have dimensions (i,n),\n",
    "where the first dimension is equal to the first dimension of the first array\n",
    "and the second dimension is equal to the second dimension of the second array.\n",
    "\n",
    "* $A$ -> Error -> $x$ has shape $(2,3)$ and $y$ has shape $(4,3)$. Since the inner dimensions do not match, this is not a valid matrix multiplication\n",
    "* $B$ -> $(2,4)$ -> $x$ has shape $(2,3)$ and $y^T$ has shape $(3,4)$. Since the inner dimensions do match, this is a valid matrix multiplication. Remember that the outputted shape will be equal to the first dimension of the first matrix (which is $2$) and the second dimension of the second matrix (which is $4$). Thus, the output shape will be $(2,4)$\n",
    "* $C$ -> Error -> $x^T$ has shape (3,2) and $y$ has shape (4,3). Since the inner dimensions do not match, this is not a valid matrix multiplication\n",
    "* $D$ -> $(4,2)$ -> $y$ has shape $(4,3)$ and $x^T$ has shape $(3,2)$. ince the inner dimensions do match, this is a valid matrix multiplication. Remember that the outputted shape will be equal to the first dimension of the first matrix (which is $4$) and the second dimension of the second matrix (which is $2$). Thus, the output shape will be $(4,2)$"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "In an effort to always be aware of the shapes of your numpy arrays, it is recommended \n",
    "that you make sure that both inputs into np.matmul are 2-dimensional. If one of your\n",
    "matrices is a single dimension, you can add a dimension using `np.expand_dims` and can\n",
    "also remove a dimension if needed by using the `.squeeze()` command."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Using np.expand_dims to add a new dimension to the front\n",
    "# and back of a 1-d vector.\n",
    "a = np.array([1,2,3])\n",
    "\n",
    "# Add the axis = n parameter to specify that the nth dimension should be \n",
    "# added as a dimension with size 1\n",
    "a_expanded_0 = np.expand_dims(a, axis = 0)\n",
    "a_expanded_1 = np.expand_dims(a, axis = 1)\n",
    "print(\"np.expand_dims(a, axis = 0) =>\\n {}\\nShape : {}\\n\".format(a_expanded_0, a_expanded_0.shape))\n",
    "print(\"np.expand_dims(a, axis = 1) =>\\n {}\\nShape : {}\\n\".format(a_expanded_1, a_expanded_1.shape))\n",
    "\n",
    "# Create b with shape (1,3,1) to display .squeeze()\n",
    "# Passing in dimension x into .squeeze(x) results in \n",
    "# a dimension with size 1 being removed\n",
    "b = np.zeros((1, 3, 1))\n",
    "\n",
    "print(\"b.shape : {}\".format(b.shape))\n",
    "print(\"b.squeeze(0).shape : {}\".format(b.squeeze(0).shape))\n",
    "print(\"b.squeeze(-1).shape : {}\".format(b.squeeze(-1).shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### EXAMPLE 2\n",
    "\n",
    "Assume we have the following matrices/vectors as `numpy` arrays:\n",
    "\n",
    "* $X$ with shape (2,3) -> matrix\n",
    "* $A$ with shape (2,) -> vector\n",
    "* $B$ with shape (2,) -> vector\n",
    "\n",
    "Compute the following values: \n",
    "\n",
    "* $Y = \\sum_{i=1}^{2} \\Bigl(A^{(i)} - B^{(i)}\\Bigr)$ -> Should be a float\n",
    "* $Z = X^T \\left(A - B\\right)$ -> Should be a vector\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Initialize X, A, and B. Their values themselves are not important\n",
    "X = np.reshape(np.arange(6) + 1, (2, 3)) # Reshape [1, 2, 3, 4, 5, 6] -> [[1, 2, 3], [4, 5, 6]]\n",
    "A = np.ones(2) # A = [1, 1]\n",
    "B = np.zeros(2) - 1 # B = [-1, -1]\n",
    "C = np.zeros(2) # C = [0, 0]\n",
    "\n",
    "print(\"X shape:\", X.shape)\n",
    "print(\"A shape:\", A.shape)\n",
    "print(\"B shape:\", B.shape)\n",
    "print(\"C shape:\", C.shape)\n",
    "\n",
    "### YOUR WORK HERE ###\n",
    "\n",
    "### YOUR WORK HERE ###"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### SOLUTION IN THE NEXT CELL"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# SOLUTION\n",
    "\n",
    "# Since A/B are the same shape, we can directly subtract them\n",
    "# We are looking for the sum over each element so we can use np.sum\n",
    "Y = np.sum(A - B) \n",
    "\n",
    "# First, let's transpose X. It is VERY IMPORTANT we do NOT use\n",
    "# np.reshape and instead use the .T to transpose\n",
    "X_T = X.T\n",
    "\n",
    "# Expand the dimensions of A and B so we can safely do matrix multiplication\n",
    "A_expanded = np.expand_dims(A, 1) # Shape is now (2,1)\n",
    "B_expanded = np.expand_dims(B, 1) # Shape is now (2,1)\n",
    "\n",
    "A_sub_B = A_expanded - B_expanded # A - B but the shape is (2,1)\n",
    "\n",
    "# Finally do the matrix multiplication\n",
    "Z = X_T @ A_sub_B # Shape is now (3,1)\n",
    "\n",
    "# Remember that our original computation was a matrix multiplied by a vector\n",
    "# This should result in a vector, or a 1d numpy array\n",
    "Z_final = Z.squeeze(1)\n",
    "\n",
    "print(\"Expected Y is:\", 4.0)\n",
    "print(\"Computed Y is:\", Y)\n",
    "\n",
    "expected_Z = np.array([10, 14, 18]).astype(np.float64)\n",
    "print(\"Expected Z is:\", expected_Z)\n",
    "print(\"Computed Z is:\", Z_final)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "However, note that if you don't pass in a 2-dimensional array and instead pass a 1-dimensional \n",
    "array into any of the matrix multiplication methods, the array will first be expanded to two\n",
    "dimensions and then the added dimension will be removed. For example, if we had `a` with shape\n",
    "(i,j) and `b` with shape (j,) (1-dimensional), then doing the operation `np.matmul(a,b)` would\n",
    "internally first expand `b` to have shape (j,1), would then matrix multiply to get a array of\n",
    "shape (i,1) and then would remove the added dimension, so the final output would be of shape (i,)\n",
    "or a 1-dimensional array\n",
    "\n",
    "Some examples of this behavior can be seen below"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Create the following variables with representative shapes\n",
    "# Each array contains all ones\n",
    "# a : (2,3)\n",
    "# b : (3,2)\n",
    "# c : (2,)\n",
    "# d : (3,)\n",
    "# e : (3,)\n",
    "\n",
    "a = np.ones((2, 3))\n",
    "b = np.ones((3, 2))\n",
    "c = np.ones(2,)\n",
    "d = np.ones(3,)\n",
    "e = np.ones(3,)\n",
    "\n",
    "ab = a @ b\n",
    "ad = a @ d\n",
    "ca = c @ a\n",
    "de = d @ e\n",
    "\n",
    "print(\"a @ b =>\\n {}\\nShape : {}\\n\".format(ab, ab.shape))\n",
    "print(\"a @ d =>\\n {}\\nShape : {}\\n\".format(ad, ad.shape))\n",
    "print(\"c @ a =>\\n {}\\nShape : {}\\n\".format(ca, ca.shape))\n",
    "print(\"d @ e =>\\n {}\\nShape : {}\\n\".format(de, de.shape))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "<a id='countvectorizer'></a>\n",
    "### Part 7. Count Vectorizer\n",
    "\n",
    "The final topic in this tutorial is going to be Scikit-Learn's CountVectorizer. This will be used in pa2 and can often be confusing. We will walk through a small example to help you understand this function."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "`CountVectorizer` in scikit-learn is a useful took for getting a list of how often each word (or ngram) appears in a document.  It converts a collection of text documents into a matrix of token (word or n-gram) counts. It's very helpful in any situation that involves counting words.  The fit() method is used to learn the vocabulary dictionary. And then the transform() method is used to actually count the words, producing a document-term matrix.  Often we use it in a supervised machine learning setup, where we use the training set to both learn the vocabulary and get counts,  and then for the test set we just get counts from the known vocabulary."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Scikit Learn's `CountVectorizer` should be initialized as a class. While there are a number of functions and use cases for this class, the primary functions we will be focusing on are the  `fit_transform` function and the `transform` function.\n",
    "\n",
    "`fit_transform` : This function is an optimized version of calling the `fit` function followed by the `transform` function. This function should be given a list of strings where each string represents a document from your full dataset. The function will then return the document-term matrix from this dataset of documents. Importantly, the vocabulary used to build this document-term matrix is generated via a combination of all the words over all the provided documents. This vocabulary is then saved by the class for later use.\n",
    "\n",
    "`transform` : This function should be given a list of strings where each string represents a document from your full dataset. The function will then return the document-term matrix from this dataset of documents. Importantly, the vocabulary used to build this document-term matrix is from the already saved vocabulary. A new vocabulary will NOT be generated. In other words, this function should only be called after `fit_transform` has been called. \n",
    "\n",
    "A great example of using these functions together would be to call `fit_transform` on the training dataset followed by `transform` on the validation/testing dataset. The vocabulary from the training dataset will then be used during validation/testing.\n",
    "\n",
    "Note that a document-term matrix is $A\\text{x}B$ matrix where there are $A$ total documents and $B$ total words in the vocabulary. Entry $i$,$j$ represents how many times word $j$ appears in document $i$. The words are ordered in alphabetical order by `CountVectorizer`.\n",
    "\n",
    "#### The output of `fit_transform` and `transform` functions is a compressed matrix. To get the full document-term matrix, you must add `.toarray()` to the end of output from the function\n",
    "\n",
    "[CountVectorizer Full Documentation](https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.CountVectorizer.html#countvectorizer)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### EXAMPLE 3\n",
    "\n",
    "You are given two lists of strings, where each string represents a document. The two lists are `train_dataset` and `test_dataset`. You have two goals:\n",
    "\n",
    "* Generate the document-term matrix for `train_dataset` using the vocabulary from `train_dataset`\n",
    "* Generate the document-term matrix for `test_dataset` using the vocabulary from `train_dataset`"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Import the proper packages\n",
    "import sklearn\n",
    "from sklearn.feature_extraction.text import CountVectorizer\n",
    "\n",
    "# Declare our count vectorizer class. \n",
    "# The variable \"vectorizer\" is now a class with many callable functions\n",
    "vectorizer = CountVectorizer()\n",
    "\n",
    "# Declare a toy example of a train dataset and test dataset\n",
    "# Note the format of our dataset: List[str] where each str is a \n",
    "# string of the full document, which could have one or many sentences.\n",
    "train_dataset = [\n",
    "    \"document one. it has the following text.\",\n",
    "    \"document two. it has two more awesome cool words.\",\n",
    "]\n",
    "\n",
    "test_dataset = [\n",
    "    \"doc 6. awesome cool words should be used one or two more times in documents.\",\n",
    "    \"doc 7. we ran out of awesome cool ideas for this not awesome one\",\n",
    "]\n",
    "\n",
    "### YOUR WORK HERE ###\n",
    "\n",
    "\n",
    "\n",
    "### YOUR WORK HERE ###"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### SOLUTION IN NEXT FEW CELLS"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Let's call fit_transform on our train_dataset and view the output\n",
    "train_output = vectorizer.fit_transform(train_dataset)\n",
    "print(\"Pure output from fit_transform is:\")\n",
    "print(train_output) "
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Notice how the above output is not in the matrix format we want. To get it into the matrix format we want, we need to add `.toarray()` to the output. "
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "train_dt_matrix = train_output.toarray()\n",
    "print(\"Train dataset document term matrix:\")\n",
    "print(train_dt_matrix)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Now let's do the same thing for the test dataset. However, remember that we want to use the vocabulary from the train dataset. Thus, we do NOT want to use `fit_transform`. Instead, we will use `transform` which will use the saved vocabulary from when we called `fit_transform` on the train dataset"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "test_dt_matrix = vectorizer.transform(test_dataset).toarray()\n",
    "print(\"Test dataset document term matrix:\")\n",
    "print(test_dt_matrix)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "Notice how the words corresponding to these frequencies are from the train dataset: \n",
    "* awesome, cool, document, following, has, it, more, one, the, text, two, words"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {
    "pycharm": {
     "name": "#%% md\n"
    }
   },
   "source": [
    "<a id='next_steps'></a>\n",
    "## Next Steps\n",
    "\n",
    "That's it! \n",
    "That's all the basic `NumPy` knowledge you need for the homeworks in this course\n",
    "(possibly you may not even need all of the tools we have introduced here).\n",
    "\n",
    "If you are interested in learning more about `NumPy`, we recommend:\n",
    "* [The NumPy User Guide](https://numpy.org/doc/stable/user/index.html)\n",
    "* [CS 231n NumPy Tutorial](https://cs231n.github.io/python-numpy-tutorial/#numpy),\n",
    "  which we based much of the content here on, with more advanced topics\n",
    "\n",
    "If you found any issues/errors in the notebook, please let us know!\n",
    "\n",
    "And if you have any further questions about using `NumPy` or confusion about any\n",
    "of the examples in the notebook, stop by office hours or post your question on \n",
    "our `Q&A` platform.\n",
    "Teaching staff will be happy to answer your questions!"
   ]
  }
 ],
 "metadata": {
  "colab": {
   "collapsed_sections": [],
   "name": "CS124 Python Exercises.ipynb",
   "provenance": []
  },
  "kernelspec": {
   "display_name": "myenv",
   "language": "python",
   "name": "myenv"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.11.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}
