{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "ab743d0e",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h1 class=\"cal cal-h1\">Lecture 02: KNN, ML Vocabulary, and K-Means &ndash; CS 189, Fall 2026</h1>\n",
    "\n",
    "\n",
    "In this notebook we build our way from raw data to our first machine learning model, and then to our\n",
    "first unsupervised model. We start with the tools we need to look at data (`pandas`, `numpy`, and\n",
    "plotting), then introduce the simplest model we can think of, **k-nearest neighbors**. Along the way\n",
    "KNN will force us to invent the ideas of **generalization**, the **train/test split**, and\n",
    "**hyperparameters**. We close with **k-means**, which solves a related problem without any labels\n",
    "at all.\n",
    "\n",
    "The main body of this notebook is what we walk through in lecture. The **Appendix** at the end goes\n",
    "much deeper on `pandas`, `numpy`, and visualization syntax; work through it after lecture.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a734c5e5",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.538925Z",
     "iopub.status.busy": "2026-08-30T12:32:43.538656Z",
     "iopub.status.idle": "2026-08-30T12:32:43.618458Z",
     "shell.execute_reply": "2026-08-30T12:32:43.617483Z"
    }
   },
   "outputs": [],
   "source": [
    "# Download data & Install Dependencies\n",
    "import os\n",
    "import requests\n",
    "\n",
    "os.makedirs(\"data\", exist_ok=True)\n",
    "data_files = {\n",
    "    \"data/penguins.csv\":\n",
    "        \"https://raw.githubusercontent.com/mwaskom/seaborn-data/master/penguins.csv\",\n",
    "}\n",
    "for path, url in data_files.items():\n",
    "    if not os.path.exists(path):\n",
    "        r = requests.get(url)\n",
    "        r.raise_for_status()\n",
    "        with open(path, \"wb\") as f:\n",
    "            f.write(r.content)\n",
    "        print(f\"Downloaded {path}\")\n",
    "    else:\n",
    "        print(f\"Found {path}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "19cc7143",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.620605Z",
     "iopub.status.busy": "2026-08-30T12:32:43.619957Z",
     "iopub.status.idle": "2026-08-30T12:32:43.931646Z",
     "shell.execute_reply": "2026-08-30T12:32:43.930486Z"
    }
   },
   "outputs": [],
   "source": [
    "import numpy as np\n",
    "import pandas as pd\n",
    "import plotly.express as px\n",
    "\n",
    "px.defaults.width = 800\n",
    "pd.set_option(\"plotting.backend\", \"plotly\")\n",
    "\n",
    "os.makedirs(\"images\", exist_ok=True)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "f368ccab",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 1: Looking at the Data</h2>\n",
    "\n",
    "Before any modeling, look at the data. This section is a fast tour of the tools; the Appendix has\n",
    "the full treatment of every function used here."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b87f19be",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">The Palmer Penguins Dataset</h3>\n",
    "\n",
    "\n",
    "Measurements of 344 penguins from three islands in the Palmer Archipelago, Antarctica. Each row is\n",
    "one penguin. We will use this single dataset for everything today.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "01b03018",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.933989Z",
     "iopub.status.busy": "2026-08-30T12:32:43.933328Z",
     "iopub.status.idle": "2026-08-30T12:32:43.947914Z",
     "shell.execute_reply": "2026-08-30T12:32:43.947060Z"
    }
   },
   "outputs": [],
   "source": [
    "penguins = pd.read_csv(\"data/penguins.csv\")\n",
    "penguins.head()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "7bb2e283",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.950000Z",
     "iopub.status.busy": "2026-08-30T12:32:43.949427Z",
     "iopub.status.idle": "2026-08-30T12:32:43.957931Z",
     "shell.execute_reply": "2026-08-30T12:32:43.957164Z"
    }
   },
   "outputs": [],
   "source": [
    "penguins.info()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "9be3c457",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.959927Z",
     "iopub.status.busy": "2026-08-30T12:32:43.959334Z",
     "iopub.status.idle": "2026-08-30T12:32:43.971720Z",
     "shell.execute_reply": "2026-08-30T12:32:43.970958Z"
    }
   },
   "outputs": [],
   "source": [
    "# describe() summarizes the numeric columns\n",
    "penguins.describe()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e7d1af0c",
   "metadata": {},
   "source": [
    "How much data do we have, and what are we looking at?\n",
    "\n",
    "- `shape` gives (rows, columns)\n",
    "- `value_counts()` counts the occurrences of each value in a column\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b06d86d7",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.973288Z",
     "iopub.status.busy": "2026-08-30T12:32:43.973094Z",
     "iopub.status.idle": "2026-08-30T12:32:43.979946Z",
     "shell.execute_reply": "2026-08-30T12:32:43.979041Z"
    }
   },
   "outputs": [],
   "source": [
    "print(\"shape:\", penguins.shape)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b7d25677",
   "metadata": {},
   "outputs": [],
   "source": [
    "penguins[\"species\"].value_counts()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "50712aae",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Selecting Data: `loc` and `iloc`</h3>\n",
    "\n",
    "\n",
    "Two ways to pull a subset out of a `DataFrame`. The distinction matters constantly, so it is worth\n",
    "getting straight now.\n",
    "\n",
    "- `iloc[]` selects by **integer position**, like indexing a numpy array. The end of a slice is *excluded*.\n",
    "- `loc[]` selects by **label**, meaning the index value and the column name. The end of a slice is *included*.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "4cd163f5",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.982281Z",
     "iopub.status.busy": "2026-08-30T12:32:43.981620Z",
     "iopub.status.idle": "2026-08-30T12:32:43.989063Z",
     "shell.execute_reply": "2026-08-30T12:32:43.988250Z"
    }
   },
   "outputs": [],
   "source": [
    "# iloc: by position. Rows 0-4, first three columns.\n",
    "penguins.iloc[0:5, 0:3]"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c4703ae6",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.991036Z",
     "iopub.status.busy": "2026-08-30T12:32:43.990484Z",
     "iopub.status.idle": "2026-08-30T12:32:43.997973Z",
     "shell.execute_reply": "2026-08-30T12:32:43.997208Z"
    }
   },
   "outputs": [],
   "source": [
    "# loc: by label. Rows 0-4 (inclusive!), named columns.\n",
    "penguins.loc[0:4, [\"species\", \"island\", \"bill_length_mm\"]]"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3b01d09d",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:43.999552Z",
     "iopub.status.busy": "2026-08-30T12:32:43.999335Z",
     "iopub.status.idle": "2026-08-30T12:32:44.004083Z",
     "shell.execute_reply": "2026-08-30T12:32:44.003322Z"
    }
   },
   "outputs": [],
   "source": [
    "# A single column is a Series; a list of columns is a DataFrame\n",
    "print(type(penguins[\"bill_length_mm\"]))\n",
    "print(type(penguins[[\"bill_length_mm\"]]))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6de7b6af",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Filtering and Missing Values</h3>\n",
    "\n",
    "\n",
    "Real data has missing information. `isna()` finds them and `dropna()` removes the rows with missing `NULL` values.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "9da04310",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.005584Z",
     "iopub.status.busy": "2026-08-30T12:32:44.005442Z",
     "iopub.status.idle": "2026-08-30T12:32:44.011484Z",
     "shell.execute_reply": "2026-08-30T12:32:44.010120Z"
    }
   },
   "outputs": [],
   "source": [
    "penguins.isna().sum()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "0b31e3d3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.013017Z",
     "iopub.status.busy": "2026-08-30T12:32:44.012871Z",
     "iopub.status.idle": "2026-08-30T12:32:44.022728Z",
     "shell.execute_reply": "2026-08-30T12:32:44.021895Z"
    }
   },
   "outputs": [],
   "source": [
    "# Boolean filtering: which penguins are both heavy and long-flippered?\n",
    "mask = (penguins[\"body_mass_g\"] > 5000) & (penguins[\"flipper_length_mm\"] > 220)\n",
    "penguins[mask].head()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "30678dbd",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.024659Z",
     "iopub.status.busy": "2026-08-30T12:32:44.024049Z",
     "iopub.status.idle": "2026-08-30T12:32:44.035353Z",
     "shell.execute_reply": "2026-08-30T12:32:44.034451Z"
    }
   },
   "outputs": [],
   "source": [
    "df = penguins.dropna()\n",
    "print(len(penguins), \"rows ->\", len(df), \"rows after dropping missing values\")\n",
    "df.head()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "aa175fc6",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Grouping and Summarizing</h3>\n",
    "\n",
    "\n",
    "`groupby()` splits the data into groups, applies a function to each, and combines the results.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c28f1844",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.037363Z",
     "iopub.status.busy": "2026-08-30T12:32:44.036738Z",
     "iopub.status.idle": "2026-08-30T12:32:44.045307Z",
     "shell.execute_reply": "2026-08-30T12:32:44.044494Z"
    }
   },
   "outputs": [],
   "source": [
    "df.groupby(\"species\")[[\"bill_length_mm\", \"flipper_length_mm\", \"body_mass_g\"]].mean().round(1)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1ba01dac",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">`numpy`: the Array Underneath</h3>\n",
    "\n",
    "\n",
    "`pandas` is built on `numpy`. Models in `scikit-learn` want numpy arrays, so this is how we hand our\n",
    "data over.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "302eca5e",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.047351Z",
     "iopub.status.busy": "2026-08-30T12:32:44.046712Z",
     "iopub.status.idle": "2026-08-30T12:32:44.052504Z",
     "shell.execute_reply": "2026-08-30T12:32:44.051718Z"
    }
   },
   "outputs": [],
   "source": [
    "FEATURES = [\"bill_length_mm\", \"flipper_length_mm\"]\n",
    "X = df[FEATURES].to_numpy()      # feature matrix, shape (n_samples, n_features)\n",
    "y = df[\"species\"].to_numpy()     # labels, shape (n_samples,)\n",
    "\n",
    "print(\"X shape:\", X.shape, \"| y shape:\", y.shape)\n",
    "print(\"X dtype:\", X.dtype, \"\\n\")\n",
    "print(\"The first three rows of X:\\n\", X[:3])"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1ef13456",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.053790Z",
     "iopub.status.busy": "2026-08-30T12:32:44.053620Z",
     "iopub.status.idle": "2026-08-30T12:32:44.058717Z",
     "shell.execute_reply": "2026-08-30T12:32:44.057916Z"
    }
   },
   "outputs": [],
   "source": [
    "# Vectorized operations act on the whole array at once, with no Python loop.\n",
    "print(\"column means:\", X.mean(axis=0).round(2))\n",
    "print(\"column stds: \", X.std(axis=0).round(2))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "99a1dd27",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Visualizing: Can We See the Species?</h3>\n",
    "\n",
    "If the species separate visually in these two dimensions, then a model has a chance.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "89a0d6f0",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.060270Z",
     "iopub.status.busy": "2026-08-30T12:32:44.060088Z",
     "iopub.status.idle": "2026-08-30T12:32:44.287075Z",
     "shell.execute_reply": "2026-08-30T12:32:44.286014Z"
    }
   },
   "outputs": [],
   "source": [
    "fig = px.scatter(\n",
    "    df, x=\"bill_length_mm\", y=\"flipper_length_mm\", color=\"species\",\n",
    "    title=\"Palmer Penguins: bill length vs flipper length\",\n",
    "    labels={\n",
    "        \"bill_length_mm\": \"Bill length (mm)\",\n",
    "        \"flipper_length_mm\": \"Flipper length (mm)\",\n",
    "        \"species\": \"Species\",\n",
    "    },\n",
    "    height=520,\n",
    ")\n",
    "# Plotly's default legend sits off to the right and often gets clipped in notebooks.\n",
    "fig.update_layout(\n",
    "    showlegend=True,\n",
    "    legend=dict(\n",
    "        title_text=\"Species\",\n",
    "        x=0.01,\n",
    "        y=0.99,\n",
    "        xanchor=\"left\",\n",
    "        yanchor=\"top\",\n",
    "        bgcolor=\"white\",\n",
    "        bordercolor=\"lightgray\",\n",
    "        borderwidth=1,\n",
    "    ),\n",
    ")\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4df78113",
   "metadata": {},
   "source": [
    "**Look at what this plot is telling us.** The three species occupy different regions. Nothing\n",
    "separates them perfectly, but points near each other tend to share a species."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "eaf9d663",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 2: The Simplest Possible Model</h2>\n",
    "\n",
    "\n",
    "We want to predict a penguin's species from its bill length and flipper length. We ask ourselves, what is the least clever thing that could possibly work?\n",
    "\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "391cd0c5",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">K-Nearest Neighbors (KNN) with k = 1</h3>\n",
    "\n",
    "\n",
    "**To predict the species of a new penguin, find the penguin in our data that is closest to it, and copy that penguin's species.**\n",
    "\n",
    "That is the entire algorithm. Notice what it does *not* have any optimization. \"Training\" is nothing more than storing the data.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "18f24c55",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:44.289480Z",
     "iopub.status.busy": "2026-08-30T12:32:44.288806Z",
     "iopub.status.idle": "2026-08-30T12:32:45.355902Z",
     "shell.execute_reply": "2026-08-30T12:32:45.354867Z"
    }
   },
   "outputs": [],
   "source": [
    "from sklearn.neighbors import KNeighborsClassifier\n",
    "\n",
    "knn1 = KNeighborsClassifier(n_neighbors=1)\n",
    "knn1.fit(X, y)          # \"training\": store the data\n",
    "print(\"Model is trained.\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3fa4a8f3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.358559Z",
     "iopub.status.busy": "2026-08-30T12:32:45.357735Z",
     "iopub.status.idle": "2026-08-30T12:32:45.364814Z",
     "shell.execute_reply": "2026-08-30T12:32:45.363989Z"
    }
   },
   "outputs": [],
   "source": [
    "# Inference: predict the species of a new penguin\n",
    "new_penguin = np.array([[45.0, 200.0]])   # bill 45mm, flipper 200mm\n",
    "print(\"Predicted species:\", knn1.predict(new_penguin)[0])"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "ec376a7c",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">How Good Is It?</h3>\n",
    "\n",
    "We ask the model to predict the species of every penguin in our dataset, and count how often it is right."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "7933447f",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.366941Z",
     "iopub.status.busy": "2026-08-30T12:32:45.366304Z",
     "iopub.status.idle": "2026-08-30T12:32:45.373438Z",
     "shell.execute_reply": "2026-08-30T12:32:45.372532Z"
    }
   },
   "outputs": [],
   "source": [
    "accuracy = knn1.score(X, y)\n",
    "print(f\"Accuracy on our data: {accuracy:.1%}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4c9c556e",
   "metadata": {},
   "source": [
    "### 100%.\n",
    "\n",
    "We just built a perfect classifier out of the dumbest idea.\n",
    "\n",
    "**Something is wrong.** What is it?"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3910bf7c",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.375297Z",
     "iopub.status.busy": "2026-08-30T12:32:45.374981Z",
     "iopub.status.idle": "2026-08-30T12:32:45.380749Z",
     "shell.execute_reply": "2026-08-30T12:32:45.379799Z"
    }
   },
   "outputs": [],
   "source": [
    "# Why 100%? Ask which point is nearest to the first penguin in our data.\n",
    "distances, indices = knn1.kneighbors(X[:1], n_neighbors=1)\n",
    "print(\"Nearest neighbour of penguin 0 is penguin index:\", indices[0][0])\n",
    "print(\"at a distance of:\", distances[0][0])"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "50f41728",
   "metadata": {},
   "source": [
    "The nearest neighbour of every penguin in our dataset **is itself**, at distance zero. We asked the model to answer questions it had already memorized the answers to.\n",
    "\n",
    "This tells us nothing about whether the model has *learned* anything. We need to evaluate it on penguins it has never seen."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "7a53d90e",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 3: Generalization and the Train/Test Split</h2>\n",
    "\n",
    "\n",
    "**Generalization** is the ability to perform well on new, unseen data drawn from the same\n",
    "distribution. To measure it, we hold data back."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "6b3f7d51",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.382589Z",
     "iopub.status.busy": "2026-08-30T12:32:45.382376Z",
     "iopub.status.idle": "2026-08-30T12:32:45.388780Z",
     "shell.execute_reply": "2026-08-30T12:32:45.387856Z"
    }
   },
   "outputs": [],
   "source": [
    "from sklearn.model_selection import train_test_split\n",
    "\n",
    "X_train, X_test, y_train, y_test = train_test_split(\n",
    "    X, y, test_size=0.2, random_state=42, stratify=y\n",
    ")\n",
    "print(f\"Training set: {X_train.shape[0]} penguins\")\n",
    "print(f\"Test set:     {X_test.shape[0]} penguins\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "ac5216bc",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.390838Z",
     "iopub.status.busy": "2026-08-30T12:32:45.390225Z",
     "iopub.status.idle": "2026-08-30T12:32:45.399783Z",
     "shell.execute_reply": "2026-08-30T12:32:45.398948Z"
    }
   },
   "outputs": [],
   "source": [
    "knn1 = KNeighborsClassifier(n_neighbors=1).fit(X_train, y_train)\n",
    "\n",
    "print(f\"Training accuracy: {knn1.score(X_train, y_train):.1%}\")\n",
    "print(f\"Test accuracy:     {knn1.score(X_test, y_test):.1%}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "42b0ca6e",
   "metadata": {},
   "source": [
    "Training accuracy is still 100%, and it always will be for $k=1$, because every training point is still its own nearest neighbour. The test accuracy is the honest number, and it is meaningfully lower.\n",
    "\n",
    "The gap between those two numbers is the thing we spend the rest of the semester managing."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "85e9e046",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 4: k Is a Choice</h2>\n",
    "\n",
    "\n",
    "Why look at only one neighbour? Let each of the $k$ nearest neighbours vote, and take the majority.\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a772c4c8",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Seeing k in the Decision Boundary</h3>\n",
    "\n",
    "\n",
    "A **decision boundary** shows what the model would predict at every point in the feature space. It makes the effect of $k$ visible.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "5d092fe3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.401823Z",
     "iopub.status.busy": "2026-08-30T12:32:45.401230Z",
     "iopub.status.idle": "2026-08-30T12:32:45.490800Z",
     "shell.execute_reply": "2026-08-30T12:32:45.489809Z"
    }
   },
   "outputs": [],
   "source": [
    "import plotly.graph_objects as go\n",
    "from sklearn.preprocessing import LabelEncoder\n",
    "\n",
    "def decision_boundary_figure(k, X_tr, y_tr, resolution=250):\n",
    "    le = LabelEncoder().fit(y_tr)\n",
    "    model = KNeighborsClassifier(n_neighbors=k).fit(X_tr, le.transform(y_tr))\n",
    "\n",
    "    pad = 1.0\n",
    "    xs = np.linspace(X_tr[:, 0].min() - pad, X_tr[:, 0].max() + pad, resolution)\n",
    "    ys = np.linspace(X_tr[:, 1].min() - pad, X_tr[:, 1].max() + pad, resolution)\n",
    "    xx, yy = np.meshgrid(xs, ys)\n",
    "    zz = model.predict(np.c_[xx.ravel(), yy.ravel()]).reshape(xx.shape)\n",
    "\n",
    "    fig = go.Figure()\n",
    "    fig.add_trace(go.Heatmap(x=xs, y=ys, z=zz, showscale=False, opacity=0.30,\n",
    "                             colorscale=[[0, \"#002675\"], [0.5, \"#FDB515\"], [1, \"#028842\"]]))\n",
    "    for i, cls in enumerate(le.classes_):\n",
    "        m = y_tr == cls\n",
    "        fig.add_trace(go.Scatter(\n",
    "            x=X_tr[m, 0], y=X_tr[m, 1], mode=\"markers\", name=cls,\n",
    "            marker=dict(size=7, line=dict(width=1, color=\"white\"),\n",
    "                        color=[\"#002675\", \"#FDB515\", \"#028842\"][i])))\n",
    "    fig.update_layout(title=f\"KNN decision boundary, k = {k}\",\n",
    "                      xaxis_title=\"Bill length (mm)\", yaxis_title=\"Flipper length (mm)\",\n",
    "                      width=800, height=520)\n",
    "    return fig\n",
    "\n",
    "decision_boundary_figure(1, X_train, y_train).show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "34c30ffc",
   "metadata": {},
   "source": [
    "**k = 1 gives a noisy boundary** with little islands around individual points. The model is contorting itself to get every single training penguin right, including the ones that are unusual.\n",
    "That is **overfitting**: fitting the noise as well as the signal."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a44fb113",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.492825Z",
     "iopub.status.busy": "2026-08-30T12:32:45.492569Z",
     "iopub.status.idle": "2026-08-30T12:32:45.698600Z",
     "shell.execute_reply": "2026-08-30T12:32:45.697536Z"
    }
   },
   "outputs": [],
   "source": [
    "decision_boundary_figure(15, X_train, y_train).show()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "db53c6be",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:45.700864Z",
     "iopub.status.busy": "2026-08-30T12:32:45.700226Z",
     "iopub.status.idle": "2026-08-30T12:32:47.002301Z",
     "shell.execute_reply": "2026-08-30T12:32:47.001083Z"
    }
   },
   "outputs": [],
   "source": [
    "decision_boundary_figure(100, X_train, y_train).show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1665976d",
   "metadata": {},
   "source": [
    "**k = 15** smooths the boundary into something that looks like a real trend.\n",
    "\n",
    "**k = 100** smooths it so much that the model is barely paying attention to the data. That is\n",
    "**underfitting**.\n",
    "\n",
    "So $k$ controls a trade-off, and somewhere in between is a good value.\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a0b9ac87",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Finding a Good k</h3>\n",
    "\n",
    "\n",
    "We cannot use the test set to pick $k$. The moment we tune anything against the test set, it stops measuring generalization and becomes just another training set.\n",
    "\n",
    "So we split again: a **validation set**, carved out of the training data.\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "4d8e86a0",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.004259Z",
     "iopub.status.busy": "2026-08-30T12:32:47.003917Z",
     "iopub.status.idle": "2026-08-30T12:32:47.010738Z",
     "shell.execute_reply": "2026-08-30T12:32:47.009732Z"
    }
   },
   "outputs": [],
   "source": [
    "X_tr, X_val, y_tr, y_val = train_test_split(\n",
    "    X_train, y_train, test_size=0.25, random_state=42, stratify=y_train\n",
    ")\n",
    "print(f\"train {len(X_tr)} | validation {len(X_val)} | test {len(X_test)}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "5250be8c",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.012311Z",
     "iopub.status.busy": "2026-08-30T12:32:47.012110Z",
     "iopub.status.idle": "2026-08-30T12:32:47.501235Z",
     "shell.execute_reply": "2026-08-30T12:32:47.500254Z"
    }
   },
   "outputs": [],
   "source": [
    "ks = list(range(1, 101))\n",
    "train_acc, val_acc = [], []\n",
    "for k in ks:\n",
    "    m = KNeighborsClassifier(n_neighbors=k).fit(X_tr, y_tr)\n",
    "    train_acc.append(m.score(X_tr, y_tr))\n",
    "    val_acc.append(m.score(X_val, y_val))\n",
    "\n",
    "best_k = ks[int(np.argmax(val_acc))]\n",
    "print(f\"Best k on the validation set: {best_k}  (validation accuracy {max(val_acc):.1%})\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "d7124395",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.503427Z",
     "iopub.status.busy": "2026-08-30T12:32:47.502772Z",
     "iopub.status.idle": "2026-08-30T12:32:47.519293Z",
     "shell.execute_reply": "2026-08-30T12:32:47.518400Z"
    }
   },
   "outputs": [],
   "source": [
    "fig = go.Figure()\n",
    "fig.add_trace(go.Scatter(x=ks, y=train_acc, name=\"Training accuracy\",\n",
    "                         line=dict(color=\"#002675\", width=3)))\n",
    "fig.add_trace(go.Scatter(x=ks, y=val_acc, name=\"Validation accuracy\",\n",
    "                         line=dict(color=\"#FDB515\", width=3)))\n",
    "fig.add_vline(x=best_k, line_dash=\"dash\", line_color=\"#028842\",\n",
    "              annotation_text=f\"best k = {best_k}\")\n",
    "fig.update_layout(title=\"Training and validation accuracy as k increases\",\n",
    "                  xaxis_title=\"k (number of neighbours)\", yaxis_title=\"Accuracy\",\n",
    "                  width=800, height=520)\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "1799f4f5",
   "metadata": {},
   "source": [
    "Read this plot carefully, because its shape recurs all semester.\n",
    "\n",
    "- On the **left** (small $k$) training accuracy is perfect and validation accuracy is lower. Overfitting.\n",
    "- On the **right** (large $k$) both accuracies fall together. Underfitting.\n",
    "- The **sweet spot** is where validation accuracy peaks.\n",
    "\n",
    "$k$ is a **hyperparameter**: a value we choose *before* training, as opposed to a **parameter**, which is learned *from* the data during training. KNN is unusual in having no parameters at all, only this one hyperparameter. That makes it a clean place to see the distinction.\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a2a4b748",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.521535Z",
     "iopub.status.busy": "2026-08-30T12:32:47.520887Z",
     "iopub.status.idle": "2026-08-30T12:32:47.527975Z",
     "shell.execute_reply": "2026-08-30T12:32:47.527237Z"
    }
   },
   "outputs": [],
   "source": [
    "# Only now, having chosen k, do we touch the test set. Once.\n",
    "final = KNeighborsClassifier(n_neighbors=best_k).fit(X_train, y_train)\n",
    "print(f\"Final test accuracy with k={best_k}: {final.score(X_test, y_test):.1%}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a4f1573d",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A Wrinkle: Distance Depends on Units</h3>\n",
    "\n",
    "\n",
    "KNN is built entirely on distance, so it cares about the scale of the features. Flipper length spans roughly 172-231 mm and bill length roughly 32-60 mm, so flipper length contributes far more to the distance simply because its numbers are bigger.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8c806c99",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.530071Z",
     "iopub.status.busy": "2026-08-30T12:32:47.529497Z",
     "iopub.status.idle": "2026-08-30T12:32:47.541198Z",
     "shell.execute_reply": "2026-08-30T12:32:47.540396Z"
    }
   },
   "outputs": [],
   "source": [
    "from sklearn.pipeline import make_pipeline\n",
    "from sklearn.preprocessing import StandardScaler\n",
    "\n",
    "scaled = make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=best_k))\n",
    "scaled.fit(X_train, y_train)\n",
    "\n",
    "print(f\"Unscaled test accuracy: {final.score(X_test, y_test):.1%}\")\n",
    "print(f\"Scaled test accuracy:   {scaled.score(X_test, y_test):.1%}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "62a50bc0",
   "metadata": {},
   "source": [
    "**Standardization** rescales each feature to have mean 0 and standard deviation 1, so every\n",
    "feature contributes to the distance on equal terms. For any distance-based model this is not\n",
    "optional.\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "aea6098a",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 5: The Same Idea for Regression</h2>\n",
    "\n",
    "\n",
    "So far we predicted a **category** (species). What if we want to predict a **number**, like body mass?\n",
    "\n",
    "The algorithm barely changes. Find the k nearest neighbours, and instead of taking a majority vote, take their **average**.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b3af2b7d",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.543342Z",
     "iopub.status.busy": "2026-08-30T12:32:47.542695Z",
     "iopub.status.idle": "2026-08-30T12:32:47.548942Z",
     "shell.execute_reply": "2026-08-30T12:32:47.548154Z"
    }
   },
   "outputs": [],
   "source": [
    "from sklearn.neighbors import KNeighborsRegressor\n",
    "\n",
    "Xr = df[[\"flipper_length_mm\"]].to_numpy()\n",
    "yr = df[\"body_mass_g\"].to_numpy()\n",
    "\n",
    "Xr_train, Xr_test, yr_train, yr_test = train_test_split(Xr, yr, test_size=0.2, random_state=42)\n",
    "print(\"Predicting body mass (g) from flipper length (mm)\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "dc7c8c84",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.551002Z",
     "iopub.status.busy": "2026-08-30T12:32:47.550390Z",
     "iopub.status.idle": "2026-08-30T12:32:47.582731Z",
     "shell.execute_reply": "2026-08-30T12:32:47.581862Z"
    }
   },
   "outputs": [],
   "source": [
    "grid = np.linspace(Xr.min(), Xr.max(), 400).reshape(-1, 1)\n",
    "\n",
    "fig = go.Figure()\n",
    "fig.add_trace(go.Scatter(x=Xr_train.ravel(), y=yr_train, mode=\"markers\", name=\"Training data\",\n",
    "                         marker=dict(color=\"lightgray\", size=6)))\n",
    "for k, colour in [(1, \"#002675\"), (25, \"#FDB515\"), (200, \"#028842\")]:\n",
    "    reg = KNeighborsRegressor(n_neighbors=k).fit(Xr_train, yr_train)\n",
    "    fig.add_trace(go.Scatter(x=grid.ravel(), y=reg.predict(grid), mode=\"lines\",\n",
    "                             name=f\"k = {k}\", line=dict(width=3, color=colour)))\n",
    "fig.update_layout(title=\"KNN regression: body mass vs flipper length\",\n",
    "                  xaxis_title=\"Flipper length (mm)\", yaxis_title=\"Body mass (g)\",\n",
    "                  width=800, height=520)\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "44f11211",
   "metadata": {},
   "source": [
    "The **same trade-off**, in a different costume. $k=1$ is a jagged step function chasing every point. $k=200$ is nearly a flat line. The middle is where the real trend lives.\n",
    "\n",
    "Overfitting and underfitting are not facts about classification. They are facts about model complexity.\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "3a2d312f",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.584313Z",
     "iopub.status.busy": "2026-08-30T12:32:47.584116Z",
     "iopub.status.idle": "2026-08-30T12:32:47.601399Z",
     "shell.execute_reply": "2026-08-30T12:32:47.600361Z"
    }
   },
   "outputs": [],
   "source": [
    "for k in [1, 25, 200]:\n",
    "    reg = KNeighborsRegressor(n_neighbors=k).fit(Xr_train, yr_train)\n",
    "    print(f\"k={k:>3}  train R^2 = {reg.score(Xr_train, yr_train):.3f}   \"\n",
    "          f\"test R^2 = {reg.score(Xr_test, yr_test):.3f}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "56a6d713",
   "metadata": {},
   "source": [
    "At **k=1** the training score is far above the test score: the model is memorizing. \n",
    "At **k=200** the two scores collapse *together*, and both are bad: the model is\n",
    "too simple to capture the trend. The same story the classification curve told, told again.\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5de0156d",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 6: What If We Had No Labels?</h2>\n",
    "\n",
    "\n",
    "Everything so far was **supervised**: every penguin came with its species attached.\n",
    "\n",
    "Now suppose a field researcher hands us the same measurements with **no species labels at all**, and\n",
    "asks: are there natural groups in here?\n",
    "\n",
    "This is **unsupervised learning**, and specifically **clustering**.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "628e248d",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.603217Z",
     "iopub.status.busy": "2026-08-30T12:32:47.602724Z",
     "iopub.status.idle": "2026-08-30T12:32:47.633470Z",
     "shell.execute_reply": "2026-08-30T12:32:47.632655Z"
    }
   },
   "outputs": [],
   "source": [
    "# Throw the labels away.\n",
    "fig = px.scatter(df, x=\"bill_length_mm\", y=\"flipper_length_mm\",\n",
    "                 title=\"The same penguins, with no labels\",\n",
    "                 labels={\"bill_length_mm\": \"Bill length (mm)\",\n",
    "                         \"flipper_length_mm\": \"Flipper length (mm)\"},\n",
    "                 height=520)\n",
    "fig.update_traces(marker=dict(color=\"gray\", size=7))\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "3d5d2ebb",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">K-Means in scikit-learn</h3>\n",
    "\n",
    "\n",
    "K-means partitions the data into K clusters, each represented by a **centroid**. It assigns each\n",
    "point to the nearest centroid, then moves each centroid to the mean of its points, and repeats.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "8318089b",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.635802Z",
     "iopub.status.busy": "2026-08-30T12:32:47.635197Z",
     "iopub.status.idle": "2026-08-30T12:32:47.700076Z",
     "shell.execute_reply": "2026-08-30T12:32:47.699177Z"
    }
   },
   "outputs": [],
   "source": [
    "from sklearn.cluster import KMeans\n",
    "\n",
    "Xs = StandardScaler().fit_transform(X)   # distance-based again, so standardize\n",
    "\n",
    "kmeans = KMeans(n_clusters=3, n_init=10, random_state=42)\n",
    "cluster = kmeans.fit_predict(Xs)\n",
    "\n",
    "df_c = df.copy()\n",
    "df_c[\"cluster\"] = cluster.astype(str)\n",
    "px.scatter(df_c, x=\"bill_length_mm\", y=\"flipper_length_mm\", color=\"cluster\",\n",
    "           title=\"K-means with K = 3\",\n",
    "           labels={\"bill_length_mm\": \"Bill length (mm)\",\n",
    "                   \"flipper_length_mm\": \"Flipper length (mm)\"},\n",
    "           height=520).show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "edd801bc",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Lloyd's Algorithm, Step by Step</h3>\n",
    "\n",
    "\n",
    "Watch the centroids move. Each iteration is two steps: assign, then update.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "16adcff3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.702342Z",
     "iopub.status.busy": "2026-08-30T12:32:47.701630Z",
     "iopub.status.idle": "2026-08-30T12:32:47.708490Z",
     "shell.execute_reply": "2026-08-30T12:32:47.707618Z"
    }
   },
   "outputs": [],
   "source": [
    "def lloyd_steps(Xs, K=3, n_iter=6, seed=1):\n",
    "    rng = np.random.default_rng(seed)\n",
    "    centres = Xs[rng.choice(len(Xs), K, replace=False)].copy()\n",
    "    history = []\n",
    "    for _ in range(n_iter):\n",
    "        d = ((Xs[:, None, :] - centres[None, :, :]) ** 2).sum(axis=2)\n",
    "        assign = d.argmin(axis=1)                       # assignment step\n",
    "        history.append((centres.copy(), assign.copy()))\n",
    "        for j in range(K):                              # update step\n",
    "            if (assign == j).any():\n",
    "                centres[j] = Xs[assign == j].mean(axis=0)\n",
    "    return history\n",
    "\n",
    "history = lloyd_steps(Xs)\n",
    "print(f\"{len(history)} iterations recorded\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "0d0a1ca7",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.710027Z",
     "iopub.status.busy": "2026-08-30T12:32:47.709869Z",
     "iopub.status.idle": "2026-08-30T12:32:47.741235Z",
     "shell.execute_reply": "2026-08-30T12:32:47.740354Z"
    }
   },
   "outputs": [],
   "source": [
    "palette = [\"#002675\", \"#FDB515\", \"#028842\"]\n",
    "fig = go.Figure()\n",
    "frames = []\n",
    "for step, (centres, assign) in enumerate(history):\n",
    "    data = []\n",
    "    for j in range(3):\n",
    "        m = assign == j\n",
    "        data.append(go.Scatter(x=Xs[m, 0], y=Xs[m, 1], mode=\"markers\",\n",
    "                               marker=dict(color=palette[j], size=6), name=f\"cluster {j}\"))\n",
    "    data.append(go.Scatter(x=centres[:, 0], y=centres[:, 1], mode=\"markers\",\n",
    "                           marker=dict(color=\"black\", size=18, symbol=\"x\"), name=\"centroids\"))\n",
    "    frames.append(go.Frame(data=data, name=str(step)))\n",
    "\n",
    "fig.add_traces(frames[0].data)\n",
    "fig.frames = frames\n",
    "fig.update_layout(\n",
    "    title=\"Lloyd's algorithm: assign, then update\",\n",
    "    xaxis_title=\"Bill length (standardized)\", yaxis_title=\"Flipper length (standardized)\",\n",
    "    width=800, height=560,\n",
    "    updatemenus=[dict(type=\"buttons\", showactive=False,\n",
    "                      buttons=[dict(label=\"Play\", method=\"animate\",\n",
    "                                    args=[None, dict(frame=dict(duration=900, redraw=True),\n",
    "                                                     fromcurrent=True)])])])\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "132b1cd2",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">Choosing K</h3>\n",
    "\n",
    "\n",
    "K is a hyperparameter, exactly like k in KNN. But there is no validation accuracy to optimize,\n",
    "because there are no labels. The common heuristic is the **elbow method**: plot the within-cluster\n",
    "sum of squares (inertia) against K and look for the bend.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "83fa7cb2",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.743351Z",
     "iopub.status.busy": "2026-08-30T12:32:47.742697Z",
     "iopub.status.idle": "2026-08-30T12:32:47.852559Z",
     "shell.execute_reply": "2026-08-30T12:32:47.851437Z"
    }
   },
   "outputs": [],
   "source": [
    "Ks = range(1, 11)\n",
    "inertia = [KMeans(n_clusters=k, n_init=10, random_state=42).fit(Xs).inertia_ for k in Ks]\n",
    "\n",
    "fig = px.line(x=list(Ks), y=inertia, markers=True,\n",
    "              title=\"Elbow method: inertia vs K\",\n",
    "              labels={\"x\": \"K (number of clusters)\", \"y\": \"Within-cluster sum of squares\"})\n",
    "fig.update_traces(line=dict(color=\"#002675\", width=3))\n",
    "fig.update_layout(width=800, height=480)\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "9be70a9b",
   "metadata": {},
   "source": [
    "The bend is around K=3, which is encouraging. But note how much judgment that reading takes.\n",
    "The elbow is a heuristic, not a criterion.\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e46744ab",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Part 7: Clusters Are Not Labels</h2>\n",
    "\n",
    "\n",
    "We have run two models on the same two columns. One was told the species; one was not. How much did\n",
    "the labels actually buy us?\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "50089a5b",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.854331Z",
     "iopub.status.busy": "2026-08-30T12:32:47.854088Z",
     "iopub.status.idle": "2026-08-30T12:32:47.867447Z",
     "shell.execute_reply": "2026-08-30T12:32:47.866571Z"
    }
   },
   "outputs": [],
   "source": [
    "comparison = pd.crosstab(df_c[\"cluster\"], df_c[\"species\"])\n",
    "comparison"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "6588fcd3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.869522Z",
     "iopub.status.busy": "2026-08-30T12:32:47.868906Z",
     "iopub.status.idle": "2026-08-30T12:32:47.909807Z",
     "shell.execute_reply": "2026-08-30T12:32:47.908903Z"
    }
   },
   "outputs": [],
   "source": [
    "fig = px.scatter(df_c, x=\"bill_length_mm\", y=\"flipper_length_mm\",\n",
    "                 color=\"species\", symbol=\"cluster\",\n",
    "                 title=\"True species (colour) vs discovered clusters (symbol)\",\n",
    "                 labels={\"bill_length_mm\": \"Bill length (mm)\",\n",
    "                         \"flipper_length_mm\": \"Flipper length (mm)\"},\n",
    "                 height=560)\n",
    "fig.show()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "5a3038bd",
   "metadata": {},
   "source": [
    "K-means, with no access to the labels at all, largely rediscovered the species.\n",
    "\n",
    "**But be careful about what that means.** The clusters are not species. K-means found groups of\n",
    "penguins that are close together in bill and flipper measurements, and in this dataset those groups\n",
    "happen to line up with species. Change the features and the clusters change. Nothing in the\n",
    "algorithm knows what a species is, and nothing guarantees the groups it finds correspond to anything\n",
    "you care about.\n",
    "\n",
    "<!-- [SLIDO POLL 4 GOES HERE] -->\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e275e402",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h2 class=\"cal cal-h2\">Appendix: Going Deeper</h2>\n",
    "\n",
    "\n",
    "The material below was not covered in lecture. Work through it on your own; it is the reference for the syntax used above and for Homework 1.\n",
    "\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b68ccc0b",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A1. `pandas` Data Structures</h3>\n",
    "\n",
    "\n",
    "A `DataFrame` is a 2-dimensional table. A `Series` is a single column. Both are built on an `Index`.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "0d235a90",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.912096Z",
     "iopub.status.busy": "2026-08-30T12:32:47.911525Z",
     "iopub.status.idle": "2026-08-30T12:32:47.916970Z",
     "shell.execute_reply": "2026-08-30T12:32:47.916158Z"
    }
   },
   "outputs": [],
   "source": [
    "s = pd.Series([3750, 3800, 3250], index=[\"p0\", \"p1\", \"p2\"], name=\"body_mass_g\")\n",
    "print(s)\n",
    "print(\"\\nindex:\", s.index.tolist())\n",
    "print(\"values:\", s.values)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "bff0624a",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.919082Z",
     "iopub.status.busy": "2026-08-30T12:32:47.918462Z",
     "iopub.status.idle": "2026-08-30T12:32:47.926308Z",
     "shell.execute_reply": "2026-08-30T12:32:47.925450Z"
    }
   },
   "outputs": [],
   "source": [
    "frame = pd.DataFrame({\n",
    "    \"species\": [\"Adelie\", \"Gentoo\", \"Chinstrap\"],\n",
    "    \"bill_length_mm\": [39.1, 46.1, 46.5],\n",
    "    \"island\": [\"Torgersen\", \"Biscoe\", \"Dream\"],\n",
    "})\n",
    "frame"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "0f7b6c4d",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A2. Exploring a `DataFrame`</h3>\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "2cf2632c",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.927901Z",
     "iopub.status.busy": "2026-08-30T12:32:47.927746Z",
     "iopub.status.idle": "2026-08-30T12:32:47.939658Z",
     "shell.execute_reply": "2026-08-30T12:32:47.938822Z"
    }
   },
   "outputs": [],
   "source": [
    "print(penguins.head(3))          # first rows\n",
    "print(penguins.tail(3))          # last rows\n",
    "print(penguins.sample(3))        # random rows\n",
    "print(penguins.columns.tolist()) # column names\n",
    "print(penguins.dtypes)           # column types\n",
    "print(penguins[\"island\"].unique())"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "f03858fe",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A3. `loc` and `iloc` in Full</h3>\n",
    "\n",
    "\n",
    "The rule to remember: `iloc` is positional and excludes the endpoint, `loc` is label-based and\n",
    "includes it.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "c37e8d09",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.941665Z",
     "iopub.status.busy": "2026-08-30T12:32:47.941017Z",
     "iopub.status.idle": "2026-08-30T12:32:47.950194Z",
     "shell.execute_reply": "2026-08-30T12:32:47.949369Z"
    }
   },
   "outputs": [],
   "source": [
    "print(penguins.iloc[0])              # row 0 as a Series\n",
    "print(penguins.iloc[0:3])            # rows 0,1,2\n",
    "print(penguins.iloc[:, 2])           # column at position 2\n",
    "print(penguins.iloc[[0, 5, 10], [0, 2]])   # arbitrary rows and columns"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "981ffbd3",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.951713Z",
     "iopub.status.busy": "2026-08-30T12:32:47.951568Z",
     "iopub.status.idle": "2026-08-30T12:32:47.961568Z",
     "shell.execute_reply": "2026-08-30T12:32:47.960633Z"
    }
   },
   "outputs": [],
   "source": [
    "print(penguins.loc[0:2])                          # rows labelled 0,1,2 (inclusive)\n",
    "print(penguins.loc[:, \"species\":\"bill_length_mm\"]) # column slice by name\n",
    "print(penguins.loc[penguins[\"species\"] == \"Gentoo\", \"body_mass_g\"].mean())"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6fec7d8d",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A4. Modifying and Sorting</h3>\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1cdb5295",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.963626Z",
     "iopub.status.busy": "2026-08-30T12:32:47.962971Z",
     "iopub.status.idle": "2026-08-30T12:32:47.974940Z",
     "shell.execute_reply": "2026-08-30T12:32:47.974111Z"
    }
   },
   "outputs": [],
   "source": [
    "tmp = df.copy()\n",
    "tmp[\"mass_kg\"] = tmp[\"body_mass_g\"] / 1000              # add a column\n",
    "tmp[\"bill_ratio\"] = tmp[\"bill_length_mm\"] / tmp[\"bill_depth_mm\"]\n",
    "tmp = tmp.drop(columns=[\"bill_ratio\"])                  # drop a column\n",
    "tmp.sort_values(\"body_mass_g\", ascending=False).head()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4234e298",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A5. Aggregation, `groupby`, and Pivot Tables</h3>\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "2b4039fb",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.976528Z",
     "iopub.status.busy": "2026-08-30T12:32:47.976380Z",
     "iopub.status.idle": "2026-08-30T12:32:47.990035Z",
     "shell.execute_reply": "2026-08-30T12:32:47.989268Z"
    }
   },
   "outputs": [],
   "source": [
    "print(df.groupby(\"species\")[\"body_mass_g\"].agg([\"mean\", \"std\", \"count\"]).round(1))\n",
    "print()\n",
    "print(df.groupby([\"species\", \"island\"])[\"flipper_length_mm\"].mean().round(1))\n",
    "print()\n",
    "print(df.pivot_table(index=\"species\", columns=\"island\",\n",
    "                     values=\"body_mass_g\", aggfunc=\"mean\").round(0))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "ccd4b599",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A6. Joining `DataFrames`</h3>\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "285a922c",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:47.992076Z",
     "iopub.status.busy": "2026-08-30T12:32:47.991488Z",
     "iopub.status.idle": "2026-08-30T12:32:48.001193Z",
     "shell.execute_reply": "2026-08-30T12:32:48.000362Z"
    }
   },
   "outputs": [],
   "source": [
    "islands = pd.DataFrame({\n",
    "    \"island\": [\"Torgersen\", \"Biscoe\", \"Dream\"],\n",
    "    \"latitude\": [-64.77, -65.43, -64.73],\n",
    "})\n",
    "merged = df.merge(islands, on=\"island\", how=\"left\")   # try how=\"inner\" and how=\"outer\"\n",
    "merged[[\"species\", \"island\", \"latitude\"]].head()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "2f687993",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A7. `numpy` Essentials</h3>\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "089cdd87",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:48.003286Z",
     "iopub.status.busy": "2026-08-30T12:32:48.002644Z",
     "iopub.status.idle": "2026-08-30T12:32:48.008596Z",
     "shell.execute_reply": "2026-08-30T12:32:48.007686Z"
    }
   },
   "outputs": [],
   "source": [
    "a = np.arange(12).reshape(3, 4)\n",
    "print(a)\n",
    "print(\"shape:\", a.shape, \"| ndim:\", a.ndim)\n",
    "print(\"row sums:\", a.sum(axis=1))\n",
    "print(\"col means:\", a.mean(axis=0))\n",
    "print(\"boolean mask:\", a[a > 6])\n",
    "print(\"broadcasting:\", (a - a.mean(axis=0)).round(2))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "b8260f5a",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:48.010036Z",
     "iopub.status.busy": "2026-08-30T12:32:48.009887Z",
     "iopub.status.idle": "2026-08-30T12:32:48.014400Z",
     "shell.execute_reply": "2026-08-30T12:32:48.013518Z"
    }
   },
   "outputs": [],
   "source": [
    "# Euclidean distance by hand, which is exactly what KNN computes\n",
    "p, q = X[0], X[1]\n",
    "print(\"manual :\", np.sqrt(((p - q) ** 2).sum()))\n",
    "print(\"numpy  :\", np.linalg.norm(p - q))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "4d43e505",
   "metadata": {},
   "source": [
    "<link rel=\"stylesheet\" href=\"berkeley.css\">\n",
    "\n",
    "<h3 class=\"cal cal-h3\">A8. Visualization Reference</h3>\n",
    "\n",
    "\n",
    "`matplotlib` and `seaborn` are the static plotting standards; `plotly` gives interactive figures,\n",
    "which is why we use it in lecture.\n",
    "\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "a7a26427",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:48.015911Z",
     "iopub.status.busy": "2026-08-30T12:32:48.015768Z",
     "iopub.status.idle": "2026-08-30T12:32:48.641285Z",
     "shell.execute_reply": "2026-08-30T12:32:48.640242Z"
    }
   },
   "outputs": [],
   "source": [
    "import matplotlib.pyplot as plt\n",
    "import seaborn as sns\n",
    "\n",
    "fig, axes = plt.subplots(1, 2, figsize=(11, 4))\n",
    "axes[0].hist(df[\"body_mass_g\"], bins=25, color=\"#002675\")\n",
    "axes[0].set_title(\"Body mass\"); axes[0].set_xlabel(\"g\")\n",
    "sns.scatterplot(data=df, x=\"bill_length_mm\", y=\"flipper_length_mm\",\n",
    "                hue=\"species\", ax=axes[1])\n",
    "axes[1].set_title(\"By species\")\n",
    "plt.tight_layout(); plt.show()"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "84c581aa",
   "metadata": {
    "execution": {
     "iopub.execute_input": "2026-08-30T12:32:48.643704Z",
     "iopub.status.busy": "2026-08-30T12:32:48.642779Z",
     "iopub.status.idle": "2026-08-30T12:32:48.753007Z",
     "shell.execute_reply": "2026-08-30T12:32:48.752082Z"
    }
   },
   "outputs": [],
   "source": [
    "# Plotly: histogram, box plot, and a faceted scatter\n",
    "px.histogram(df, x=\"body_mass_g\", color=\"species\", nbins=30, height=420).show()\n",
    "px.box(df, x=\"species\", y=\"flipper_length_mm\", color=\"species\", height=420).show()\n",
    "px.scatter(df, x=\"bill_length_mm\", y=\"flipper_length_mm\",\n",
    "           color=\"species\", facet_col=\"island\", height=420).show()"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.9.6"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
