{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "e84b9b65-4fb3-45a1-8b33-8f0021a7b185",
   "metadata": {},
   "source": "## Tokenization\n\nLecture 2 | CMU ANLP Spring 2026 | Instructor: Sean Welleck\n\nThis is a notebook for [CMU CS11-711 Advanced NLP](https://cmu-l3.github.io/anlp-spring2026/) that explores tokenizers.\n\n*Acknowledgements*: based on \"[Let's build the GPT Tokenizer](https://www.youtube.com/watch?v=zduSFxRajkE)\" by Andrej Karpathy, and its associated notebook."
  },
  {
   "cell_type": "markdown",
   "id": "60ab473a-0ea8-4edf-9d70-e9c4cb01bc82",
   "metadata": {},
   "source": [
    "### Ascii, Unicode, and UTF-8\n",
    "\n",
    "We want to support tokenizing any string of text, which includes text in a variety of languages, code, Latex, etc."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 1,
   "id": "ed3a2127-67bf-43cd-b87e-0518589e76f3",
   "metadata": {},
   "outputs": [
    {
     "data": {
      "text/plain": [
       "'元気ですか。Hello!'"
      ]
     },
     "execution_count": 1,
     "metadata": {},
     "output_type": "execute_result"
    }
   ],
   "source": [
    "\"元気ですか。Hello!\""
   ]
  },
  {
   "cell_type": "markdown",
   "id": "88067c8b-53b0-49e3-bd10-f1f0fabc2dfc",
   "metadata": {},
   "source": [
    "Tokenization requires choosing a vocabulary, i.e. the set of possible tokens.\n",
    "\n",
    "The set of [ASCII](https://en.wikipedia.org/wiki/ASCII) characters (e.g. a-zA-Z0-9) is not expressive: it cannot, for instance, represent a Japanese sentence.\n",
    "\n",
    "A potential expressive vocabulary is the set of [Unicode](https://en.wikipedia.org/wiki/Unicode) characters:"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 2,
   "id": "62099efa-46f0-46aa-9551-aa562e201541",
   "metadata": {},
   "outputs": [
    {
     "data": {
      "text/plain": [
       "[20803, 27671, 12391, 12377, 12363, 12290, 72, 101, 108, 108, 111, 33]"
      ]
     },
     "execution_count": 2,
     "metadata": {},
     "output_type": "execute_result"
    }
   ],
   "source": [
    "[ord(x) for x in \"元気ですか。Hello!\"]"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "dd501ad1-3ea6-4f0c-954f-58b0d8451aac",
   "metadata": {},
   "source": [
    "However, there are around [150,000](https://en.wikipedia.org/wiki/Unicode) characters, and the Unicode standard may change. Hence we would have a large vocabulary, would need to support changes to the Unicode standard, and it would be inefficient. For example, the string above requires 12 tokens."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "e870324f-da1d-4e95-88e3-1fcd6d8d2a65",
   "metadata": {},
   "source": [
    "A third option is to observe that UTF-8 maps Unicode into byte strings. There are 256 such byte strings, i.e. a vocabulary of 256 tokens."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 3,
   "id": "2dfe345d-7f78-47f3-a74a-4cac18d553f6",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[229, 133, 131, 230, 176, 151, 227, 129, 167, 227, 129, 153, 227, 129, 139, 227, 128, 130, 72, 101, 108, 108, 111, 33]\n"
     ]
    }
   ],
   "source": [
    "utf = \"元気ですか。Hello!\".encode(\"utf-8\")\n",
    "\n",
    "print([x for x in utf])"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "15176204-3305-4b16-bfdc-226f12d9ae7e",
   "metadata": {},
   "source": [
    "This vocabulary is expressive: it can represent any unicode sequence. \n",
    "\n",
    "However, it is inefficient. For example, the string above now requires 24 tokens."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "babd4a7a-6096-4f8b-bc7c-ddb14ddd0080",
   "metadata": {},
   "source": [
    "### Byte Pair Encoding"
   ]
  },
  {
   "attachments": {},
   "cell_type": "markdown",
   "id": "e6c764ba-c740-42b2-81bc-c005f77ef53b",
   "metadata": {},
   "source": [
    "A middle ground is to start with the size-256 UTF-8 vocabulary, and merge elements of the vocabulary to add additional vocabulary elements. This is the approach taken in [GPT-2 (see section 2.2)](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf).\n",
    "\n",
    "We will walk through one step of this merging procedure. Then, we will implement an outer loop. The resulting algorithm is called [Byte Pair Encoding (BPE)](https://en.wikipedia.org/wiki/Byte_pair_encoding).\n",
    "\n",
    "First, here is an example text and its tokens."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 4,
   "id": "6fc5b481-443a-4545-81e0-b15b604d0374",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "119\n",
      "119\n"
     ]
    }
   ],
   "source": [
    "text = \"\"\"Hello, world! Here is some example text to test\n",
    "the BPE algorithm. It is not very interesting, but it will\n",
    "do the job.\n",
    "\"\"\" \n",
    "\n",
    "tokens = text.encode(\"utf-8\")\n",
    "tokens = list(map(int, tokens))\n",
    "\n",
    "print(len(text))\n",
    "print(len(tokens))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "c33bd473",
   "metadata": {},
   "source": [
    "Next, we count the occurrences of each consecutive token pair (i.e., bigram).\n",
    "\n",
    "Let's look at the top pair."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 5,
   "id": "a242a5e2-5845-4970-8f3c-331e7c315fa5",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "(101, 32) \"e\" \" \"\n"
     ]
    }
   ],
   "source": [
    "def get_stats(ids):\n",
    "    counts = {}\n",
    "    for pair in zip(ids, ids[1:]):\n",
    "        counts[pair] = counts.get(pair, 0) + 1\n",
    "    return counts\n",
    "    \n",
    "stats = get_stats(tokens)\n",
    "# print(sorted(((v,k) for k,v in stats.items()), reverse=True))\n",
    "\n",
    "top_pair = max(stats, key=stats.get)\n",
    "print(top_pair, '\"' + chr(top_pair[0]) + '\"', '\"' + chr(top_pair[1]) + '\"')"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "9b83cb45",
   "metadata": {},
   "source": [
    "Now we want to **merge**. Specifically, we designate a new token. We replace all occurrences of the top pair with the new token."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 6,
   "id": "5f1fe690-346a-493d-9490-4d2905cb1427",
   "metadata": {},
   "outputs": [],
   "source": [
    "def merge(ids, pair, idx):\n",
    "    new_ids = []\n",
    "    i = 0\n",
    "    while i < len(ids):\n",
    "        if i < len(ids) - 1 and ids[i] == pair[0] and ids[i+1] == pair[1]:\n",
    "            new_ids.append(idx)\n",
    "            i += 2\n",
    "        else:\n",
    "            new_ids.append(ids[i])\n",
    "            i += 1\n",
    "    return new_ids"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 7,
   "id": "1b2279a3-edc7-4c52-99dd-10bff2b118cd",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[72, 101, 108, 108, 111, 44, 32, 119, 111, 114, 108, 100, 33, 32, 72, 101, 114, 256, 105, 115, 32, 115, 111, 109, 256, 101, 120, 97, 109, 112, 108, 256, 116, 101, 120, 116, 32, 116, 111, 32, 116, 101, 115, 116, 10, 116, 104, 256, 66, 80, 69, 32, 97, 108, 103, 111, 114, 105, 116, 104, 109, 46, 32, 73, 116, 32, 105, 115, 32, 110, 111, 116, 32, 118, 101, 114, 121, 32, 105, 110, 116, 101, 114, 101, 115, 116, 105, 110, 103, 44, 32, 98, 117, 116, 32, 105, 116, 32, 119, 105, 108, 108, 10, 100, 111, 32, 116, 104, 256, 106, 111, 98, 46, 10]\n",
      "119 114\n"
     ]
    }
   ],
   "source": [
    "tokens2 = merge(tokens, top_pair, 256)\n",
    "print(tokens2)\n",
    "print(len(tokens), len(tokens2))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b6ee442d-b905-482b-a701-bc988bae680a",
   "metadata": {},
   "source": [
    "Now let's put it into a loop, and run it on a longer sequence of text."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 8,
   "id": "8f1e903c-9808-468b-8775-5508e2eeb496",
   "metadata": {
    "jupyter": {
     "source_hidden": true
    }
   },
   "outputs": [],
   "source": [
    "training_text = \"\"\"Hello, world! \n",
    "Here is some example text to test\n",
    "the BPE algorithm. It is not very \n",
    "interesting, but it will do the job.\n",
    "\"\"\" \n",
    "\n",
    "# Use this one for a longer example. \n",
    "training_text = \"\"\"大阪マラソン2025大会要項\n",
    "大会名称\t大阪マラソン2025　～OSAKA MARATHON 2025～（第13回大阪マラソン）\n",
    "（英文名：Osaka Marathon 2025）\n",
    "兼 ジャパンマラソンチャンピオンシップシリーズ・男子G1/女子G2\n",
    "兼 東京2025世界陸上競技選手権大会 日本代表選手選考競技会（男子）\n",
    "主催\t大阪府、大阪市、（公財）大阪陸上競技協会\n",
    "共催\t読売新聞社、毎日新聞社、NHK、（公財）日本陸上競技連盟\n",
    "主管\t（公財）大阪陸上競技協会\n",
    "後援\t大阪市地域振興会、大阪府商店街連合会、大阪府商店街振興組合連合会、大阪市商店会総連盟、（公社）関西経済連合会、大阪商工会議所、（一社）関西経済同友会、（公財）大阪観光局、（公財）大阪府スポーツ協会、大阪府体育連合、大阪府スポーツ推進委員協議会、大阪市スポーツ協会、大阪市体育厚生協会、大阪市スポーツ推進委員協議会、（一財）大阪スポーツみどり財団、大阪府障がい者スポーツ協会、（社福）大阪市障害者福祉・スポーツ協会、（一社）大阪府医師会、（一社）大阪府病院協会、（公社）大阪府看護協会、国土交通省近畿地方整備局、国土交通省近畿運輸局、阪神高速道路（株）、（社福）読売光と愛の事業団、（特非）大阪ライフサポート協会、大阪府教育委員会、大阪市教育委員会、（株）報知新聞社、讀賣テレビ放送（株）、（株）毎日放送、（株）スポーツニッポン新聞社（予定、順不同）\n",
    "テレビ放送\tNHK、読売テレビ、毎日放送\n",
    "種目\t\n",
    "マラソン(42.195㎞）\n",
    "720<なにわ>マラソン(ランの部7.2㎞・車いすの部720m）\n",
    "開催日時\t\n",
    "2025年（令和7年）2月24日（月・振替休日）\n",
    "\n",
    "  9:15／\n",
    "マラソン第1ウェーブスタート、以降第2ウェーブ・第3ウェーブを順次スタート\n",
    "720<なにわ>マラソン（ランの部）第3ウェーブよりスタート\n",
    "10:30／\n",
    "720<なにわ>マラソン（車いすの部）スタート\n",
    "11:00／\n",
    "720<なにわ>マラソン（車いすの部）終了\n",
    "11:05／\n",
    "720<なにわ>マラソン（ランの部）終了\n",
    "16:15／\n",
    "マラソン終了\n",
    "コース\t\n",
    "マラソン\n",
    "大阪府庁前をスタートし、大阪城公園内をフィニッシュとする大阪マラソンコース（日本陸上競技連盟（日本陸連）・ワールドアスレチックス（WA）／国際マラソン・ディスタンスレース協会（AIMS）公認コース）\n",
    "720<なにわ>マラソン\n",
    "ランの部：大阪府庁前をスタートし、こども本の森中之島前をフィニッシュとするコース\n",
    "車いすの部：城見１交差点をスタートし、大阪城公園内をフィニッシュとするコース\n",
    "競技規則\t最新のWA競技規則並びに日本陸連規則及び本大会規定によります。\n",
    "なお、本大会はWA認定のラベルレースのため、WAロードレースラベリング規定が適用されます。また、WAの規則により、ドーピング検査を実施します。\n",
    "スタート方法\t\n",
    "混雑緩和と選手安全対策のためウェーブ（時間差）スタートを実施します。日本陸連登録の有無に関わらず、申込時の記録証タイム（自己ベストタイム）の申告等を参考にして、ウェーブスタート順やスタート整列ブロックを設定します。記録証タイムと予想タイムの両方が未申告の場合は、最終ウェーブの最後尾ブロックからのスタートとします。\n",
    "なお、設定されたウェーブよりも前方からスタートした場合は、失格とします。\n",
    "\n",
    "【ウェーブスタート順と整列ブロック分けの優先順位】\n",
    "\n",
    "招待選手・エリートランナー\n",
    " 自己ベストタイム（グロスタイム又はネットタイム）をお持ちの方（2021年（令和3年）8月1日以降で確認できるマラソン大会での記録であること。虚偽の場合は出場を取り消します）\n",
    "予想タイムを申告した方\n",
    "自己ベストタイム・予想タイムの両方を申告しなかった方\n",
    "グループエントリーの方は、申込時の自己ベストタイムと予想タイムが最も遅い方と同じブロックからのスタートとなります。\n",
    "720<なにわ>マラソン（ランの部）は第3ウェーブからのスタートとなります。\n",
    "制限時間\t\n",
    "マラソン：7時間（競技終了時刻16:15）\n",
    "制限時間は第1ウェーブの号砲を基準とします。\n",
    "\n",
    "720<なにわ>マラソン\n",
    "ランの部：1時間20分（競技終了時刻11:05）\n",
    "制限時間は第3ウェーブの号砲（9:45）を基準とします。\n",
    "\n",
    "車いすの部：30分（競技終了時刻11:00）\n",
    "仮装\t日本陸連登録競技者は仮装を禁止します。Aブロックにおいては日本陸連の登録の有無に関わらず、仮装を禁止します。加えて、他のランナーや沿道の方に不快感を与える服装や行為は認めません。\n",
    "参加資格\t\n",
    "マラソン：2006年（平成18年）4月1日以前に生まれた方\n",
    "①日本陸連登録競技者（2024年度の登録者）\n",
    "②日本陸連に登録していないランナー\n",
    "①②ともに単独での走行が困難な方は伴走者（1名）をつけることができます。\n",
    "（伴走者は、コース上の指定された箇所で１回に限り交代できます。なお、盲導犬等の動物の伴走は認めません。）\n",
    "競技終了時刻までに完走できる方。\n",
    "720<なにわ>マラソン\n",
    "ランの部　：2009年（平成21年）4月1日以前に生まれた方\n",
    "競技終了時刻までに完走できる方。\n",
    "車いすの部：2012年（平成24年）4月1日以前に生まれた方。\n",
    "身体障害者手帳を所持し車いすを使用される方。\n",
    "使用できる車いすは、生活用車いす及び競技用車いす（陸上競技用のレーサーは除く）とします。\n",
    "電動車いすは不可とします。\n",
    "競技終了時刻までに完走できる方。\n",
    "自力走行で競技終了時刻までに完走が困難とされる方は、介助する伴走者（1名）をつけることができます。\n",
    "定員\t34,000人（＜１＞マラソン：31,970人、＜２＞720〈なにわ〉マラソン（ランの部：2,000人、車いすの部：30人））\n",
    "申込区分\n",
    "（エリート部門を除く）\t\n",
    "マラソン\n",
    "①一般ランナー（個人）国内・国外\n",
    "②一般ランナー（グループ2人～7人）国内・国外\n",
    "①～②合計28,420人\n",
    "\n",
    "③障がい者ランナー（50人）国内\n",
    "身体障害者手帳、精神障害者保健福祉手帳、療育手帳のいずれかをお持ちの方。\n",
    "落選した場合は一般ランナーとして再抽選を行います（再エントリーは不要）。\n",
    "④市民アスリート（1,500人/先着順）国内\n",
    "年代・性別毎に設定した基準タイム以内の記録（日本陸連公認又はAIMS公認コースで2021年（令和3年）8月1日以降のグロスタイム又はネットタイム）を有する方に限ります。基準タイムは記録証に記載の大会当日満年齢によります。\n",
    "18～39歳：3時間（男子）・3時間45分（女子）、\n",
    "40～49歳：3時間10分（男子）・3時間50分（女子）、\n",
    "50～59歳：3時間25分（男子）・4時間10分（女子）、\n",
    "60～69歳：3時間50分（男子）・4時間40分（女子）、\n",
    "70歳～　：4時間30分（男子）・5時間20分（女子）\n",
    "資格審査に通らなかった場合は、事務手数料を差し引いて参加料等を返金します。一般ランナーとして参加を希望する場合は、2024年8月28日（水）17時までにご自身でエントリー手続をお願いします。\n",
    "⑤大阪スポーツ応援ランナー（大阪府・大阪市各400人、計800人/先着順）国内\n",
    "ふるさと納税制度を活用し、「大阪スポーツ応援ランナー」への寄附として、大阪府「なみはやスポーツ振興基金」又は大阪市「大阪市スポーツ振興基金」へ10万円以上寄附された方（もしくは寄附者が指名した方）に1名分の出走権を進呈します（参加料は別途必要）。\n",
    "10万円以上の寄附を頂いた方に、エントリーコードをメールでお送りします。ランネットよりエントリーコードを入力の上、大会エントリー手続を完了してください（寄附申込だけでは、大阪マラソンへ参加ができませんので注意してください）。\n",
    "⑥チャリティランナー（1,000人/先着順）国内・国外\n",
    "参加料とは別に2口以上（1口500円）の寄附を行うとともにファンドレイジングにより、別途69,000円以上の寄附を集めていただきます。\n",
    "ファンドレイジング期間内［2024年（令和6年）7月23日（火）10:00〜12月13日（金）17:00］に寄附金総額が最低寄附金額（7万円）に達しなかった場合は、不足金額を参加者ご自身に寄附いただきます。（クレジットカードへ不足額を自動的に課金します。最低寄附金額に到達しないことで、参加を取り消すことはできません）\n",
    "⑦万博チケット付きランナー（200人/先着順）国内\n",
    "大阪・関西万博2025 超早期購入割引チケット付きエントリーとなります。\n",
    "720<なにわ>マラソン\n",
    "①ランの部（2,000人）国内・国外\n",
    "②車いすの部（30人）国内\n",
    "申込方法\t\n",
    "国内\n",
    "ランネット（https://runnet.jp/runtes/）を通じて、インターネットにて申込を受け付けます。\n",
    "国外\n",
    "JTBスポーツステーション（https://jtbsports.jp）を通じて、インターネットにて申込を受け付けます。\n",
    "申込期間\t\n",
    "マラソン\n",
    "①②一般ランナー③障がい者ランナー\n",
    "2024年（令和6年）7月23日（火）10:00〜8月28日（水）17:00\n",
    "\n",
    "④市民アスリート\n",
    "2024年（令和6年）7月22日（月）10:00〜7月24日（水）17:00（先行募集・先着順）\n",
    "\n",
    "⑤大阪スポーツ応援ランナー\n",
    "寄附申込期間：2024年（令和6年）7月23日（火）10:00〜10月2日（水）17:00（先着順）\n",
    "※但し、納付書による寄附は8月16日（金）17:00まで\n",
    "エントリー期間：エントリーコード発行日〜10月16日（水）17:00\n",
    "\n",
    "⑥チャリティランナー\n",
    "2024年（令和6年）7月23日（火）10:00〜10月16日（水）17:00（先着順）\n",
    "※ファンドレイジング期間は2024年（令和6年）7月23日（火）10:00〜12月13日（金）17:00\n",
    "\n",
    "⑦万博チケット付きランナー\n",
    "2024年（令和6年）7月23日（火）10:00〜8月28日（水）17:00（先着順）\n",
    "\n",
    "720<なにわ>マラソン\n",
    "①ランの部\n",
    "2024年（令和6年）7月23日（火）10:00〜8月28日（水）17:00\n",
    "\n",
    "②車いすの部\n",
    "2024年（令和6年）9月1日（日）10:00～9月30日（月）17:00\n",
    "\n",
    "参加者決定方法\t\n",
    "マラソン\n",
    "①②一般ランナー、③障がい者ランナー\n",
    "申込者数が定員を超えた場合は抽選を行い、抽選結果は、9月26日（木）にメールで通知します。（状況により追加当選通知を行う場合があります。）\n",
    "\n",
    "当選通知の際に指定する期日までに支払手続を完了した時に、参加権利が確定します。\n",
    "\n",
    "④市民アスリート、⑤大阪スポーツ応援ランナー、⑥チャリティランナー、⑦万博チケット付きランナー\n",
    "先着順とし、定員になり次第締め切ります。\n",
    "\n",
    "但し、④市民アスリートは資格審査結果に不備のある方のみ、申込締切後概ね3週間以内にメールで通知します。\n",
    "\n",
    "720<なにわ>マラソン\n",
    "①ランの部\n",
    "申込者数が定員を超えた場合は抽選を行い、抽選結果は、9月26日（木）にメールで通知します。（状況により追加当選通知を行う場合があります）\n",
    "②車いすの部\n",
    "申込者数が定員を超えた場合は抽選を行い、抽選結果は、10月中にメールで通知します。（状況により追加当選通知を行う場合があります）\n",
    "\n",
    "当選通知の際に指定する期日までに支払手続を完了した時に、参加権利が確定します。\n",
    "\n",
    "参加料等\t\n",
    "マラソン\n",
    "①一般ランナー（個人）国内：16,000円/国外：18,000円\n",
    "②一般ランナー（グループ2～7人）参加者1名につき、国内：16,500円/国外：18,500円\n",
    "③障がい者ランナー、④市民アスリート、⑤大阪スポーツ応援ランナー 国内：16,000円\n",
    "⑥チャリティランナー　国内：16,000円　国外：18,000円\n",
    "⑦万博チケット付きランナー　国内：22,000円\n",
    "参加料とは別に事務手数料（国内：決済金額の5.5％、国外：決済金額の11%）及び参加者1人につき2口以上（1口500円）のチャリティ募金が必要です。\n",
    "720<なにわ>マラソン\n",
    "①ランの部　国内：5,000円/国外：6,000円\n",
    "\n",
    "②車いすの部　国内：1,000円\n",
    "\n",
    "参加料とは別に事務手数料（国内：決済金額の5.5％、国外：決済金額の11%※決済金額が4,000円以下の場合は一律220円）及び参加者1人につき2口以上（1口500円）のチャリティ募金が必要です。\n",
    "参加料の支払方法\t\n",
    "クレジットカード等による即時決済またはコンビニ支払い。\n",
    "\n",
    "(1)マラソンの④市民アスリートはクレジットカード等による即時決済のみ。\n",
    "\n",
    "表彰\t\n",
    "表彰は、グロスタイムにより次のとおり行います。(720<なにわ>マラソンの表彰はありません）\n",
    "\n",
    "◎総合男女の各1位～8位を表彰します。\n",
    "◎市民ランナー賞として、招待選手、エリートランナーを除く男女各1位を表彰します。\n",
    "◎シカゴマラソン賞として、招待選手、エリートランナー、連携大会の代表選手、第12回大会までの同賞受賞者を除く大阪府内在住者の男女各1位を表彰します。\n",
    "◎プラハマラソン賞として、招待選手、エリートランナー、連携大会の代表選手を除く、大阪府内在住者の男女各2位を表彰します。ただし、男女各1位がシカゴマラソン賞に該当しない場合は男女各1位を表彰します。\n",
    "別途ネットタイムにより、招待選手・エリートランナー・総合入賞者を除く年代別5歳刻みの男女各1位～3位と、最高齢で完走されたスーパーシニア賞に対して賞状を後日送付します。\n",
    "\n",
    "ランナー受付\n",
    "（大阪マラソンEXPO2025）\t\n",
    "日程／2025年（令和7年）2月22日（土）～2月23日（日・祝）\n",
    "\n",
    "場所／インテックス大阪\n",
    "\n",
    "時間／ランナー受付：2月22日（土）11:00〜19:00、2月23日（日・祝）10:00〜18:00\n",
    "展示エリア：2月22日（土）11:00〜19:30（最終入場 19:00）、2月23日（日・祝）10:00〜18:30（最終入場18:00）\n",
    "\n",
    "大会当日（2月24日(月・振替休日)）の受付は行いません。\n",
    "受付時に、本人確認を行うため、ランナー本人以外の代理受付は認めません。\n",
    "受付の際、顔写真付きの本人確認書類を持参してください。\n",
    "障がい者ランナーは申込時に登録した手帳（身体障害者手帳、精神障害者保健福祉手帳、療育手帳のいずれか）を持参してください。\n",
    "720＜なにわ＞マラソン（車いすの部）の受付は大会当日（9:00～9:40）に行います。\n",
    "ドーピング\n",
    "コントロール\tWAアンチ・ドーピング規程もしくは日本アンチ・ドーピング規程に基づいてドーピング検査を行います。大会前又は後の検査においては、尿又は血液（あるいはその両方）の採取を行います。該当者は指示に従って検査を受けてください。\n",
    "TUE申請\t\n",
    "禁止表国際基準で定められる禁止物質・禁止方法を病気の治療目的で使わざるを得ない競技者は、治療使用特例（TUE）の申請を行ってください。詳細については、下記ウェブサイトを確認してください。\n",
    "\n",
    "▼日本陸連医事委員会\n",
    "（https://www.jaaf.or.jp/about/resist/medical/）\n",
    "\n",
    "▼日本アンチ・ドーピング機構\n",
    "（https://www.playtruejapan.org/）\n",
    "禁止物質・禁止方法についてTUEが付与されている場合には、その証明書（コピーで可）をドーピング検査の際に検査員へ提出してください。\n",
    "\n",
    "参加注意事項\t\n",
    "主催者は競技中の事故は応急処置に限り対応します。主催者に重大な過失がある場合を除き、補償は加入した傷害保険（見舞金補償）の範囲内となります。\n",
    "チャリティプログラムの趣旨に賛同できない方の参加はご遠慮ください。\n",
    "参加資格を満たさない方は参加できません。申込者以外の参加は認めません。\n",
    "氏名、参加資格、記録証等の虚偽や不正が判明した場合は失格とし、参加を認めません。また今後の大会の申込も認めません。\n",
    "申込後の連絡は、登録されたメールアドレスへの配信で行います。参加者の機器等の不具合やメール設定の不備、メールアドレスの変更等により主催者からのメールを受信できなかった場合、主催者は一切の責任を負いません。\n",
    "会場へは公共の交通機関等をご利用ください。720＜なにわ＞マラソン（車いすの部）は別途案内します。但し、交通機関の遅延等により参加できなかった場合、主催者は一切の責任を負いません。\n",
    "体調管理を行った上で参加してください。\n",
    "貴重品や手荷物の紛失や盗難、損傷などについて主催者は一切責任を負いません。\n",
    "広告目的で大会会場（コース上を含む）において、企業名、商品名等を意味する図案や文字、商標等を掲出したり、身につけて表示することは認めません。\n",
    "大会の映像、写真、記事、位置情報、参加者の氏名、年齢、居住地（都道府県名又は市区町村名）、記録等のテレビ、新聞、雑誌、インターネット、主催者が発行する印刷物等への掲載権と肖像権は主催者に帰属します。\n",
    "本大会は、「AbbottWMM Wanda Age Group World Championships」の予選会となる「AbbottWMM Wanda Age Group World Rankings」の対象大会であることから、大会結果（氏名・年齢・性別・国籍・記録）をアボット・ワールドマラソンメジャーズを運営するアメリカ合衆国所在のWORLD MARATHON MAJORS LLCへ提供します。「AbbottWMM Wanda Age Group World Championships」出場の対象となったランナーにはその旨の通知がWORLD MARATHON MAJORS LLCから届きます。なお、通知を受け取るには、あらかじめご自身でWORLD MARATHON MAJORS LLCに登録する必要があります。詳細はアボット・ワールドマラソンメジャーズについてをご確認ください。\n",
    "大会の写真等について、主催者が委託する者がランナー向けに販売することがあります。\n",
    "上記のほか、大会に関する事項については主催者の指示に従ってください。\n",
    "申込注意事項\t\n",
    "ご利用の機器等（OS、ブラウザソフト、回線を含む）によって申込みできないことがあります。機器等の不具合等による申込みの遅れについて、主催者は一切の責任を負いません。\n",
    "複数の申込区分での重複申込が判明した場合はすべての申込を失格とし、参加を認めません。 同一人物による重複申込が判明した場合はすべての申込を失格とし、参加を認めません。\n",
    "申込後の自己都合による申込内容の変更、キャンセル、参加料の返金はできません。ただし、エントリーサイトのマイページで変更できる変更については、この限りではありません。\n",
    "申込状況、抽選結果等の問合せについては、一切応じられません。\n",
    "主催者は参加料等の領収書の発行は行いません。クレジットカード会社発行の利用明細書又は請求書をご利用ください。コンビニ払いの場合は、お振込みの控えを領収書に代えさせていただきます。なお、ランネットからエントリーした場合は、ランネットのマイページより領収書の発行が可能です。\n",
    "エリート部門への参加を希望する場合は、エリート部門でご応募ください。エリートランナーの募集は、別途行います。\n",
    "開催可否・中止判断\t地震、風水害などの災害、感染症拡大、警察・消防の対応が必要な事故等の発生により、安全な大会運営が困難と判断した場合には、大会を中止します。大会中止の場合の参加料等については、中止までに要した経費等を差し引いた上で返金の有無および金額を決定します。\n",
    "個人情報について\t\n",
    "主催者は個人情報の保護法令を遵守し、大阪マラソン組織委員会のプライバシーポリシーに従って、参加者の個人情報を取り扱います。プライバシーポリシーについては、下記ウェブサイトを確認してください。\n",
    "\n",
    "▼プライバシーポリシー\n",
    "https://www.osaka-marathon.com/2025/policy/\n",
    "\n",
    "主催者または大阪マラソンコールセンターから申込内容に関する確認連絡をさせていただくことがあります。\n",
    "取材に承諾いただいた参加者に限り、主催者が、共催者及び関係メディア等に個人情報を提供させていただきます。\n",
    "その他\t感染症対策について、国、大阪府、（公財）日本スポーツ協会及び（公財）日本陸上競技連盟等から方針又はガイドラインが示された場合には、それらに沿って対策を行います。\n",
    "Guide to the Osaka Marathon 2025\n",
    "Race Name\tOsaka Marathon 2025 (13th Osaka Marathon)\n",
    "also serves as: Japan Marathon Championship Series Men's G1/Women's G2\n",
    "World Athletics Championships Tokyo 2025 Marathon, Japan National Team qualifying trials (Men)\n",
    "Race Organizers\tOsaka Prefectural Government, Osaka City, Osaka Athletics (OA)\n",
    "Co-organizers\tThe Yomiuri Shimbun, The Mainichi Newspapers, Japan Broadcasting Corporation (NHK), Japan Association of Athletics Federations (JAAF)\n",
    "Managing Organization\tOsaka Athletics (OA)\n",
    "Operational Supporter\tOsaka Para Athletics Association\n",
    "Supporting Organizations\tOsaka City Community Promotion Association, Osaka Prefecture Shopping District Association, Osaka Prefectural Federation of Shopping Center Promotion Associations, Osaka City Shopping Streets Association, Kansai Economic Federation, Osaka Chamber of Commerce and Industry, Kansai Association of Corporate Executives, Osaka Convention & Tourism Bureau, Osaka Prefectural Sport Association, Osaka Athletic Association, Osaka Sport Promotion Council, Osaka City Sport Association, Osaka City Sports and Welfare Association, Osaka City Sport Promotion Council, Osaka City Sports and Greenery Association, Osaka Para-Sports Association, Osaka City Welfare and Sports Association for Persons with Disabilities, Osaka Medical Association, Osaka Hospital Association, Osaka Nursing Association, Kinki Regional Development Bureau of the Ministry of Land, Infrastructure, Transport and Tourism, Kinki District Transport Bureau of the Ministry of Land, Infrastructure, Transport and Tourism, Hanshin Expressway Company Limited, Yomiuri Light and Humanity Association, Osaka Life Support Association, Osaka Prefectural Board of Education, Osaka City Board of Education, The Hochi Shimbun Inc., Yomiuri Telecasting Corporation, Mainichi Broadcasting System, Sports Nippon Newspapers (planned, in no particular order)\n",
    "Official Sponsors\tOsaka Metro Co., Ltd.; OPTAGE Inc.; Mizuno Corporation; Duskin Co., Ltd.; Daiwa House Industry Co., Ltd.; MUFG Bank, Ltd.; others\n",
    "Telecast\tNHK, Yomiuri Telecasting Corporation, Mainichi Broadcasting System\n",
    "Event\t\n",
    "Marathon (42.195 km)\n",
    "720 <Naniwa> Marathon (Runners: 7.2 km, Wheelchairs: 720 m)\n",
    "Date & Time\t\n",
    "Monday, February 24, 2025\n",
    "9:15: Marathon Wave 1 starts. This is followed by Wave 2 and Wave 3, which start sequentially.\n",
    "720 <Naniwa> Marathon (Runners) starts with Wave 3.\n",
    "10:30: 720 <Naniwa> Marathon (Wheelchairs) starts.\n",
    "11:00: 720 <Naniwa> Marathon (Wheelchairs) finishes.\n",
    "11:05: 720 <Naniwa> Marathon (Runners) finishes.\n",
    "16:15: Marathon finishes.\n",
    "Course\t\n",
    "Marathon:\n",
    "Osaka Marathon course starting in front of the Osaka Prefectural Government Building and finishing within Osaka Castle (official course of the Japan Association of Athletics Federations (JAAF), World Athletics (WA), and the Association of International Marathons and Distance Races (AIMS))\n",
    "720 <Naniwa> Marathon :\n",
    "Runners: Course starting in front of the Osaka Prefectural Government Building and finishing in front of Nakanoshima Children’s Book Forest\n",
    "Wheelchairs: Course starting at the Shiromi 1 Intersection and finishing inside Osaka Castle Park\n",
    "Competition Rules\tThe race will be conducted in accordance with the latest rules and regulations of the WA, the JAAF, and this event.\n",
    "Because the Osaka Marathon is a Label Road Race certified by the WA, the WA Road Race Label Regulations also apply. Doping control tests will be conducted in accordance with WA Rules. The wheelchair marathon follows the rules and regulations of the World Para Athletics (WPA) and this event.\n",
    "Line-up at Start\t\n",
    "To relieve congestion and protect athletes’ safety, a wave (corral) start will be used. A wave-start schedule and waiting blocks will be set up according to the certified time on the record certificate (personal best time) submitted at the time of application, regardless of whether or not you have JAAF membership. Those who have not specified both a certified time on the record certificate and an estimated finish time will start from the rearmost block of the last wave.\n",
    "Please note that starting from a point ahead of your set wave will result in disqualification.\n",
    "\n",
    "[Priority order for wave starts and waiting blocks]\n",
    "\n",
    "Invited athletes / Elite runners\n",
    "Those who have a personal best time (gross time or net time) (The time must be recorded at a confirmable race held on or after August 1, 2021. If it is found to be a false report, the entry will be canceled.)\n",
    "Those who reported an estimated finish time\n",
    "Those who reported neither a personal best time nor an estimated finish time\n",
    "* Group entries will start from the same block as the participant in the same group with the slowest personal best time and estimated time at application.\n",
    "* 720 <Naniwa> Marathon (Runners) shall start from Wave 3.\n",
    "Time Limit\t\n",
    "Marathon: 7 hours (The competition ends at 16:15.)\n",
    "* Time limits are based on the timing of the starting gun for Wave 1.\n",
    "\n",
    "720 <Naniwa> Marathon:\n",
    "Runners: 1 hour and 20 minutes (The competition ends at 11:05.)\n",
    "* Time limits are based on the timing of the starting gun (9:45) for Wave 3.\n",
    "\n",
    "Wheelchairs: 30 minutes (The competition ends at 11:00.)\n",
    "Costume-clad Participation\tJAAF members are prohibited from participating in the race dressed in costume. Costume-clad runners are not allowed to start from Block A regardless of having JAAF membership or not. Costumes and behavior that cause discomfort to other runners or people along the route are prohibited.\n",
    "Qualifications\t\n",
    "<1> Marathon: Persons born on or before April 1, 2006\n",
    "\n",
    "JAAF members (FY2024 registered members)\n",
    "Non-JAAF members\n",
    "* In both (1) and (2), those who have difficulty running on their own may each be accompanied by one escort runner. (An escort runner can change only once at a designated point on the course. Running with a guide dog or any other animal is not permitted.)\n",
    "* Those who are capable of completing the race within the race time limits.\n",
    "<2> 720<Naniwa>Marathon（Runners）: Persons born on or before April 1, 2009\n",
    "\n",
    "Maximum Number of Participants\t34,000 (Marathon: 31,970 / 720 <Naniwa> Marathon (Runners): 2,000)\n",
    "Application Category\n",
    "(excluding elite runners)\t\n",
    "<1> Marathon:\n",
    "\n",
    "①General runners (individuals), ②General runners (2- to 7-person groups)\n",
    "① to ② total 28,420\n",
    "③Charity Runners 1,000\n",
    "<2> 720 <Naniwa> Marathon:\n",
    "\n",
    "①Runners 2,000\n",
    "How to Apply\t\n",
    "Domestic entries\n",
    "Applications are accepted on the Internet through Runnet\n",
    "(https://runnet.jp/runtes/).\n",
    "Overseas entries\n",
    "Applications are accepted on the Internet through JTB Sports Station\n",
    "(https://jtbsports.jp).\n",
    "Application Period\t\n",
    "<1> Marathon:\n",
    "\n",
    "①②General runners: From 10:00 on July 23 (Tue) to 17:00 August 28（Tue）, 2024\n",
    "③Charity Runner: From 10:00 on July 23 (Tue) to 17:00 October 16（Wed）, 2024\n",
    "* First come, first served. Applications will no longer be accepted when capacity is reached.\n",
    "\n",
    "<2> 720 <Naniwa> Marathon:\n",
    "\n",
    "④Runners: From 10:00 on July 23 (Tue) to 17:00 August 28（Tue）, 2024\n",
    "How Participants are Selected\t①②④If the number of applicants exceeds the capacity, a drawing will be held. Notification of the drawing results will be provided on September 26, 2024 (Thursday) by email.(Depending on the situation, a notification may be sent to those who are additionally selected as participants.)\n",
    "③First come, first served. Applications will no longer be accepted when capacity is reached\n",
    "Entry Fee\t\n",
    "<1> Marathon:\n",
    "\n",
    "①General runners (individuals) Domestic: 16,000 yen / Overseas: 18,000 yen\n",
    "②General runners (2- to 7-person groups), per participant Domestic: 16,500 yen / Overseas: 18,500 yen\n",
    "③Charity runners Domestic: 16,000 yen / Overseas: 18,000 yen\n",
    "<2> 720 <Naniwa> Marathon:\n",
    "\n",
    "④Runners Domestic: 5,000 yen / Overseas: 6,000 yen\n",
    "* For both marathons, a separate administrative fee (5.5% of the amount of payment for domestic entries, and 11% for overseas entries) and at least two charity donations per participant (500 yen per donation) are required in addition to the entry fee.\n",
    "Payment Method\t\n",
    "Payment must be made immediately by credit card, etc., or at a convenience store.\n",
    "\n",
    "Awards\t\n",
    "The awards will be granted based on the gross time as follows. (No awards are issued for the 720 <Naniwa> Marathon.)\n",
    "\n",
    "◎The top eight male and female finishers overall will be awarded.\n",
    "◎The Civic Runner Award is granted to male and female runners finishing in first place, excluding invited athletes and elite runners.\n",
    "◎The Chicago Marathon Award is granted to male and female runners who live in Osaka Prefecture and finish in first place, excluding invited athletes, elite runners, representative runners for affiliate marathons, and the Chicago Marathon Award winners up through the 12th Osaka Marathon.\n",
    "◎The Prague Marathon Award is granted to male and female runners who live in Osaka Prefecture and finish in second place, excluding invited athletes, elite runners, and representative runners for affiliate marathons. However, if a male or female runner finishing first place is not eligible for the Chicago Marathon Award, they shall receive the Prague Marathon Award instead.\n",
    "* A certificate will also be sent separately at a later date to the top three male and female runners per five-year age bracket, excluding invited athletes, elite runners, and overall prize winners, according to their net time. The oldest runner to complete the marathon will be sent the Super Senior Award at a later date.\n",
    "\n",
    "Runners' Registration\n",
    "(Osaka Marathon EXPO 2025)\t\n",
    "Period: Two days on February 22 (Sat) and 23 (Sun, holiday), 2025\n",
    "\n",
    "Venue: INTEX Osaka\n",
    "\n",
    "Time/Runner registration: February 22 (Sat) 11:00 to 19:00, February 23 (Sun, holiday) 10:00 to 18:00\n",
    "Exhibition area: February 22 (Sat) 11:00 to 19:30 (last admission: 19:00), February 23 (Sun, holiday) 10:00 to 18:30 (last admission: 18:00)\n",
    "\n",
    "* Registration will not be accepted on the day of the race (Monday, February 24). Please bring a photo identification document with you for registration.Your registration will not be accepted if you do not bring a photo identification document with you.\n",
    "* In order to verify the identity of the runner at the time of registration, no substitutes other than the runner himself/herself will be allowed to register on behalf of the runner.\n",
    "Doping Control\tDoping control tests will be conducted in accordance with the latest WA Anti-Doping Rules, the World Anti-Doping Code, and the Japan Anti-Doping Code. For doping control tests carried out before or after the race, a sample of urine and/or blood will be collected. The relevant people should take the test as instructed.\n",
    "TUE Application\t\n",
    "Athletes who have to use prohibited substances/methods included on the Prohibited List for the purpose of treating a disease must file an application for a Therapeutic Use Exemption (TUE). For details, please refer to the websites below.\n",
    "\n",
    "▼JAAF Medical Committee\n",
    "(https://www.jaaf.or.jp/about/resist/medical/)\n",
    "\n",
    "▼Japan Anti-Doping Agency (JADA)\n",
    "(https://www.playtruejapan.org/)\n",
    "\n",
    "If a TUE has been issued for a prohibited substance/method, submit a certificate of proof (copy acceptable) to the inspector during doping control testing.\n",
    "\n",
    "Notes for Participating in the Race\t\n",
    "The organizer will only provide first aid in the event of an accident during the race. Except in the case of gross negligence by the organizer, compensation will be within the range of the engaged accident insurance (solatium payment).\n",
    "The event organizing body will not accept applications from those who disagree with the concept of charity programs.\n",
    "Those who do not meet the qualifications cannot participate in the race. No one but the selected applicants can participate in the race.\n",
    "If any falsification of your name/qualifications/record certificate or other dishonest act is identified, you will be disqualified and you will not be allowed to participate in the race. Furthermore, applications for future races will also not be allowed.\n",
    "After application, notifications will be sent to the registered email address. The event organizing body will take no responsibility if you cannot receive email messages sent by the event organizing body due to a malfunction in your device, incorrect email settings, a change in your email address, or the like.\n",
    "Please use public transportation to reach the venue. Please note that the event organizing body will take no responsibility if you cannot participate in the race due to transportation delays.\n",
    "Please be sure to manage your health before participating in the race.\n",
    "The event organizing body will take no responsibility for any loss of, theft of, or damage to your belongings, including valuables and baggage.\n",
    "It is prohibited to post or wear any designs, letters, trademarks, or the like representing company names or product names, etc. for advertising purposes within the event site (including the race course).\n",
    "Usage and portrait rights of the following items belong to the event organizing body when they are used for TV broadcasting, newspapers, magazines, the Internet, and printed materials published by the event organizing body: images, photographs, and articles covering the race; location information; and entrants’ names, ages, addresses (prefecture or city) and records.\n",
    "This race is covered by the AbbottWMM Wanda Age Group World Rankings, which is a qualifier for the AbbottWMM Wanda Age Group World Championships. Therefore, race results (names, ages, genders, nationalities, and records) are provided to World Marathon Majors LLC, a U.S.-based company that manages the Abbott World Marathon Majors. Runners who qualify to participate in the AbbottWMM Wanda Age Group World Championships will receive a notification to that effect from World Marathon Majors LLC. To receive this notification, you yourself must first register with World Marathon Majors LLC. For more information, see the About Abbott World Marathon Majors.\n",
    "Photographs of the race may be sold to the runners by a party entrusted by the event organizing body.\n",
    "For matters relating to the race, even if they are not included in the aforementioned notes, please follow instructions given by the event organizing body.\n",
    "Notes for Applications\t\n",
    "You may be unable to apply depending on the device you are using (including OS, browser software, and network configuration). The event organizing body will take no responsibility for any delay in applications due to equipment malfunctions or other problems.\n",
    "You cannot apply for multiple application categories. If duplicate applications by the same person are found, all applications will be disqualified, and you will not be allowed to participate in the race.\n",
    "You cannot change or cancel your application for personal reasons after application. However, this does not apply to changes that can be made on My Page.\n",
    "We will not respond to any inquiries regarding application status, drawing results, and other matters.\n",
    "No receipts will be issued for the entry fee and other charges. Please use your credit card statement or bill issued by a credit card company as your receipt. In the case of payment at a convenience store, use the copy of the payment slip as the receipt. Please note that those who enter the event through Runnet can obtain a receipt from My Page in Runnet.\n",
    "If you wish to participate in the race as an elite runner, please apply in the category of “elite runners.” Applications for elite runners will be accepted separately.\n",
    "Judgment on Whether the Race Can Be Held or Will Be Canceled\tIf the safe operation of the race is judged to be difficult due to an earthquake, storm, flood, or any other natural disaster; the spread of an infectious disease; accident requiring a police or fire department response; etc., the race will be canceled. The amount of refund of the entry fee, etc. in the case of race cancellation will be determined after deducting the expenses incurred before the cancellation decision was made.\n",
    "Personal Information\t\n",
    "The event organizing body abides by laws and regulations related to the protection of personal information and handles the personal information of the participants in accordance with the privacy policy of the Osaka Marathon Organizing Committee. For the privacy policy, please refer to the website below.\n",
    "\n",
    "▼Privacy Policy\n",
    "https://www.osaka-marathon.com/2025/en/policy/\n",
    "\n",
    "* The event organizing body or the Osaka Marathon Call Center may contact participants to confirm the information stated on their application forms.\n",
    "* The event organizing body will not provide any personal information on participants to the co-organizers or related mass media without the participant's consent.\n",
    "Other Information\tIn the event that the Japanese government, Osaka Prefecture, the Japan Sport Association, the Japan Association of Athletics Federations, or other applicable body issues policies or guidelines regarding countermeasures against infectious diseases, appropriate measures will be taken in accordance with those policies or guidelines.\n",
    "\"\"\"\n",
    "# https://www.osaka-marathon.com/2025\n"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 9,
   "id": "0b19d4e9-f7ef-4430-9d9a-2b16050ba1bc",
   "metadata": {},
   "outputs": [
    {
     "data": {
      "text/plain": [
       "39298"
      ]
     },
     "execution_count": 9,
     "metadata": {},
     "output_type": "execute_result"
    }
   ],
   "source": [
    "# UTF-8 vocabulary\n",
    "tokens = training_text.encode(\"utf-8\")\n",
    "tokens = list(map(int, tokens))\n",
    "\n",
    "# Character-level vocabulary\n",
    "# tokens = list(training_text)\n",
    "len(tokens)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 10,
   "id": "e6099917-96bf-4d25-bc94-263f389a5963",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "pair: (227, 129) freq: 1457\n",
      "merging (227, 129) into a new token 256\n",
      "\n",
      "pair: (227, 131) freq: 985\n",
      "merging (227, 131) into a new token 257\n",
      "\n",
      "pair: (227, 130) freq: 709\n",
      "merging (227, 130) into a new token 258\n",
      "\n",
      "pair: (101, 32) freq: 469\n",
      "merging (101, 32) into a new token 259\n",
      "\n",
      "pair: (239, 188) freq: 384\n",
      "merging (239, 188) into a new token 260\n",
      "\n",
      "pair: (227, 128) freq: 337\n",
      "merging (227, 128) into a new token 261\n",
      "\n",
      "pair: (116, 104) freq: 287\n",
      "merging (116, 104) into a new token 262\n",
      "\n",
      "pair: (111, 110) freq: 279\n",
      "merging (111, 110) into a new token 263\n",
      "\n",
      "pair: (116, 105) freq: 253\n",
      "merging (116, 105) into a new token 264\n",
      "\n",
      "pair: (97, 110) freq: 199\n",
      "merging (97, 110) into a new token 265\n",
      "\n",
      "pair: (111, 114) freq: 193\n",
      "merging (111, 114) into a new token 266\n",
      "\n",
      "pair: (101, 114) freq: 192\n",
      "merging (101, 114) into a new token 267\n",
      "\n",
      "pair: (261, 129) freq: 187\n",
      "merging (261, 129) into a new token 268\n",
      "\n",
      "pair: (115, 32) freq: 187\n",
      "merging (115, 32) into a new token 269\n",
      "\n",
      "pair: (116, 32) freq: 186\n",
      "merging (116, 32) into a new token 270\n",
      "\n",
      "pair: (105, 110) freq: 178\n",
      "merging (105, 110) into a new token 271\n",
      "\n",
      "pair: (257, 188) freq: 176\n",
      "merging (257, 188) into a new token 272\n",
      "\n",
      "pair: (256, 174) freq: 174\n",
      "merging (256, 174) into a new token 273\n",
      "\n",
      "pair: (100, 32) freq: 174\n",
      "merging (100, 32) into a new token 274\n",
      "\n",
      "pair: (97, 114) freq: 165\n",
      "merging (97, 114) into a new token 275\n",
      "\n",
      "pair: (260, 137) freq: 163\n",
      "merging (260, 137) into a new token 276\n",
      "\n",
      "pair: (260, 136) freq: 160\n",
      "merging (260, 136) into a new token 277\n",
      "\n",
      "pair: (44, 32) freq: 157\n",
      "merging (44, 32) into a new token 278\n",
      "\n",
      "pair: (257, 179) freq: 145\n",
      "merging (257, 179) into a new token 279\n",
      "\n",
      "pair: (262, 259) freq: 143\n",
      "merging (262, 259) into a new token 280\n",
      "\n",
      "pair: (256, 171) freq: 130\n",
      "merging (256, 171) into a new token 281\n",
      "\n",
      "pair: (264, 263) freq: 126\n",
      "merging (264, 263) into a new token 282\n",
      "\n",
      "pair: (101, 110) freq: 115\n",
      "merging (101, 110) into a new token 283\n",
      "\n",
      "pair: (121, 32) freq: 114\n",
      "merging (121, 32) into a new token 284\n",
      "\n",
      "pair: (258, 146) freq: 113\n",
      "merging (258, 146) into a new token 285\n",
      "\n",
      "pair: (48, 48) freq: 113\n",
      "merging (48, 48) into a new token 286\n",
      "\n",
      "pair: (261, 130) freq: 113\n",
      "merging (261, 130) into a new token 287\n",
      "\n",
      "pair: (256, 153) freq: 112\n",
      "merging (256, 153) into a new token 288\n",
      "\n",
      "pair: (257, 169) freq: 111\n",
      "merging (257, 169) into a new token 289\n",
      "\n",
      "pair: (256, 190) freq: 109\n",
      "merging (256, 190) into a new token 290\n",
      "\n",
      "pair: (256, 175) freq: 106\n",
      "merging (256, 175) into a new token 291\n",
      "\n",
      "pair: (97, 282) freq: 106\n",
      "merging (97, 282) into a new token 292\n",
      "\n",
      "pair: (229, 164) freq: 102\n",
      "merging (229, 164) into a new token 293\n",
      "\n",
      "pair: (272, 257) freq: 101\n",
      "merging (272, 257) into a new token 294\n",
      "\n",
      "pair: (256, 132) freq: 101\n",
      "merging (256, 132) into a new token 295\n",
      "\n",
      "pair: (114, 101) freq: 101\n",
      "merging (114, 101) into a new token 296\n",
      "\n",
      "pair: (32, 280) freq: 101\n",
      "merging (32, 280) into a new token 297\n",
      "\n",
      "pair: (108, 32) freq: 95\n",
      "merging (108, 32) into a new token 298\n",
      "\n",
      "pair: (228, 184) freq: 91\n",
      "merging (228, 184) into a new token 299\n",
      "\n"
     ]
    }
   ],
   "source": [
    "# Settings\n",
    "starting_vocab_size = 256\n",
    "vocab_size = 300\n",
    "num_merges = vocab_size - starting_vocab_size\n",
    "ids = list(tokens)\n",
    "\n",
    "# Full BPE algorithm\n",
    "def get_stats(ids):\n",
    "    counts = {}\n",
    "    for pair in zip(ids, ids[1:]):\n",
    "        counts[pair] = counts.get(pair, 0) + 1\n",
    "    return counts\n",
    "\n",
    "def merge(ids, pair, idx):\n",
    "    new_ids = []\n",
    "    i = 0\n",
    "    while i < len(ids):\n",
    "        if i < len(ids) - 1 and ids[i] == pair[0] and ids[i+1] == pair[1]:\n",
    "            new_ids.append(idx)\n",
    "            i += 2\n",
    "        else:\n",
    "            new_ids.append(ids[i])\n",
    "            i += 1\n",
    "    return new_ids\n",
    "\n",
    "\n",
    "merges = {}\n",
    "for i in range(num_merges):\n",
    "    stats = get_stats(ids)\n",
    "    pair = max(stats, key=stats.get)\n",
    "    idx = starting_vocab_size + i\n",
    "    print(f\"pair: {pair} freq: {stats[pair]}\")\n",
    "    print(f\"merging {pair} into a new token {idx}\\n\")\n",
    "    ids = merge(ids, pair, idx)\n",
    "    merges[pair] = idx\n"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6743e722",
   "metadata": {},
   "source": [
    "We can see the merges occurring and the new tokens being added to the vocabulary.\n",
    "\n",
    "We can interpret each merge as \"compression\". That is, two tokens are replaced with 1, hence reducing the length of the sequence. Let us look at how the length changed, and the compression ratio."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 11,
   "id": "8ab7cfa0-ed1b-4b52-b3a5-f3c7cd2c9ec3",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "original tokens length: 39298\n",
      "new tokens length: 29324\n",
      "compression ratio: 1.34X\n"
     ]
    }
   ],
   "source": [
    "print(\"original tokens length:\", len(tokens))\n",
    "print(\"new tokens length:\", len(ids))\n",
    "print(f\"compression ratio: {len(tokens) / len(ids):.2f}X\")"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "a66a6cc9-cbd4-4fed-8fbc-ac0d8d4791a2",
   "metadata": {},
   "source": [
    "### Tiktoken\n",
    "\n",
    "Now let's look at some tools and tokenizers in practice.\n",
    "\n",
    "OpenAI released a tokenization library called Tiktoken that can be used for **inference**: i.e., given a trained tokenizer, encode or decode sequences with it.\n",
    "\n",
    "The library cannot be used for **training** a tokenizer, i.e. running BPE on a corpus."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 12,
   "id": "06b40e23-691d-43aa-aecf-cb5466f30f4b",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[15496, 11, 23294, 241, 22174, 28618, 2515, 94, 31676]\n",
      "[9906, 11, 220, 90115]\n"
     ]
    }
   ],
   "source": [
    "# !pip install tiktoken\n",
    "import tiktoken\n",
    "\n",
    "enc = tiktoken.get_encoding(\"gpt2\")\n",
    "print(enc.encode(\"Hello, こんにちは\"))\n",
    "\n",
    "enc = tiktoken.get_encoding(\"cl100k_base\")\n",
    "print(enc.encode(\"Hello, こんにちは\"))"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "f2894a0f-102d-4735-9827-5fd726586231",
   "metadata": {},
   "source": [
    "### GPT-2 Tokenizer\n",
    "\n",
    "OpenAI also released their tokenization inference code for GPT-2.\n",
    "\n",
    "https://github.com/openai/gpt-2/blob/master/src/encoder.py\n",
    "\n",
    "You can take a look; it has code for encoding and decoding given the trained GPT-2 tokenizer.\n",
    "\n",
    "They also have some interesting hacks and implementation details. You can watch Andrej Karpathy's video for a fun exploration of these details."
   ]
  },
  {
   "cell_type": "markdown",
   "id": "91422236-5a20-492a-adaa-70e965820119",
   "metadata": {},
   "source": [
    "### Sentencepiece\n",
    "\n",
    "Another common library is Sentencepiece:\n",
    "\n",
    "https://github.com/google/sentencepiece\n",
    "\n",
    "It can do training and inference. However, by default it does BPE on Unicode characters, rather than UTF-8 byte strings. Hence, a tokenizer trained with the default configuration could have \"out of vocabulary\" tokens; i.e., the tokenizer is not expressive. \n",
    "\n",
    "To solve this, you can set an option `byte_fallback=True` in order for the tokenizer to encode an out of vocabulary token as a UTF-8 byte sequence.\n",
    "\n",
    "Let us **train** a tokenizer. We put the mixed English/Japanese text from before into `toy_tokenizer_txt.txt`, and run Sentencepiece."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 13,
   "id": "6c3b973c-c7bd-4abd-b3e7-713b0cd5a3a9",
   "metadata": {},
   "outputs": [],
   "source": [
    "import sentencepiece as spm\n",
    "\n",
    "with open(\"toy_tokenizer_txt.txt\", \"w\", encoding=\"utf-8\") as f:\n",
    "    f.write(training_text)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "602e9326-790f-4f4a-b9fb-222cfbb62e08",
   "metadata": {},
   "outputs": [],
   "source": [
    "import os\n",
    "\n",
    "# There are a lot of sentencepiece options.\n",
    "# If you are considering training your own tokenizer for a serious project\n",
    "# or production, make sure you are aware of the exact setting you are running.\n",
    "# Here we just make sure that we have set the vocab size and byte_fallback.\n",
    "# For fun, we also split digits to show the use of such a rule.\n",
    "options = dict(\n",
    "  # input spec\n",
    "  input=\"toy_tokenizer_txt.txt\",\n",
    "  input_format=\"text\",\n",
    "  model_prefix=\"toy_tokenizer\",\n",
    "  model_type=\"bpe\",\n",
    "  vocab_size=2048,\n",
    "  byte_fallback=True,\n",
    "  split_digits=True,\n",
    "  num_threads=os.cpu_count(),\n",
    ")\n",
    "\n",
    "spm.SentencePieceTrainer.train(**options)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 15,
   "id": "2ec3ddf8-bfdd-476b-aed1-6777ed667df6",
   "metadata": {
    "scrolled": true
   },
   "outputs": [
    {
     "data": {
      "text/plain": [
       "[['<unk>', 0],\n",
       " ['<s>', 1],\n",
       " ['</s>', 2],\n",
       " ['<0x00>', 3],\n",
       " ['<0x01>', 4],\n",
       " ['<0x02>', 5],\n",
       " ['<0x03>', 6],\n",
       " ['<0x04>', 7],\n",
       " ['<0x05>', 8],\n",
       " ['<0x06>', 9],\n",
       " ['<0x07>', 10],\n",
       " ['<0x08>', 11],\n",
       " ['<0x09>', 12],\n",
       " ['<0x0A>', 13],\n",
       " ['<0x0B>', 14],\n",
       " ['<0x0C>', 15],\n",
       " ['<0x0D>', 16],\n",
       " ['<0x0E>', 17],\n",
       " ['<0x0F>', 18],\n",
       " ['<0x10>', 19],\n",
       " ['<0x11>', 20],\n",
       " ['<0x12>', 21],\n",
       " ['<0x13>', 22],\n",
       " ['<0x14>', 23],\n",
       " ['<0x15>', 24],\n",
       " ['<0x16>', 25],\n",
       " ['<0x17>', 26],\n",
       " ['<0x18>', 27],\n",
       " ['<0x19>', 28],\n",
       " ['<0x1A>', 29],\n",
       " ['<0x1B>', 30],\n",
       " ['<0x1C>', 31],\n",
       " ['<0x1D>', 32],\n",
       " ['<0x1E>', 33],\n",
       " ['<0x1F>', 34],\n",
       " ['<0x20>', 35],\n",
       " ['<0x21>', 36],\n",
       " ['<0x22>', 37],\n",
       " ['<0x23>', 38],\n",
       " ['<0x24>', 39],\n",
       " ['<0x25>', 40],\n",
       " ['<0x26>', 41],\n",
       " ['<0x27>', 42],\n",
       " ['<0x28>', 43],\n",
       " ['<0x29>', 44],\n",
       " ['<0x2A>', 45],\n",
       " ['<0x2B>', 46],\n",
       " ['<0x2C>', 47],\n",
       " ['<0x2D>', 48],\n",
       " ['<0x2E>', 49],\n",
       " ['<0x2F>', 50],\n",
       " ['<0x30>', 51],\n",
       " ['<0x31>', 52],\n",
       " ['<0x32>', 53],\n",
       " ['<0x33>', 54],\n",
       " ['<0x34>', 55],\n",
       " ['<0x35>', 56],\n",
       " ['<0x36>', 57],\n",
       " ['<0x37>', 58],\n",
       " ['<0x38>', 59],\n",
       " ['<0x39>', 60],\n",
       " ['<0x3A>', 61],\n",
       " ['<0x3B>', 62],\n",
       " ['<0x3C>', 63],\n",
       " ['<0x3D>', 64],\n",
       " ['<0x3E>', 65],\n",
       " ['<0x3F>', 66],\n",
       " ['<0x40>', 67],\n",
       " ['<0x41>', 68],\n",
       " ['<0x42>', 69],\n",
       " ['<0x43>', 70],\n",
       " ['<0x44>', 71],\n",
       " ['<0x45>', 72],\n",
       " ['<0x46>', 73],\n",
       " ['<0x47>', 74],\n",
       " ['<0x48>', 75],\n",
       " ['<0x49>', 76],\n",
       " ['<0x4A>', 77],\n",
       " ['<0x4B>', 78],\n",
       " ['<0x4C>', 79],\n",
       " ['<0x4D>', 80],\n",
       " ['<0x4E>', 81],\n",
       " ['<0x4F>', 82],\n",
       " ['<0x50>', 83],\n",
       " ['<0x51>', 84],\n",
       " ['<0x52>', 85],\n",
       " ['<0x53>', 86],\n",
       " ['<0x54>', 87],\n",
       " ['<0x55>', 88],\n",
       " ['<0x56>', 89],\n",
       " ['<0x57>', 90],\n",
       " ['<0x58>', 91],\n",
       " ['<0x59>', 92],\n",
       " ['<0x5A>', 93],\n",
       " ['<0x5B>', 94],\n",
       " ['<0x5C>', 95],\n",
       " ['<0x5D>', 96],\n",
       " ['<0x5E>', 97],\n",
       " ['<0x5F>', 98],\n",
       " ['<0x60>', 99],\n",
       " ['<0x61>', 100],\n",
       " ['<0x62>', 101],\n",
       " ['<0x63>', 102],\n",
       " ['<0x64>', 103],\n",
       " ['<0x65>', 104],\n",
       " ['<0x66>', 105],\n",
       " ['<0x67>', 106],\n",
       " ['<0x68>', 107],\n",
       " ['<0x69>', 108],\n",
       " ['<0x6A>', 109],\n",
       " ['<0x6B>', 110],\n",
       " ['<0x6C>', 111],\n",
       " ['<0x6D>', 112],\n",
       " ['<0x6E>', 113],\n",
       " ['<0x6F>', 114],\n",
       " ['<0x70>', 115],\n",
       " ['<0x71>', 116],\n",
       " ['<0x72>', 117],\n",
       " ['<0x73>', 118],\n",
       " ['<0x74>', 119],\n",
       " ['<0x75>', 120],\n",
       " ['<0x76>', 121],\n",
       " ['<0x77>', 122],\n",
       " ['<0x78>', 123],\n",
       " ['<0x79>', 124],\n",
       " ['<0x7A>', 125],\n",
       " ['<0x7B>', 126],\n",
       " ['<0x7C>', 127],\n",
       " ['<0x7D>', 128],\n",
       " ['<0x7E>', 129],\n",
       " ['<0x7F>', 130],\n",
       " ['<0x80>', 131],\n",
       " ['<0x81>', 132],\n",
       " ['<0x82>', 133],\n",
       " ['<0x83>', 134],\n",
       " ['<0x84>', 135],\n",
       " ['<0x85>', 136],\n",
       " ['<0x86>', 137],\n",
       " ['<0x87>', 138],\n",
       " ['<0x88>', 139],\n",
       " ['<0x89>', 140],\n",
       " ['<0x8A>', 141],\n",
       " ['<0x8B>', 142],\n",
       " ['<0x8C>', 143],\n",
       " ['<0x8D>', 144],\n",
       " ['<0x8E>', 145],\n",
       " ['<0x8F>', 146],\n",
       " ['<0x90>', 147],\n",
       " ['<0x91>', 148],\n",
       " ['<0x92>', 149],\n",
       " ['<0x93>', 150],\n",
       " ['<0x94>', 151],\n",
       " ['<0x95>', 152],\n",
       " ['<0x96>', 153],\n",
       " ['<0x97>', 154],\n",
       " ['<0x98>', 155],\n",
       " ['<0x99>', 156],\n",
       " ['<0x9A>', 157],\n",
       " ['<0x9B>', 158],\n",
       " ['<0x9C>', 159],\n",
       " ['<0x9D>', 160],\n",
       " ['<0x9E>', 161],\n",
       " ['<0x9F>', 162],\n",
       " ['<0xA0>', 163],\n",
       " ['<0xA1>', 164],\n",
       " ['<0xA2>', 165],\n",
       " ['<0xA3>', 166],\n",
       " ['<0xA4>', 167],\n",
       " ['<0xA5>', 168],\n",
       " ['<0xA6>', 169],\n",
       " ['<0xA7>', 170],\n",
       " ['<0xA8>', 171],\n",
       " ['<0xA9>', 172],\n",
       " ['<0xAA>', 173],\n",
       " ['<0xAB>', 174],\n",
       " ['<0xAC>', 175],\n",
       " ['<0xAD>', 176],\n",
       " ['<0xAE>', 177],\n",
       " ['<0xAF>', 178],\n",
       " ['<0xB0>', 179],\n",
       " ['<0xB1>', 180],\n",
       " ['<0xB2>', 181],\n",
       " ['<0xB3>', 182],\n",
       " ['<0xB4>', 183],\n",
       " ['<0xB5>', 184],\n",
       " ['<0xB6>', 185],\n",
       " ['<0xB7>', 186],\n",
       " ['<0xB8>', 187],\n",
       " ['<0xB9>', 188],\n",
       " ['<0xBA>', 189],\n",
       " ['<0xBB>', 190],\n",
       " ['<0xBC>', 191],\n",
       " ['<0xBD>', 192],\n",
       " ['<0xBE>', 193],\n",
       " ['<0xBF>', 194],\n",
       " ['<0xC0>', 195],\n",
       " ['<0xC1>', 196],\n",
       " ['<0xC2>', 197],\n",
       " ['<0xC3>', 198],\n",
       " ['<0xC4>', 199],\n",
       " ['<0xC5>', 200],\n",
       " ['<0xC6>', 201],\n",
       " ['<0xC7>', 202],\n",
       " ['<0xC8>', 203],\n",
       " ['<0xC9>', 204],\n",
       " ['<0xCA>', 205],\n",
       " ['<0xCB>', 206],\n",
       " ['<0xCC>', 207],\n",
       " ['<0xCD>', 208],\n",
       " ['<0xCE>', 209],\n",
       " ['<0xCF>', 210],\n",
       " ['<0xD0>', 211],\n",
       " ['<0xD1>', 212],\n",
       " ['<0xD2>', 213],\n",
       " ['<0xD3>', 214],\n",
       " ['<0xD4>', 215],\n",
       " ['<0xD5>', 216],\n",
       " ['<0xD6>', 217],\n",
       " ['<0xD7>', 218],\n",
       " ['<0xD8>', 219],\n",
       " ['<0xD9>', 220],\n",
       " ['<0xDA>', 221],\n",
       " ['<0xDB>', 222],\n",
       " ['<0xDC>', 223],\n",
       " ['<0xDD>', 224],\n",
       " ['<0xDE>', 225],\n",
       " ['<0xDF>', 226],\n",
       " ['<0xE0>', 227],\n",
       " ['<0xE1>', 228],\n",
       " ['<0xE2>', 229],\n",
       " ['<0xE3>', 230],\n",
       " ['<0xE4>', 231],\n",
       " ['<0xE5>', 232],\n",
       " ['<0xE6>', 233],\n",
       " ['<0xE7>', 234],\n",
       " ['<0xE8>', 235],\n",
       " ['<0xE9>', 236],\n",
       " ['<0xEA>', 237],\n",
       " ['<0xEB>', 238],\n",
       " ['<0xEC>', 239],\n",
       " ['<0xED>', 240],\n",
       " ['<0xEE>', 241],\n",
       " ['<0xEF>', 242],\n",
       " ['<0xF0>', 243],\n",
       " ['<0xF1>', 244],\n",
       " ['<0xF2>', 245],\n",
       " ['<0xF3>', 246],\n",
       " ['<0xF4>', 247],\n",
       " ['<0xF5>', 248],\n",
       " ['<0xF6>', 249],\n",
       " ['<0xF7>', 250],\n",
       " ['<0xF8>', 251],\n",
       " ['<0xF9>', 252],\n",
       " ['<0xFA>', 253],\n",
       " ['<0xFB>', 254],\n",
       " ['<0xFC>', 255],\n",
       " ['<0xFD>', 256],\n",
       " ['<0xFE>', 257],\n",
       " ['<0xFF>', 258],\n",
       " ['th', 259],\n",
       " ['on', 260],\n",
       " ['ti', 261],\n",
       " ['▁a', 262],\n",
       " ['or', 263],\n",
       " ['er', 264],\n",
       " ['in', 265],\n",
       " ['▁th', 266],\n",
       " ['▁the', 267],\n",
       " ['ar', 268],\n",
       " ['an', 269],\n",
       " ['re', 270],\n",
       " ['tion', 271],\n",
       " ['te', 272],\n",
       " ['en', 273],\n",
       " ['▁b', 274],\n",
       " ['ation', 275],\n",
       " ['st', 276],\n",
       " ['▁w', 277],\n",
       " ['▁c', 278],\n",
       " ['▁f', 279],\n",
       " ['se', 280],\n",
       " ['▁p', 281],\n",
       " ['le', 282],\n",
       " ['ce', 283],\n",
       " ['▁A', 284],\n",
       " ['li', 285],\n",
       " ['un', 286],\n",
       " ['▁o', 287],\n",
       " ['▁(', 288],\n",
       " ['▁M', 289],\n",
       " ['ers', 290],\n",
       " ['▁an', 291],\n",
       " ['▁t', 292],\n",
       " ['ます', 293],\n",
       " ['ing', 294],\n",
       " ['ll', 295],\n",
       " ['ro', 296],\n",
       " ['ent', 297],\n",
       " ['▁r', 298],\n",
       " ['▁in', 299],\n",
       " ['▁wi', 300],\n",
       " ['me', 301],\n",
       " ['sa', 302],\n",
       " ['ara', 303],\n",
       " ['▁of', 304],\n",
       " ['▁re', 305],\n",
       " ['ci', 306],\n",
       " ['▁O', 307],\n",
       " ['▁to', 308],\n",
       " ['▁C', 309],\n",
       " ['ラン', 310],\n",
       " ['▁be', 311],\n",
       " ['unn', 312],\n",
       " ['▁and', 313],\n",
       " ['▁e', 314],\n",
       " ['thon', 315],\n",
       " ['arathon', 316],\n",
       " ['大阪', 317],\n",
       " ['▁d', 318],\n",
       " ['is', 319],\n",
       " ['▁T', 320],\n",
       " ['ed', 321],\n",
       " ['▁n', 322],\n",
       " ['ka', 323],\n",
       " ['▁or', 324],\n",
       " ['saka', 325],\n",
       " ['al', 326],\n",
       " ['ou', 327],\n",
       " ['ソン', 328],\n",
       " ['マラ', 329],\n",
       " ['ted', 330],\n",
       " ['マラソン', 331],\n",
       " ['▁Marathon', 332],\n",
       " ['▁P', 333],\n",
       " ['▁Osaka', 334],\n",
       " ['▁W', 335],\n",
       " ['▁will', 336],\n",
       " ['he', 337],\n",
       " ['pp', 338],\n",
       " ['ナー', 339],\n",
       " ['ランナー', 340],\n",
       " ['ma', 341],\n",
       " ['cation', 342],\n",
       " ['参加', 343],\n",
       " ['ani', 344],\n",
       " ['unners', 345],\n",
       " ['ace', 346],\n",
       " ['ot', 347],\n",
       " ['します', 348],\n",
       " ['fi', 349],\n",
       " ['it', 350],\n",
       " ['ss', 351],\n",
       " ['ート', 352],\n",
       " ['ve', 353],\n",
       " ['▁y', 354],\n",
       " ['▁on', 355],\n",
       " ['hi', 356],\n",
       " ['▁S', 357],\n",
       " ['▁R', 358],\n",
       " ['申込', 359],\n",
       " ['es', 360],\n",
       " ['場合', 361],\n",
       " ['arti', 362],\n",
       " ['ts', 363],\n",
       " ['ate', 364],\n",
       " ['▁for', 365],\n",
       " ['ho', 366],\n",
       " ['▁I', 367],\n",
       " ['▁s', 368],\n",
       " ['pan', 369],\n",
       " ['por', 370],\n",
       " ['ppli', 371],\n",
       " ['ng', 372],\n",
       " ['om', 373],\n",
       " ['▁F', 374],\n",
       " ['の部', 375],\n",
       " ['大会', 376],\n",
       " ['▁race', 377],\n",
       " ['ld', 378],\n",
       " ['ct', 379],\n",
       " ['ha', 380],\n",
       " ['od', 381],\n",
       " ['oci', 382],\n",
       " ['our', 383],\n",
       " ['▁Ass', 384],\n",
       " ['▁Associ', 385],\n",
       " ['▁runners', 386],\n",
       " ['pplication', 387],\n",
       " ['▁Association', 388],\n",
       " ['ri', 389],\n",
       " ['せん', 390],\n",
       " ['競技', 391],\n",
       " ['▁st', 392],\n",
       " ['ません', 393],\n",
       " ['▁<', 394],\n",
       " ['ット', 395],\n",
       " ['ase', 396],\n",
       " ['aniz', 397],\n",
       " ['ganiz', 398],\n",
       " ['artici', 399],\n",
       " ['▁J', 400],\n",
       " ['▁L', 401],\n",
       " ['スタ', 402],\n",
       " ['ard', 403],\n",
       " ['▁ac', 404],\n",
       " ['▁ev', 405],\n",
       " ['▁is', 406],\n",
       " ['▁ti', 407],\n",
       " ['ay', 408],\n",
       " ['ir', 409],\n",
       " ['によ', 410],\n",
       " ['国内', 411],\n",
       " ['ist', 412],\n",
       " ['▁ma', 413],\n",
       " ['ting', 414],\n",
       " ['▁event', 415],\n",
       " ['ur', 416],\n",
       " ['▁m', 417],\n",
       " ['いす', 418],\n",
       " ['いて', 419],\n",
       " ['する', 420],\n",
       " ['イム', 421],\n",
       " ['tic', 422],\n",
       " ['ver', 423],\n",
       " ['▁In', 424],\n",
       " ['場合は', 425],\n",
       " ['大阪府', 426],\n",
       " ['車いす', 427],\n",
       " ['▁are', 428],\n",
       " ['▁not', 429],\n",
       " ['▁time', 430],\n",
       " ['▁with', 431],\n",
       " ['et', 432],\n",
       " ['ge', 433],\n",
       " ['id', 434],\n",
       " ['ow', 435],\n",
       " ['ps', 436],\n",
       " ['▁g', 437],\n",
       " ['、(', 438],\n",
       " ['して', 439],\n",
       " ['でき', 440],\n",
       " ['なに', 441],\n",
       " ['につ', 442],\n",
       " ['スポ', 443],\n",
       " ['ーツ', 444],\n",
       " ['時間', 445],\n",
       " ['cor', 446],\n",
       " ['ity', 447],\n",
       " ['rom', 448],\n",
       " ['tes', 449],\n",
       " ['なにわ', 450],\n",
       " ['ment', 451],\n",
       " ['thle', 452],\n",
       " ['▁The', 453],\n",
       " ['スポーツ', 454],\n",
       " ['▁partici', 455],\n",
       " ['as', 456],\n",
       " ['fe', 457],\n",
       " ['▁B', 458],\n",
       " ['▁D', 459],\n",
       " ['▁N', 460],\n",
       " ['くだ', 461],\n",
       " ['さい', 462],\n",
       " ['した', 463],\n",
       " ['主催', 464],\n",
       " ['日本', 465],\n",
       " ['for', 466],\n",
       " ['ons', 467],\n",
       " ['▁no', 468],\n",
       " ['tifi', 469],\n",
       " ['ください', 470],\n",
       " ['スタート', 471],\n",
       " ['at', 472],\n",
       " ['ue', 473],\n",
       " ['ul', 474],\n",
       " ['wa', 475],\n",
       " ['rou', 476],\n",
       " ['▁at', 477],\n",
       " ['▁by', 478],\n",
       " ['ります', 479],\n",
       " ['タイム', 480],\n",
       " ['リート', 481],\n",
       " ['主催者', 482],\n",
       " ['ance', 483],\n",
       " ['orld', 484],\n",
       " ['▁con', 485],\n",
       " ['▁pro', 486],\n",
       " ['▁reg', 487],\n",
       " ['▁World', 488],\n",
       " ['▁organiz', 489],\n",
       " ['el', 490],\n",
       " ['op', 491],\n",
       " ['qu', 492],\n",
       " ['ra', 493],\n",
       " ['▁G', 494],\n",
       " ['を行', 495],\n",
       " ['ウェ', 496],\n",
       " ['ール', 497],\n",
       " ['ody', 498],\n",
       " ['ure', 499],\n",
       " ['Nani', 500],\n",
       " ['inis', 501],\n",
       " ['ther', 502],\n",
       " ['▁who', 503],\n",
       " ['▁body', 504],\n",
       " ['Naniwa', 505],\n",
       " ['▁finis', 506],\n",
       " ['▁application', 507],\n",
       " ['tt', 508],\n",
       " ['▁*', 509],\n",
       " ['令和', 510],\n",
       " ['寄附', 511],\n",
       " ['ter', 512],\n",
       " ['cord', 513],\n",
       " ['について', 514],\n",
       " ['車いすの部', 515],\n",
       " ['▁organizing', 516],\n",
       " ['bi', 517],\n",
       " ['cl', 518],\n",
       " ['mb', 519],\n",
       " ['ru', 520],\n",
       " ['ud', 521],\n",
       " ['ww', 522],\n",
       " ['▁E', 523],\n",
       " ['リー', 524],\n",
       " ['ント', 525],\n",
       " ['ーブ', 526],\n",
       " ['協会', 527],\n",
       " ['ave', 528],\n",
       " ['mes', 529],\n",
       " ['日本陸', 530],\n",
       " ['apan', 531],\n",
       " ['fect', 532],\n",
       " ['▁ent', 533],\n",
       " ['▁you', 534],\n",
       " ['ウェーブ', 535],\n",
       " ['ランの部', 536],\n",
       " ['erson', 537],\n",
       " ['),', 538],\n",
       " ['ad', 539],\n",
       " ['ly', 540],\n",
       " ['sp', 541],\n",
       " ['ty', 542],\n",
       " ['コー', 543],\n",
       " ['ング', 544],\n",
       " ['国外', 545],\n",
       " ['通知', 546],\n",
       " ['cep', 547],\n",
       " ['▁se', 548],\n",
       " ['エント', 549],\n",
       " ['大阪市', 550],\n",
       " ['port', 551],\n",
       " ['ward', 552],\n",
       " ['lease', 553],\n",
       " ['unner', 554],\n",
       " ['エントリー', 555],\n",
       " ['refect', 556],\n",
       " ['▁Athle', 557],\n",
       " ['▁finish', 558],\n",
       " ['//', 559],\n",
       " ['▁H', 560],\n",
       " ['▁l', 561],\n",
       " ['から', 562],\n",
       " ['こと', 563],\n",
       " ['され', 564],\n",
       " ['まで', 565],\n",
       " ['終了', 566],\n",
       " ['選手', 567],\n",
       " ['://', 568],\n",
       " ['ali', 569],\n",
       " ['htt', 570],\n",
       " ['ran', 571],\n",
       " ['▁fe', 572],\n",
       " ['▁申込', 573],\n",
       " ['により', 574],\n",
       " ['ネット', 575],\n",
       " ['参加料', 576],\n",
       " ['clud', 577],\n",
       " ['llow', 578],\n",
       " ['https', 579],\n",
       " ['▁star', 580],\n",
       " ['▁your', 581],\n",
       " ['▁マラソン', 582],\n",
       " ['▁Prefect', 583],\n",
       " ['▁Athletic', 584],\n",
       " ['AA', 585],\n",
       " ['ab', 586],\n",
       " ['cy', 587],\n",
       " ['eb', 588],\n",
       " ['ic', 589],\n",
       " ['ry', 590],\n",
       " ['▁u', 591],\n",
       " ['いた', 592],\n",
       " ['先着', 593],\n",
       " ['又は', 594],\n",
       " ['登録', 595],\n",
       " ['認め', 596],\n",
       " ['金額', 597],\n",
       " ['AAF', 598],\n",
       " ['and', 599],\n",
       " ['mit', 600],\n",
       " ['ust', 601],\n",
       " ['▁Co', 602],\n",
       " ['▁Ma', 603],\n",
       " ['▁ad', 604],\n",
       " ['▁me', 605],\n",
       " ['います', 606],\n",
       " ['メール', 607],\n",
       " ['先着順', 608],\n",
       " ['参加者', 609],\n",
       " ['irst', 610],\n",
       " ['lite', 611],\n",
       " ['▁can', 612],\n",
       " ['▁com', 613],\n",
       " ['▁yen', 614],\n",
       " ['エリート', 615],\n",
       " ['日本陸連', 616],\n",
       " ['oping', 617],\n",
       " ['▁City', 618],\n",
       " ['▁from', 619],\n",
       " ['認めません', 620],\n",
       " ['ations', 621],\n",
       " ['▁Japan', 622],\n",
       " ['▁accep', 623],\n",
       " ['▁other', 624],\n",
       " ['Runners', 625],\n",
       " ['ersonal', 626],\n",
       " ['▁particip', 627],\n",
       " ['▁Athletics', 628],\n",
       " ['▁Application', 629],\n",
       " ['ak', 630],\n",
       " ['ap', 631],\n",
       " ['bb', 632],\n",
       " ['if', 633],\n",
       " ['mp', 634],\n",
       " ['oc', 635],\n",
       " ['vi', 636],\n",
       " ['▁◎', 637],\n",
       " ['があ', 638],\n",
       " ['チャ', 639],\n",
       " ['一般', 640],\n",
       " ['手帳', 641],\n",
       " ['抽選', 642],\n",
       " ['等の', 643],\n",
       " ['表彰', 644],\n",
       " ['記録', 645],\n",
       " ['The', 646],\n",
       " ['cei', 647],\n",
       " ['ges', 648],\n",
       " ['les', 649],\n",
       " ['oli', 650],\n",
       " ['ott', 651],\n",
       " ['▁as', 652],\n",
       " ['▁co', 653],\n",
       " ['までに', 654],\n",
       " ['コース', 655],\n",
       " ['cted', 656],\n",
       " ['form', 657],\n",
       " ['roup', 658],\n",
       " ['▁any', 659],\n",
       " ['▁cer', 660],\n",
       " ['bbott', 661],\n",
       " ['ports', 662],\n",
       " ['quali', 663],\n",
       " ['▁that', 664],\n",
       " ['▁Award', 665],\n",
       " ['▁recei', 666],\n",
       " ['してください', 667],\n",
       " ['一般ランナー', 668],\n",
       " ['大阪マラソン', 669],\n",
       " ['cluding', 670],\n",
       " ['▁regist', 671],\n",
       " ['▁runner', 672],\n",
       " ['▁certifi', 673],\n",
       " ['formation', 674],\n",
       " ['▁participan', 675],\n",
       " [').', 676],\n",
       " ['.)', 677],\n",
       " ['ch', 678],\n",
       " ['gh', 679],\n",
       " ['ub', 680],\n",
       " ['▁Y', 681],\n",
       " ['った', 682],\n",
       " ['ない', 683],\n",
       " ['アス', 684],\n",
       " ['ック', 685],\n",
       " ['ティ', 686],\n",
       " ['ドー', 687],\n",
       " ['公財', 688],\n",
       " ['受付', 689],\n",
       " ['市民', 690],\n",
       " ['時刻', 691],\n",
       " ['決済', 692],\n",
       " ['男女', 693],\n",
       " ['男子', 694],\n",
       " ['確認', 695],\n",
       " ['禁止', 696],\n",
       " ['ary', 697],\n",
       " ['ast', 698],\n",
       " ['ose', 699],\n",
       " ['res', 700],\n",
       " ['▁MA', 701],\n",
       " ['▁ch', 702],\n",
       " ['▁fo', 703],\n",
       " ['▁ha', 704],\n",
       " ['▁li', 705],\n",
       " ['チャリ', 706],\n",
       " ['ドーピ', 707],\n",
       " ['上競技', 708],\n",
       " ['ebru', 709],\n",
       " ['fore', 710],\n",
       " ['hibi', 711],\n",
       " ['ural', 712],\n",
       " ['▁Cha', 713],\n",
       " ['▁For', 714],\n",
       " ['主催者は', 715],\n",
       " ['競技終了', 716],\n",
       " ['ourse', 717],\n",
       " ['▁Race', 718],\n",
       " ['▁Wave', 719],\n",
       " ['があります', 720],\n",
       " ['チャリティ', 721],\n",
       " ['ドーピング', 722],\n",
       " ['▁Febru', 723],\n",
       " ['▁Group', 724],\n",
       " ['▁elite', 725],\n",
       " ['▁first', 726],\n",
       " ['競技終了時刻', 727],\n",
       " ['▁Please', 728],\n",
       " ['▁accord', 729],\n",
       " ['▁February', 730],\n",
       " ['▁accepted', 731],\n",
       " ['▁personal', 732],\n",
       " ['tification', 733],\n",
       " ['▁participate', 734],\n",
       " [')、', 735],\n",
       " [')・', 736],\n",
       " ['.,', 737],\n",
       " ['LC', 738],\n",
       " ['MM', 739],\n",
       " ['OR', 740],\n",
       " ['UE', 741],\n",
       " ['WA', 742],\n",
       " ['ag', 743],\n",
       " ['di', 744],\n",
       " ['du', 745],\n",
       " ['il', 746],\n",
       " ['jp', 747],\n",
       " ['si', 748],\n",
       " ['▁/', 749],\n",
       " ['▁▼', 750],\n",
       " ['い者', 751],\n",
       " ['って', 752],\n",
       " ['とな', 753],\n",
       " ['る方', 754],\n",
       " ['を除', 755],\n",
       " ['シー', 756],\n",
       " ['ドレ', 757],\n",
       " ['付き', 758],\n",
       " ['以上', 759],\n",
       " ['個人', 760],\n",
       " ['及び', 761],\n",
       " ['女子', 762],\n",
       " ['委員', 763],\n",
       " ['害者', 764],\n",
       " ['方法', 765],\n",
       " ['検査', 766],\n",
       " ['等を', 767],\n",
       " ['結果', 768],\n",
       " ['障が', 769],\n",
       " ['WMM', 770],\n",
       " ['ake', 771],\n",
       " ['cas', 772],\n",
       " ['cha', 773],\n",
       " ['day', 774],\n",
       " ['end', 775],\n",
       " ['mpi', 776],\n",
       " ['ors', 777],\n",
       " ['vid', 778],\n",
       " ['www', 779],\n",
       " ['▁If', 780],\n",
       " ['▁WA', 781],\n",
       " ['▁af', 782],\n",
       " ['▁es', 783],\n",
       " ['▁ex', 784],\n",
       " ['▁sa', 785],\n",
       " ['▁国内', 786],\n",
       " ['いただ', 787],\n",
       " ['された', 788],\n",
       " ['として', 789],\n",
       " ['による', 790],\n",
       " ['参加を', 791],\n",
       " ['男女各', 792],\n",
       " ['障害者', 793],\n",
       " ['anda', 794],\n",
       " ['cate', 795],\n",
       " ['ener', 796],\n",
       " ['heel', 797],\n",
       " ['hose', 798],\n",
       " ['male', 799],\n",
       " ['omes', 800],\n",
       " ['pply', 801],\n",
       " ['ries', 802],\n",
       " ['seas', 803],\n",
       " ['▁Age', 804],\n",
       " ['▁Com', 805],\n",
       " ['▁LLC', 806],\n",
       " ['▁use', 807],\n",
       " ['いません', 808],\n",
       " ['とします', 809],\n",
       " ['市民アス', 810],\n",
       " ['障がい者', 811],\n",
       " ['chair', 812],\n",
       " ['mbers', 813],\n",
       " ['onshi', 814],\n",
       " ['▁comp', 815],\n",
       " ['▁male', 816],\n",
       " ['▁resp', 817],\n",
       " ['した場合は', 818],\n",
       " ['については', 819],\n",
       " ['arting', 820],\n",
       " ['ayment', 821],\n",
       " ['llowed', 822],\n",
       " ['ration', 823],\n",
       " ['▁Those', 824],\n",
       " ['▁Wanda', 825],\n",
       " ['▁athle', 826],\n",
       " ['▁start', 827],\n",
       " ['大阪スポーツ', 828],\n",
       " ['casting', 829],\n",
       " ['hibited', 830],\n",
       " ['omestic', 831],\n",
       " ['qualifi', 832],\n",
       " ['verseas', 833],\n",
       " ['▁Champi', 834],\n",
       " ['▁female', 835],\n",
       " ['市民アスリート', 836],\n",
       " ['bbottWMM', 837],\n",
       " ['▁members', 838],\n",
       " ['エリートランナー', 839],\n",
       " ['heelchair', 840],\n",
       " ['▁athletes', 841],\n",
       " ['▁starting', 842],\n",
       " ['▁Championshi', 843],\n",
       " ['▁Prefectural', 844],\n",
       " ['▁information', 845],\n",
       " ['▁Applications', 846],\n",
       " ['.;', 847],\n",
       " ['/)', 848],\n",
       " ['au', 849],\n",
       " ['eg', 850],\n",
       " ['so', 851],\n",
       " ['su', 852],\n",
       " ['。(', 853],\n",
       " ['なお', 854],\n",
       " ['なか', 855],\n",
       " ['には', 856],\n",
       " ['は別', 857],\n",
       " ['を受', 858],\n",
       " ['イン', 859],\n",
       " ['チケ', 860],\n",
       " ['ブロ', 861],\n",
       " ['一切', 862],\n",
       " ['万博', 863],\n",
       " ['位を', 864],\n",
       " ['基準', 865],\n",
       " ['完走', 866],\n",
       " ['定員', 867],\n",
       " ['当選', 868],\n",
       " ['必要', 869],\n",
       " ['応援', 870],\n",
       " ['情報', 871],\n",
       " ['招待', 872],\n",
       " ['新聞', 873],\n",
       " ['日以', 874],\n",
       " ['期間', 875],\n",
       " ['発行', 876],\n",
       " ['自己', 877],\n",
       " ['資格', 878],\n",
       " ['Tue', 879],\n",
       " ['all', 880],\n",
       " ['ime', 881],\n",
       " ['loc', 882],\n",
       " ['not', 883],\n",
       " ['ong', 884],\n",
       " ['red', 885],\n",
       " ['rol', 886],\n",
       " ['ven', 887],\n",
       " ['▁Ex', 888],\n",
       " ['▁No', 889],\n",
       " ['▁al', 890],\n",
       " ['▁de', 891],\n",
       " ['▁di', 892],\n",
       " ['▁if', 893],\n",
       " ['きます', 894],\n",
       " ['を行う', 895],\n",
       " ['を除く', 896],\n",
       " ['大会の', 897],\n",
       " ['bili', 898],\n",
       " ['eder', 899],\n",
       " ['jors', 900],\n",
       " ['lock', 901],\n",
       " ['mail', 902],\n",
       " ['rans', 903],\n",
       " ['trol', 904],\n",
       " ['ules', 905],\n",
       " ['▁Cor', 906],\n",
       " ['▁dis', 907],\n",
       " ['▁may', 908],\n",
       " ['▁per', 909],\n",
       " ['▁sub', 910],\n",
       " ['▁参加料', 911],\n",
       " ['なかった', 912],\n",
       " ['チケット', 913],\n",
       " ['ブロック', 914],\n",
       " ['メールで', 915],\n",
       " ['位を表彰', 916],\n",
       " ['当選通知', 917],\n",
       " ['招待選手', 918],\n",
       " ['決済金額', 919],\n",
       " ['Gener', 920],\n",
       " ['arity', 921],\n",
       " ['ating', 922],\n",
       " ['ction', 923],\n",
       " ['olicy', 924],\n",
       " ['unnet', 925],\n",
       " ['vited', 926],\n",
       " ['▁JAAF', 927],\n",
       " ['▁Time', 928],\n",
       " ['▁have', 929],\n",
       " ['▁issu', 930],\n",
       " ['▁sent', 931],\n",
       " ['▁take', 932],\n",
       " ['▁wave', 933],\n",
       " ['▁Feder', 934],\n",
       " ['▁Inter', 935],\n",
       " ['▁Sport', 936],\n",
       " ['▁apply', 937],\n",
       " ['▁cance', 938],\n",
       " ['▁email', 939],\n",
       " ['▁those', 940],\n",
       " ['は認めません', 941],\n",
       " ['スポーツ協会', 942],\n",
       " ['チケット付き', 943],\n",
       " ['応援ランナー', 944],\n",
       " ['General', 945],\n",
       " ['▁Corpor', 946],\n",
       " ['▁Majors', 947],\n",
       " ['▁Sports', 948],\n",
       " ['▁before', 949],\n",
       " ['▁cannot', 950],\n",
       " ['▁record', 951],\n",
       " ['位を表彰します', 952],\n",
       " ['ransport', 953],\n",
       " ['▁Runners', 954],\n",
       " ['▁allowed', 955],\n",
       " ['▁entries', 956],\n",
       " ['障がい者ランナー', 957],\n",
       " ['▁Domestic', 958],\n",
       " ['▁Overseas', 959],\n",
       " ['▁marathon', 960],\n",
       " ['heelchairs', 961],\n",
       " ['▁finishing', 962],\n",
       " ['▁accordance', 963],\n",
       " ['▁prohibited', 964],\n",
       " ['▁certificate', 965],\n",
       " ['大阪スポーツ応援ランナー', 966],\n",
       " ['▁applications', 967],\n",
       " ['▁Championships', 968],\n",
       " ['AT', 969],\n",
       " ['HK', 970],\n",
       " ['HO', 971],\n",
       " ['de', 972],\n",
       " ['im', 973],\n",
       " ['iu', 974],\n",
       " ['km', 975],\n",
       " ['ms', 976],\n",
       " ['tc', 977],\n",
       " ['td', 978],\n",
       " ['▁K', 979],\n",
       " ['その', 980],\n",
       " ['とし', 981],\n",
       " ['に関', 982],\n",
       " ['に限', 983],\n",
       " ['の方', 984],\n",
       " ['の際', 985],\n",
       " ['への', 986],\n",
       " ['また', 987],\n",
       " ['より', 988],\n",
       " ['を負', 989],\n",
       " ['カー', 990],\n",
       " ['クレ', 991],\n",
       " ['スト', 992],\n",
       " ['ター', 993],\n",
       " ['テレ', 994],\n",
       " ['ニッ', 995],\n",
       " ['プラ', 996],\n",
       " ['中止', 997],\n",
       " ['予想', 998],\n",
       " ['交通', 999],\n",
       " ...]"
      ]
     },
     "execution_count": 15,
     "metadata": {},
     "output_type": "execute_result"
    }
   ],
   "source": [
    "sp = spm.SentencePieceProcessor()\n",
    "sp.load('toy_tokenizer.model')\n",
    "vocab = [[sp.id_to_piece(idx), idx] for idx in range(sp.get_piece_size())]\n",
    "vocab"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 16,
   "id": "95d7ca09-6ceb-4b23-bd2b-a575c0415ead",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[1290, 295, 1370, 1387, 1364, 1560, 1485, 1398, 1811, 1404, 1364, 329, 1437, 582, 960]\n",
      "['▁he', 'll', 'o', ',', '▁', 'こ', 'ん', 'に', 'ち', 'は', '▁', 'マラ', 'ソ', '▁マラソン', '▁marathon']\n"
     ]
    }
   ],
   "source": [
    "ids = sp.encode(\"hello, こんにちは マラソ マラソン marathon\")\n",
    "print(ids)\n",
    "\n",
    "print([sp.id_to_piece(idx) for idx in ids])"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 17,
   "id": "d1898934",
   "metadata": {},
   "outputs": [
    {
     "name": "stdout",
     "output_type": "stream",
     "text": [
      "[1364, 1386, 1384, 1419, 1364, 1386, 1364, 1384, 1386, 1396, 1471, 1419, 1364, 1461, 1471, 1419, 1364, 39, 1386, 1384, 1419, 1364, 1471, 1461, 1384, 1419, 46, 1386, 1384, 1415, 1429, 1419, 1396, 1419, 1384, 1419]\n",
      "['▁', '1', '2', '3', '▁', '1', '▁', '2', '1', '.', '9', '3', '▁', '8', '9', '3', '▁', '<0x24>', '1', '2', '3', '▁', '9', '8', '2', '3', '<0x2B>', '1', '2', '/', '4', '3', '.', '3', '2', '3']\n"
     ]
    }
   ],
   "source": [
    "# What happens if you set `split_digits=False`?\n",
    "ids = sp.encode(\"123 1 21.93 893 $123 9823+12/43.323\")\n",
    "print(ids)\n",
    "\n",
    "print([sp.id_to_piece(idx) for idx in ids])"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "6f1405d0",
   "metadata": {},
   "source": [
    "### Recap\n",
    "\n",
    "#### Tokenizers\n",
    "A **tokenizer** encodes raw text into a sequence of discrete tokens, and vice-versa (decoding).\n",
    "\n",
    "We saw the *BPE* algorithm for training a tokenizer. Training means that we iteratively find the most frequent token pairs and merge them into a new vocabulary element.\n",
    "\n",
    "We can then encode and decode tokens with the trained tokenizer.\n",
    "\n",
    "#### Practical considerations\n",
    "\n",
    "1. **Dataset:** Since we train a tokenizer on a dataset, the composition of the dataset has a large impact on the vocabulary. Consider the following:\n",
    "    - What if the tokenizer training dataset only had English text?\n",
    "    - What if the only code in the tokenizer training dataset was Python code?\n",
    "\n",
    "2. **Vocabulary size**: We need to choose a vocabulary size as a hyperparameter. As we **increase the vocab size**:\n",
    "    - Shorter sequence length for a given string of text.\n",
    "    - Size of token embedding table and language model head increase.\n",
    "    - Fewer examples per token; each token may be undertrained.\n",
    "    - Less computation per semantic unit\n",
    "\n",
    "If we decrease the vocab size (e.g. to UTF-8 bytes in the limit), the sequence length increases.\n",
    "\n",
    "3. **Odd phenomena** follow from using BPE tokenization. For example:\n",
    "    - Some numeric tokens may be grouped (e.g. `123`) while others arent (e.g. `134` -> `1, 3, 4`). \n",
    "    - If we end a prompt in a space, e.g. `world. `, the following token may begin in a character, e.g. `world. Next` gets tokenized into `world. `, `Next`. However, the training set may have used ` Next`, so the `Next` token is out of distribution. \n",
    "    \n",
    "You can check out Andrej's video for an excellent discussion of these and other phenomena."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "5133af89",
   "metadata": {},
   "outputs": [],
   "source": []
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "anlp",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.10.18"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}