|
191 | 191 | "source": [ |
192 | 192 | "ray.timeline(filename=\"timeline03.json\")" |
193 | 193 | ] |
| 194 | + }, |
| 195 | + { |
| 196 | + "cell_type": "markdown", |
| 197 | + "metadata": {}, |
| 198 | + "source": [ |
| 199 | + "### Application: NLP Pipeline\n", |
| 200 | + "\n", |
| 201 | + "To demonstrate the practical applications of nested remote functions, we create a NLP pipeline that analyzes some interesting Wikipedia pages and we speed up this pipeline using nested remote functions. Though this is only a toy example, this example can easily be scaled up with very little changes! \n", |
| 202 | + "\n", |
| 203 | + "For this NLP pipeline, we first use the function `parse_wikipedia` to parse a Wikipedia page on a given topic. Within `parse_wikipedia`, we first tokenize the page using the `tokenize` function and then feed the tokens to `entity_recognizer`, a named entity recognizer. Not satisfied with just the named entities of each language, we also search for the Wikipedia pages of the first 10 unique named entities, recursively until we hit a recursion depth of 2. Finally, we return a pandas dataframe containing all the named entities we found.\n", |
| 204 | + "\n", |
| 205 | + "As an example, we parse the wikipedia entries for python, java, and c++, the languages that are used in ray!\n", |
| 206 | + "\n", |
| 207 | + "**NOTE:** We use the `en_core_web_sm`, a pre-trained model provided by `spacy`. Make sure you have it downloaded by running the following piece of code within a jupyter notebook cell.\n", |
| 208 | + "\n", |
| 209 | + "```\n", |
| 210 | + "!python -m spacy download en_core_web_sm\n", |
| 211 | + "```\n", |
| 212 | + "After running the above code, you may need to restart the notebook kernel." |
| 213 | + ] |
| 214 | + }, |
| 215 | + { |
| 216 | + "cell_type": "code", |
| 217 | + "execution_count": null, |
| 218 | + "metadata": {}, |
| 219 | + "outputs": [], |
| 220 | + "source": [ |
| 221 | + "import modin.pandas as pd\n", |
| 222 | + "import spacy\n", |
| 223 | + "import wikipedia\n", |
| 224 | + "\n", |
| 225 | + "MAX_LINKS = 2\n", |
| 226 | + "MAX_DEPTH = 2" |
| 227 | + ] |
| 228 | + }, |
| 229 | + { |
| 230 | + "cell_type": "code", |
| 231 | + "execution_count": null, |
| 232 | + "metadata": {}, |
| 233 | + "outputs": [], |
| 234 | + "source": [ |
| 235 | + "def tokenize(text):\n", |
| 236 | + " time.sleep(2)\n", |
| 237 | + " nlp = spacy.load(\"en_core_web_sm\")\n", |
| 238 | + " return nlp(text)\n", |
| 239 | + " \n", |
| 240 | + "def entity_recognizer(tokens, topic):\n", |
| 241 | + " time.sleep(2)\n", |
| 242 | + " results = []\n", |
| 243 | + " for token in tokens.ents:\n", |
| 244 | + " results.append([topic, token.text, token.lemma_, token.label_])\n", |
| 245 | + " \n", |
| 246 | + " return results\n", |
| 247 | + "\n", |
| 248 | + "def recursive_wiki_scraper(topic, depth=0):\n", |
| 249 | + " try:\n", |
| 250 | + " wiki_page = wikipedia.page(topic)\n", |
| 251 | + " except:\n", |
| 252 | + " return []\n", |
| 253 | + " wiki_links = wiki_page.links[:MAX_LINKS]\n", |
| 254 | + "\n", |
| 255 | + " page_tokens = tokenize(wiki_page.content)\n", |
| 256 | + " topic_result = entity_recognizer(page_tokens, topic)\n", |
| 257 | + " result = []\n", |
| 258 | + " \n", |
| 259 | + " if depth < MAX_DEPTH:\n", |
| 260 | + " for link in wiki_links:\n", |
| 261 | + " result.extend(recursive_wiki_scraper(link, depth+1))\n", |
| 262 | + " \n", |
| 263 | + " result.extend(topic_result)\n", |
| 264 | + " return result" |
| 265 | + ] |
| 266 | + }, |
| 267 | + { |
| 268 | + "cell_type": "markdown", |
| 269 | + "metadata": {}, |
| 270 | + "source": [ |
| 271 | + "Now let's try and get some information on the languages that ray is built on!" |
| 272 | + ] |
| 273 | + }, |
| 274 | + { |
| 275 | + "cell_type": "code", |
| 276 | + "execution_count": null, |
| 277 | + "metadata": {}, |
| 278 | + "outputs": [], |
| 279 | + "source": [ |
| 280 | + "start = time.time()\n", |
| 281 | + "\n", |
| 282 | + "languages = [\"Python\", \"Java programming\", \"C++\"]\n", |
| 283 | + "results = []\n", |
| 284 | + "for lang in languages:\n", |
| 285 | + " results.extend(recursive_wiki_scraper(lang))\n", |
| 286 | + " \n", |
| 287 | + "duration = time.time() - start\n", |
| 288 | + "print(\"Constructing the dataframe took {} seconds.\".format(duration))" |
| 289 | + ] |
| 290 | + }, |
| 291 | + { |
| 292 | + "cell_type": "markdown", |
| 293 | + "metadata": {}, |
| 294 | + "source": [ |
| 295 | + "**Exercise:** Speed up the above NLP pipeline using ray and its nested remote functions. To do so, it is recommended that you only make `tokenize` and `entity_recognizer` be remote functions and not the `recursive_wiki_scraper`. Try and understand why this is the case. Below you should find the a sample of the results. Feel free to explore the data to find interesting associations to the programming languages that ray is written in!" |
| 296 | + ] |
| 297 | + }, |
| 298 | + { |
| 299 | + "cell_type": "code", |
| 300 | + "execution_count": null, |
| 301 | + "metadata": {}, |
| 302 | + "outputs": [], |
| 303 | + "source": [ |
| 304 | + "df = pd.DataFrame(results, columns=[\"topic\", \"text\", \"lemma\", \"label\"])\n", |
| 305 | + "df.sample(10)" |
| 306 | + ] |
194 | 307 | } |
195 | 308 | ], |
196 | 309 | "metadata": { |
|
0 commit comments