
19tools reviewed
5guides
About Khushboo
Khushboo Gupta is a PhD student in computer science at the University of Illinois Chicago, where she researches natural language processing. As a graduate teaching assistant she has taught Program Design, Data Structures, Introduction to Data Science and Natural Language Processing. Before her PhD she was a software development engineer at Amazon Web Services and a software engineer at Pacific Northwest National Laboratory, and she holds an MS in computer science from Syracuse University. On binarytranslator.ai she reviews the guides on text encoding, data structures and number systems.
In Khushboo's words
Every NLP model starts with text turned into numbers. If the encoding is wrong at that first step, nothing after it can be right.
Tools Khushboo reviewed
Guides
Binary basicsWhat is Base64? How Base64 encoding works
Base64 writes any data with 64 safe characters. See the bit-by-bit steps, the alphabet, why strings end in = or ==, why the result is a third bigger and where you will meet it.10 min read
ProgrammingPython bytes to string and string to bytes
Convert bytes to a string with decode() and a string to bytes with encode(). Includes a live encoder, the str(b) bug, UnicodeDecodeError fixes, int.to_bytes() and hex.11 min read
ProgrammingPython ord() and chr(): character to number and back
ord() returns the Unicode number of a character and chr() turns a number back into a character. Examples for ASCII, accented letters and emoji, plus the errors and how to fix them.9 min read
Binary basicsUnicode, UTF-8 and ASCII explained
ASCII, Unicode and UTF-8 explained: the 128 ASCII codes, Unicode code points, how UTF-8 stores them in 1 to 4 bytes, UTF-16 compared, and what causes garbled text.8 min read
Binary basicsWhat is binary code?
Binary code stores numbers, text, colors, sound and programs with only 0 and 1. See how bits and bytes work, why computers chose two digits and what Hello looks like in binary.8 min readQuestions for Khushboo
- Where does encoding come up in NLP?
- Before a model sees any text, the text is stored as bytes, usually UTF-8, and then split into tokens. A file read with the wrong encoding turns accented letters or emoji into junk, and the model learns from the junk.
- What do students struggle with in data structures?
- Thinking in terms of memory. Once they see that an integer takes a fixed number of bytes and a character takes one to four bytes in UTF-8, a lot of questions about size and speed answer themselves.
- What do you check when you review a guide?
- That each step follows from the one before, and that a student can check every example by hand. If I can't reproduce an example in a few lines of Python, it needs to be clearer.