Email Labeling with AI
▶ Watch the originalThe idea
The argument is that for email labeling, instead of feeding all labels to one Large Language Model (LLM) at once, it’s more effective to pull each label individually and have the LLM assess the likelihood that the label applies to the email. This approach ensures that the AI is less likely to be biased towards the first or last labels, making the labeling process more reliable. By processing each label in parallel, the system can scale easily, whether you have just a few labels or many, without affecting performance. This method not only improves accuracy but also enhances the scalability of the system, making it a robust solution for managing and organizing emails efficiently.
Why it works
- The argument is that by feeding each label to an LLM individually, the system avoids bias towards the first or last label, making the labeling more reliable.
- Each LLM focuses solely on whether a specific label fits the email, outputting a likelihood score that is then compared to determine the best fit.
- This parallel processing approach is more scalable, as the number of labels does not significantly impact the system's performance, making it efficient and reliable regardless of the label count.
The playbook
- Identify Your Labels: First, define all the labels you need for your email categorization. Think through common types of emails, such as work-related, personal, promotional, or urgent.
- Set Up the Database: Create a database to store both your emails and the labels. Ensure that the system can efficiently handle and update these entries as new emails come in.
- Prepare Your Language Models: For each label, prepare a language model that can assess the probability of an email belonging to that category. This might involve training or selecting pre-trained models that can handle text classification tasks.
- Parallel Processing: When a new email arrives, save it to the database and then pass it through each label’s language model in parallel. Each model will output a score indicating the likelihood that the email belongs to that label.
- Assign the Highest Score: After processing, the label with the highest score will be assigned to the email. This ensures that the most relevant label is chosen, improving the accuracy and reliability of your email categorization.
Where people get it wrong
Where people get it wrong is in how they approach the auto-labeling process. Many start by feeding all the labels to a single large language model (LLM) at once. This method often fails because the LLM tends to pick from the first few or last few labels, rather than considering the middle labels. Instead, you should label each email individually by running each label through the LLM one at a time. This ensures a more accurate and reliable categorization.
Another common mistake is not saving the emails in a database first. Without this step, you cannot efficiently pull and process each label. Always save the emails before running them through the LLM. This approach not only improves accuracy but also scales better, whether you have three or forty-five labels.
Lastly, treating AI as a magical sentient being is a significant misstep. AI is a tool, and it needs to be used as such. By treating it like reliable infrastructure, you can ensure that the auto-labeling process is consistent and reliable. This means framing the problem correctly and leveraging the tools available to you in a straightforward, efficient manner.
Do this next
- When a new email arrives, log it into your database.
- For each label you've set up, use a separate LLM to determine the likelihood that the label applies to the email.
- Assign the label with the highest score to the email.
- Run these processes in parallel for efficiency and reliability.