Building Named Entity Recognition for Code-Mixed Text with XLM-RoBERTa
Named Entity Recognition (NER) is used to identify different types of entities in text, such as peop 2026-10-7 06:0:2 Author: hackernoon.com(查看原文) 阅读量:2 收藏

Named Entity Recognition (NER) is used to identify different types of entities in text, such as people, places, and organizations.

One of the most difficult parts of this project was the limited availability of rich datasets for Roman Urdu–English code-mixed text. Roman Urdu–English code-mixed text means that different languages, such as English and Roman Urdu, can be used within the same sentence.

This makes NER challenging because Roman Urdu can be written in many different ways. For example, one person may write a name as “Mehwish,” while another person may write it as “Mevish.” People can also spell Roman Urdu words differently according to their own writing style. Because of these spelling variations and the limited amount of available data, it is difficult for a model to learn all possible variations and make reliable predictions on different types of code-mixed text.

The Problem with Code-Mixed NER

The main challenge in this project was the many variations in Roman Urdu spelling. Different people can use their own way of writing the same Roman Urdu word, which makes it difficult for a model to learn consistent patterns.

Another challenge is that Roman Urdu and English can appear in the same sentence. The language can switch within a single sentence, so the model needs to learn patterns from the context rather than simply memorizing names. For example, it needs to learn whether a word or phrase represents a person, location, or organization based on how it is used in the sentence.

The hardest part was the limited availability of Roman Urdu–English code-mixed datasets. With a relatively limited source of training data, it becomes more difficult for the model to learn the different spelling and language patterns that occur in real-world text. My project therefore focuses on recognizing three entity types — Person, Location, and Organization — while handling these variations through pattern learning rather than simply memorizing entity names.

Dataset and Entity Types

I used the BIO format for my dataset, where each token is assigned a label. The B label represents the beginning of an entity, the I label represents a token inside or continuing an entity, and the O label represents a token outside an entity.

For example, if the text contains the name “Mehwish Faryad,” it can be represented as Mehwish — B-PER and Faryad — I-PER. A token that does not belong to an entity receives the O label.

The dataset contains 4,348 sentences, 138,879 tokens, and 7,936 named entities. The project focuses on three entity types: Person (PER), Location (LOC), and Organization (ORG).

Using BIO labels allows the model to learn not only which tokens belong to entities, but also where an entity begins and where a multi-token entity continues.

Why XLM-RoBERTa?

I needed a model that could work well with code-mixed text, so I considered models that had been pretrained on multilingual data. I selected XLM-RoBERTa instead of using a standard BERT-based approach because my text contains both Roman Urdu and English, making multilingual representation important for the problem.

XLM-RoBERTa is a multilingual transformer model that I used for the Named Entity Recognition task. I fine-tuned the model for token classification so that it could predict BIO labels for each token in the input text.

This multilingual approach was relevant to my problem because the dataset contains Roman Urdu–English code-mixed text. The goal was to help the model learn patterns from the context rather than simply memorize entity names, while handling the language and spelling variations present in the dataset.

System Architecture

Figure 1: System architecture of the NER applicationFigure 1: System architecture of the NER application

The system starts with the user's Roman Urdu–English code-mixed text. The text is first preprocessed to normalize the input and remove unnecessary characters. It is then divided into manageable text segments before being passed to the XLM-RoBERTa tokenizer. The tokenized input is processed by the fine-tuned XLM-RoBERTa model, which predicts BIO labels for the tokens. These predictions are then reconstructed into named entities such as Person, Location, and Organization. The Flask API handles the model inference and returns the results to the web interface, where the recognized entities can be displayed to the user.

Training and Implementation

I trained my model using Google Colab and used my local laptop for development and testing. I used Hugging Face Transformers and tokenizers for model training and text processing. I started with publicly available datasets and cleaned the data before training. I removed noise from the dataset and kept the user-comment text that was relevant to the NER task. I then converted the text into BIO tagging format using Python. During preprocessing, I removed sentences that contained only `O` tags while keeping a smaller number of them so that the model could still learn the difference between entities and non-entities. After preparing the BIO-formatted dataset, I used a tokenizer and aligned the labels with the tokenized input.

I followed the general transformer-based NER approach described in Train a NER Transformer Model with Just a Few Lines of Code via spaCy 3, while implementing my own training pipeline with XLM-RoBERTa. The model was trained to predict seven labels: `B-PER`, `I-PER`, `B-LOC`, `I-LOC`, `B-ORG`, `I-ORG`, and `O`. Initially, I experimented with an 80/20 data split. During this experiment, I observed overfitting, and the model was not learning the patterns as expected. I therefore changed my evaluation approach and used K-fold cross-validation. I also went back through the dataset, cleaned the data again, and continued working with the prepared dataset before training the model again. Another practical problem I faced was GPU availability and stability during training. My laptop GPU caused problems several times while working with the training dataset, which made local training difficult. Because of this, I used Google Colab for training and my local laptop mainly for development and testing.

Problems I Encountered

One of the first challenges I faced was the quality of the dataset. The dataset was very noisy and contained a large number of sentences that had only `O` tags. These sentences did not provide useful entity information for learning the NER patterns, so I reduced the number of such sentences while keeping some of them for the model to learn non-entity tokens. I also faced a problem when I initially used an 80/20 train-validation split. During training, I observed overfitting, so I changed my approach and used K-fold cross-validation. I also went back to the dataset and performed additional cleaning before training the model again. Another challenge was the GPU available on my laptop.

Training the model required significant GPU resources, and after training the model once, I often needed to retrain it after making updates. In some cases, I reached the GPU limit and could not continue training immediately. Because of this limitation, I used Google Colab for model training. Creating and checking the BIO labels was another difficult part of the project. Roman Urdu words can have different spellings, and some words were incorrectly labeled during the annotation process. I therefore had to manually check and review the labels to correct mistakes before using the data for training.

The model also sometimes produced incorrect predictions. For example, a location name could also be used as a person's name, and the model could predict it as `PER` instead of `LOC`. Similar errors occurred when the model confused Person, Location, and Organization entities. I also noticed a difference between the predictions I received in Google Colab and the predictions from my Flask application. The model was predicting entities more accurately during my Colab experiments, while the Flask application sometimes produced different or incorrect predictions. To investigate this issue, I worked on loading the trained model and tokenizer correctly in the Flask application and tested the inference process with new input text.

Evaluation

I evaluated the model using K-fold cross-validation, a confusion matrix, and training and validation loss. I used these different evaluation views to understand not only how accurately the model classified entity labels, but also how consistently it performed across different validation splits. The K-fold results showed relatively consistent performance across the five folds. The F1 score remained approximately between 86% and 89% across the folds, while the highest recall was approximately 90.5% in Fold 5. This gave me a better view of how the model performed across different portions of the dataset rather than relying only on a single 80/20 split. I also examined a confusion matrix for the BIO entity labels while excluding the `O` label.

Most of the predictions were concentrated along the diagonal, showing that the model correctly classified most of the entity labels. Some confusion still occurred between different entity types, particularly when a name could be interpreted as more than one type of entity. The training and validation loss curves provided another view of the model's learning behavior. Training loss decreased continuously across the epochs, while validation loss decreased at first and then became relatively stable. This showed that the model continued improving on the training data while the improvement on the validation data slowed down.

Overall, the evaluation showed that the model was able to learn useful patterns for Person, Location, and Organization recognition. It also showed why evaluating the model from multiple perspectives was important. K-fold cross-validation helped me examine consistency across different data splits, the confusion matrix showed which BIO labels were being confused, and the loss curves helped me understand the relationship between training and validation performance.

Figure 2: Validation performance across the five K-fold iterations.Figure 2: Validation performance across the five K-fold iterations.

Figure 3: Confusion matrix for BIO entity labels, excluding the O label.Figure 3: Confusion matrix for BIO entity labels, excluding the O label.

Figure 4: Training and validation loss across training epochs.Figure 4: Training and validation loss across training epochs.

Turning the Model into a Web Application

After training the NER model, I integrated it into a web application using Flask. The purpose of the application was to allow a user to enter Roman Urdu–English code-mixed text and receive the detected Person, Location, and Organization entities. When a user submits new text, the Flask backend first preprocesses the input. The preprocessing includes Unicode normalization, removal of unnecessary control or invisible characters, conversion to lowercase, and whitespace cleanup. This helps keep the input consistent with the processing used during inference. The application then passes the processed text to the trained XLM-RoBERTa model.

Before prediction, the text is divided into smaller segments when necessary. This is useful for longer inputs because the model uses a maximum input length, so processing the text in smaller segments helps avoid losing information from longer text. The tokenizer converts the input into the format required by XLM-RoBERTa. The model then predicts BIO labels for the tokens. The backend reconstructs these predictions into complete entities and returns information such as the entity text, entity type, position, and prediction confidence. For example, if the input contains a person's name, a location, and an organization, the application can return them as `PER`, `LOC`, and `ORG` entities. The results are then displayed through the web interface.

I also connected the application to MongoDB to store analysis history and prediction results. This allowed previously analyzed text and its results to be stored rather than showing only the current prediction. The final application therefore connects several parts of the project: user input from the web interface, text preprocessing, text segmentation, the XLM-RoBERTa tokenizer and trained model, entity reconstruction, Flask-based inference, and MongoDB storage.The project source code is available on GitHub

What I Learned

This project taught me that working with Roman Urdu–English code-mixed text is more challenging than working with standard English text. Roman Urdu can have many different spelling variations, which makes it difficult for a model to learn consistent patterns from the data. I also learned that dataset cleaning and labeling are some of the most important parts of an NER project. If the dataset contains too much noise or incorrect labels, even a strong model cannot learn the patterns properly. Preparing and reviewing the dataset therefore became an important part of my work rather than just a preprocessing step.

Another lesson was that model evaluation is not only about training the model and checking one result. My initial 80/20 split showed overfitting, so I used K-fold cross-validation to evaluate the model across different portions of the dataset. This gave me a better understanding of how consistently the model was learning the entity patterns. I also learned how to take a trained machine learning model beyond a notebook and integrate it into a functional web application. Connecting the trained model with Flask, the frontend, and MongoDB helped me understand how machine learning inference works as part of a complete application rather than as an isolated model.

Finally, I learned that model performance can change when moving from the training environment to an application. The predictions I observed in Google Colab were sometimes different from the predictions in my Flask application, which made deployment and inference another important part of the project. If I continue working on this project, I would like to train the model on a larger and more diverse dataset, improve the accuracy of entity recognition, reduce incorrect predictions, and experiment with other suitable models to determine whether they can handle Roman Urdu–English code-mixed text more accurately.

Conclusion

Building a Named Entity Recognition system for Roman Urdu–English code-mixed text showed me that a successful machine learning project depends on much more than selecting a model. Dataset cleaning, BIO labeling, evaluation, and deployment all had a major effect on the final system. Using XLM-RoBERTa gave me a practical way to work with multilingual and code-mixed text, while K-fold cross-validation helped me evaluate the model across different portions of the dataset. I also learned that a model that performs well during training does not automatically behave the same way when integrated into a real web application. The project also helped me understand the complete process of taking an NLP model from a dataset to a working application using Python, Flask, and MongoDB. There is still room to improve the system, particularly by using a larger and more diverse dataset, reducing incorrect entity predictions, and exploring other models. For me, the most important lesson was that building an ML application is an iterative process. Data preparation, model training, evaluation, debugging, and deployment all require testing and refinement.


文章来源: https://hackernoon.com/building-named-entity-recognition-for-code-mixed-text-with-xlm-roberta?source=rss
如有侵权请联系:admin#unsafe.sh