Bidyarthi Paul
Prospective Ph.D. candidate (Fall-2027)
Ahsanullah University of Science and Technology (AUST)

About Me

I am a CS graduate working in Natural Language Processing (NLP), with a focus on Bengali and other low-resource language technologies. My research spans Bengali mathematical reasoning, regional dialect processing, named entity recognition, and harmful content detection. More recently, I have developed a growing interest in LLM reasoning and AI safety, particularly in understanding how large language models reason, fail, and behave in safety-critical settings.

🎓 I am actively looking for PhD opportunities for Fall 2027 in NLP, LLM reasoning, and AI safety. My long-term goal is to contribute to building robust, interpretable, and reliable AI systems, especially for low-resource and multilingual settings. If my research interests align with your group, I would be glad to connect 📩.


Education
  • Ahsanullah University of Science and Technology (AUST)
    B.Sc. in Computer Science and Engineering (CSE)
    Dec. 2020 - Jan. 2025
Experience
  • Southeast University, Dhaka, Bangladesh
    Adjunct Lecturer
    23rd February, 2025 – Present
  • ELITE Research Lab LLC, Remote
    Research Assistant
    25th May, 2026 – Present
Honors & Awards
  • Received a research grant from the Institute of Research & Training (IRT), Southeast University for ANCHOLIK-NER.
    2025
  • Dean's List of Honor, 2nd Position
    2024
  • Consistently received merit based waivers in the 1st, 2nd, 3rd, 5th, 6th, and 7th semesters based on top academic ranking and merit within the CSE Department.
    2020 - 2024
  • Scholars' School and College Science Fair 2nd Runners up
    2016
  • General Scholarship in JSC (Junior School Certificate)
    2014
News
2026
Joined ELITE Research Lab LLC as Research Assistant
May 24
Selected Publications (view all )
ANCHOLIK-NER: A benchmark dataset for bangla regional named entity recognition
PLOS ONE
ANCHOLIK-NER: A benchmark dataset for bangla regional named entity recognition

Bidyarthi Paul; Faika Fairuj Preotee; Shuvashis Sarker; Shamim Rahim Refat; Shifat Islam; Tashreef Muhammad; Mohammad Ashraful Hoque; Shahriar Manzoor.

PLOS ONE 2026

TL;DR: We introduce ANCHOLIK-NER, the first benchmark dataset for Bangla regional named entity recognition, and show that BanglaBERT provides the strongest baseline across five major dialect regions.

#Bangla NLP#Named Entity Recognition#Regional Dialects#Dataset#BanglaBERT

ANCHOLIK-NER: A benchmark dataset for bangla regional named entity recognition

Bidyarthi Paul; Faika Fairuj Preotee; Shuvashis Sarker; Shamim Rahim Refat; Shifat Islam; Tashreef Muhammad; Mohammad Ashraful Hoque; Shahriar Manzoor.

PLOS ONE 2026

TL;DR: We introduce ANCHOLIK-NER, the first benchmark dataset for Bangla regional named entity recognition, and show that BanglaBERT provides the strongest baseline across five major dialect regions.

#Bangla NLP#Named Entity Recognition#Regional Dialects#Dataset#BanglaBERT

PLOS ONE
BIDWESH: A Bangla regional based hate speech detection dataset
arXiv
BIDWESH: A Bangla regional based hate speech detection dataset

Azizul Hakim Fayaz; MD Uddin; Rayhan Uddin Bhuiyan; Zakia Sultana; Md Samiul Islam; Bidyarthi Paul; Tashreef Muhammad; Shahriar Manzoor.

arXiv 2025

TL;DR: We introduce BIDWESH, the first multi-dialectal Bangla hate speech dataset, covering Barishal, Noakhali, and Chittagong dialects to support dialect-aware hate speech detection and fairer content moderation.

#Bangla NLP#Hate Speech Detection#Regional Dialects#Dataset#Low-resource NLP

BIDWESH: A Bangla regional based hate speech detection dataset

Azizul Hakim Fayaz; MD Uddin; Rayhan Uddin Bhuiyan; Zakia Sultana; Md Samiul Islam; Bidyarthi Paul; Tashreef Muhammad; Shahriar Manzoor.

arXiv 2025

TL;DR: We introduce BIDWESH, the first multi-dialectal Bangla hate speech dataset, covering Barishal, Noakhali, and Chittagong dialects to support dialect-aware hate speech detection and fairer content moderation.

#Bangla NLP#Hate Speech Detection#Regional Dialects#Dataset#Low-resource NLP

arXiv
Leveraging large language models for bengali math word problem solving with chain of thought reasoning
TALLIP
Leveraging large language models for bengali math word problem solving with chain of thought reasoning

Bidyarthi Paul; Jalisha Jashim Era; Mirazur Rahman Zim; Tahmid Sattar Aothoi; Faisal Muhammad Shah.

ACM TALLIP 2025

TL;DR: We introduce SOMADHAN, a Bengali math word problem dataset with step-by-step solutions, and show that chain-of-thought prompting substantially improves LLM performance on complex Bengali mathematical reasoning tasks.

#Bangla NLP#Math Word Problems#Large Language Models#Chain-of-Thought#Reasoning

Leveraging large language models for bengali math word problem solving with chain of thought reasoning

Bidyarthi Paul; Jalisha Jashim Era; Mirazur Rahman Zim; Tahmid Sattar Aothoi; Faisal Muhammad Shah.

ACM TALLIP 2025

TL;DR: We introduce SOMADHAN, a Bengali math word problem dataset with step-by-step solutions, and show that chain-of-thought prompting substantially improves LLM performance on complex Bengali mathematical reasoning tasks.

#Bangla NLP#Math Word Problems#Large Language Models#Chain-of-Thought#Reasoning

TALLIP
Empowering bengali education with ai: Solving bengali math word problems through transformer models
ICCIT
Empowering bengali education with ai: Solving bengali math word problems through transformer models

Jalisha Jashim Era; Bidyarthi Paul; Tahmid Sattar Aothoi; Mirazur Rahman Zim; Faisal Muhammad Shah.

ICCIT 2024

TL;DR: We solve Bengali math word problems using transformer-based models and introduce the PatiGonit dataset, showing that mT5 achieves the best performance for translating problems into equations.

#Bangla NLP#Math Word Problems#Transformer Models#Bengali Education#Low-resource NLP

Empowering bengali education with ai: Solving bengali math word problems through transformer models

Jalisha Jashim Era; Bidyarthi Paul; Tahmid Sattar Aothoi; Mirazur Rahman Zim; Faisal Muhammad Shah.

ICCIT 2024

TL;DR: We solve Bengali math word problems using transformer-based models and introduce the PatiGonit dataset, showing that mT5 achieves the best performance for translating problems into equations.

#Bangla NLP#Math Word Problems#Transformer Models#Bengali Education#Low-resource NLP

ICCIT
Bridging dialects: Translating standard bangla to regional variants using neural models
ICCIT
Bridging dialects: Translating standard bangla to regional variants using neural models

Md Arafat Alam Khandaker; Ziyan Shirin Raha; Bidyarthi Paul; Tashreef Muhammad.

ICCIT 2024

TL;DR: We translate standard Bangla into multiple regional dialects using neural machine translation models and show that BanglaT5 performs best for capturing dialect-specific linguistic variations.

#Bangla NLP#Dialect Translation#Neural Machine Translation#BanglaT5#Low-resource NLP

Bridging dialects: Translating standard bangla to regional variants using neural models

Md Arafat Alam Khandaker; Ziyan Shirin Raha; Bidyarthi Paul; Tashreef Muhammad.

ICCIT 2024

TL;DR: We translate standard Bangla into multiple regional dialects using neural machine translation models and show that BanglaT5 performs best for capturing dialect-specific linguistic variations.

#Bangla NLP#Dialect Translation#Neural Machine Translation#BanglaT5#Low-resource NLP

ICCIT
Analyzing emotions in Bangla social media comments using machine learning and LIME
MIET
Analyzing emotions in Bangla social media comments using machine learning and LIME

Bidyarthi Paul; SM Musfiqur Rahman; Dipta Biswas; Md Ziaul Hasan; Md Zahid Hossain.

MIET 2024

TL;DR: We analyze Bangla social media comments for emotion detection using machine learning and deep learning models, and use LIME to explain model predictions for better interpretability.

#Bangla NLP#Emotion Analysis#Machine Learning#LIME#Social Media Analysis

Analyzing emotions in Bangla social media comments using machine learning and LIME

Bidyarthi Paul; SM Musfiqur Rahman; Dipta Biswas; Md Ziaul Hasan; Md Zahid Hossain.

MIET 2024

TL;DR: We analyze Bangla social media comments for emotion detection using machine learning and deep learning models, and use LIME to explain model predictions for better interpretability.

#Bangla NLP#Emotion Analysis#Machine Learning#LIME#Social Media Analysis

MIET
Improving Bangla regional dialect detection using BERT, LLMs, and XAI
COMPAS
Improving Bangla regional dialect detection using BERT, LLMs, and XAI

Bidyarthi Paul; Faika Fairuj Preotee; Shuvashis Sarker; Tashreef Muhammad.

COMPAS 2024

TL;DR: We improve Bangla regional dialect detection by comparing BERT-based models and large language models, and use LIME to interpret model predictions across multiple dialect regions.

#Bangla NLP#Dialect Detection#BERT#LLM#XAI

Improving Bangla regional dialect detection using BERT, LLMs, and XAI

Bidyarthi Paul; Faika Fairuj Preotee; Shuvashis Sarker; Tashreef Muhammad.

COMPAS 2024

TL;DR: We improve Bangla regional dialect detection by comparing BERT-based models and large language models, and use LIME to interpret model predictions across multiple dialect regions.

#Bangla NLP#Dialect Detection#BERT#LLM#XAI

COMPAS
All publications