Experience
Professional Experience
- Graduate Student Researcher
SAILI Lab, Université de Sherbrooke, Quebec, Canada
- Period: September 2023 – August 2025
- Conceptualized and formalized “Open-Target Stance Detection”, a new generative task that enables LLMs to identify both targets and stances without predefined lists, addressing a critical gap in real-world, unstructured data analysis.
- Designed and executed a comparative study of 8+ state-of-the-art models (GPT, Gemini, Llama-3, Mistral) across three benchmark datasets, processing over 17,000 text samples on a high-performance computing (HPC) server.
- Developed “BTSD,” a semantic similarity metric based on BERTweet embeddings, evaluate generative target quality; achieved a Pearson correlation of 0.85 with human judgment, significantly outperforming standard metrics like SemSim (0.57).
- Developed and iterated on prompt strategies, including Chain-of-Thought (CoT) and multi-stage zero-shot instruction, to improve model performance in detecting non-explicit targets by over 15% in complex reasoning scenarios.
- Conducted a granular diagnostic of 500+ model failures, categorizing errors into 8 distinct categories; quantified a specific performance drop in “Non-explicit” target scenarios, where model accuracy fell significantly compared to “Explicit” mentions.
- Managed a team of 3 annotators to establish a gold-standard validation set for generative targets; ensured statistical rigor by calculating a Krippendorff’s alpha of 0.76 and a Fleiss’ kappa of 0.664 for inter-rater agreement.
- Validated that generative LLMs outperform traditional Target-Stance Extraction (TSE) models by 104% in human evaluation relevance scores (0.690 for GPT-4o vs. 0.338 for the TSE baseline).
- Managed the full R&D lifecycle over a 24-month period, from initial literature synthesis and hypothesis formulation to experimental execution, technical documentation, and peer-reviewed publication to ACL.
- Research Engineer (Speech & NLP)
AIMS Lab, UIU, Dhaka, Bangladesh
- Period: February 2023 – May 2023
- Architected a national-scale data acquisition and preprocessing pipeline to train a clinical Named Entity Recognition (NER) model for automated medical scribing, projected to streamline workflows for 100K clinicians nationwide.
- Developed a conversation audio recorder to deploy across 6 distinct clinical locations, establishing a distributed data collection network.
- Secured high-level project clearance and formalized data-sharing protocols with BMRC, DGHS, and UNICEF.
- Engineered an automated audio processing pipeline that compresses raw clinical recordings into 320K bitrate MP3s at a 44K sample rate to ensure optimal acoustic feature density for model training.
- Implemented a redundant storage system using Amazon S3 for centralized cloud access and password-encrypted local storage for immediate clinical data protection.
- Developed a robust anonymization workflow utilizing vocal tuning to disable voice identification systems, ensuring patient privacy while maintaining the linguistic utility of the audio.
- Orchestrated a REST API-driven synchronization between the recording hardware, the DGHS Server, and the Clinic EHR management system to capture critical metadata (Doctor ID, Patient ID, Location ID).
- Established a multi-stage Human-in-the-Loop workflow involving Manual Transcription and Cross-Checking to generate high-accuracy ground truth data.
- Designed an Audio Segmentation process to create paired datasets for acoustic modeling and integrated 2 Domain Experts to perform gold-standard clinical information annotation.
- Implemented a systematic Inter-Annotator Agreement (IAA) verification step to ensure the reliability of labeled data prior to Custom NER model training.
- Machine Learning Research Engineer
Intelsense AI, Dhaka, Bangladesh
- Period: September 2021 – April 2022
- Implemented a G2P model for Bengali and gained state-of-the-art accuracy (99%) on unseen data.
- Prepared large-scale (nearly 600 hours) audio data for better Bengali ASR training.
- Speech synthesis: Implemented Coqui TTS models for low-resourced language like Bengali.
- Conversational AI: Developed AI-driven chatbots using Rasa Open Source.
- Bengali transcriber: Prepared the annotated corpus for the Bengali transcriber; already in use.
- Machine Learning Research Intern
Intelsense AI, Dhaka, Bangladesh
- Period: June 2021 – August 2021
- Developed the pipeline for Bengali text normalization and punctuation restoration.
- Reviewed the literature of the related technologies.
Voluntary Service
- Reviewer at COLING 2025 (Abu Dhabi, UAE)
- Period: Oct 2024 – Nov 2024
- Manager, Committee of External Affairs at the Association of the Muslims of University of Sherbrooke (AMUS), Canada
- Period: Nov 2024 – Nov 2025
- Student Volunteer at EACL (Dubrovnik, Croatia)
- Period: May 2023 – May 2023
- Helped people find the rooms, their poster, etc. in GatherTown during the virtual poster sessions
- Communication Responsible at Mozilla, Bangladesh
- Period: January 2018 – January 2018
- Affiliated with Ahsanullah University of Science and Technology (AUST)
Honours & Awards
- Marc-Andr´e Roy Scholarship from the Université de Sherbrooke (March 2024)
- Diversity & Inclusion Award (Registration fee waiver, travel grant, hotel accommodation) from EACL (Dubrovnik, Croatia); sponsored by Amazon Science (2023)
- Robi-Datathon 2.0 Finalist (Top 6% among 358 Teams); organized by Robi Axiata Limited (2022)
- Game Showcasing Competition 1st Runner-Up; organized by AUST IDC (Spring 2018)
Contests & Participations
Top