IUB Academic Repository
    • Login
    View Item 
    •   IUBAR Home
    • School of Engineering, Technology & Sciences
    • Computer Science and Engineering
    • Undergraduate Thesis
    • View Item
    •   IUBAR Home
    • School of Engineering, Technology & Sciences
    • Computer Science and Engineering
    • Undergraduate Thesis
    • View Item
    JavaScript is disabled for your browser. Some features of this site may not work without it.

    Advancing Bengali Dialect Identification (DiD) Through The BengDiDa Dataset & Dialect Classification

    Thumbnail
    View/Open
    Thesis Report_ 1920392_1920392.pdf (4.841Mb)
    Date
    2025-05
    Author
    Kabir, Md Zakaria
    Chowdhury, Md. Zannat
    Metadata
    Show full item record
    Abstract
    Dialect Identification (DID) is necessary for recognizing linguistic diversity alongside improving speech technologies, especially for Bengali dialects which are considerably diverse in nature due to influence inflicted by geographical, cultural, and social factors. Phonetic similarities to adjacent dialects, the scarcity of comprehensive datasets, and the dynamic nature of speech sounds raise the difficulty of identifying these dialects. In this paper, we introduce the BengDiDa dataset, an extensive Bengali dialect speech corpus. Beng-DiDa includes 48,000 audio samples from 20 distinct dialects, designed to support the development of advanced modeling techniques. This research also looks at the efficacy of Convolutional Recurrent Neural Networks for the DiD task. The results highlight the importance of effective feature extraction and the management of spatial and channel-wise dependencies, thereby advancing automatic speech recognition and contributing to the preservation of linguistic heritage. Among evaluated architectures, a fine-tuned ResNet-50 backbone raises accuracy from 80.8 % (F1 = 0.83) to 85.3 % (F1 = 0.85), yet a purpose-built four-layer CNN + BiGRU achieves the best score of 90.3 % (F1 = 0.90) on raw MFCC/GFCC/RPLP features, underscoring the advantage of lightweight, task- specific designs.
    URI
    http://ar.iub.edu.bd/handle/11348/1004
    Collections
    • Undergraduate Thesis [47]
    Publisher:
    IUB
    Type:
    Thesis
    Keywords:
    Bengali Dialect, Bangla Accent Speech Dataset, Regional Language, Dialect Identification, Dialect Identification, Speech Recognition, Multipurpose Dataset, Deep Learning, CRNN

    Copyright © 2026  IUB Academic Repository.
    IUB Repository | Contact Us | Send Feedback
    Maintained by  Library Information Technology (LIT)
    LIT
     

     

    Browse

    All of IUBARCommunities & CollectionsBy Issue DateAuthorsTitlesSubjectsThis CollectionBy Issue DateAuthorsTitlesSubjects

    My Account

    LoginRegister

    Statistics

    View Usage Statistics

    Copyright © 2026  IUB Academic Repository.
    IUB Repository | Contact Us | Send Feedback
    Maintained by  Library Information Technology (LIT)
    LIT