Skip to main content
GalsenAI

Collaborative project

AfriQA

Creation of an open-access evaluation dataset for answering questions in African languages.

  • Question-réponse
  • Langues africaines
  • Corpus

The project

AfriQA is a collaborative cross-lingual Question-Answering (QA) dataset spanning 10 African languages. It aims to bridge the gap in natural language processing for low-resource languages by providing a high-quality benchmark for evaluating models in African languages.

Historically, models are evaluated on datasets translated from English, which introduces cultural and linguistic biases. AfriQA addresses this by having native speakers write the questions directly.

African languages
10
Questions
12,000+

Goals

  • Evaluate the performance of QA models on African languages
  • Produce a dataset natively written by native speakers, without going through translation
  • Integrate these languages into state-of-the-art NLP research

Results and publications

Continue

Other projects

All projects
  • Masakhane

    Collaborative

    Research community working on NLP for African languages.

    • NLP
    • Traduction automatique
    • Recherche ouverte
    Active
  • Waxal

    GalsenAI

    A voice recognition project for Wolof.

    • Reconnaissance vocale
    • Langues africaines
    • Corpus
    Active
  • Adama

    GalsenAI

    A Wolof text-to-speech model.

    • Synthèse vocale
    • Langues africaines
    Active