Skip to content

Speech dataset to test speech to text transcription #18

Description

@qian-chu

Basic Information

Select the dataset status:

  • I have collected the data and want to upload it to the OSF sample data repository (https://doi.org/10.17605/OSF.IO/3N85H).
  • I would like the maintainers to collect or obtain this dataset.

If you have the data, which formats are available?

  • Pupil Cloud
  • Native
  • Both (Pupil Cloud and Native) — Recommended

Dataset Description

Two short recordings (<2 minutes each) in which a participant reads a short text aloud, one in English and one in German.

Benefits and Use Cases

This dataset would serve as a benchmark for testing speech-to-text transcription workflows using AI-based tools (e.g., Whisper).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

sample_dataQuestions regarding sample data or suggestion for new sample data

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions