Inviting an author to review:
Find an author and click ‘Invite to review selected article’ near their name.
Search for authorsSearch for similar articles
8
views
0
recommends
+1 Recommend
0 collections
    0
    shares
      • Record: found
      • Abstract: found
      • Article: found
      Is Open Access

      Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions

      Preprint
      ,

      Read this article at

      Bookmark
          There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

          Abstract

          In this paper, we tackle the singing voice phoneme segmentation problem in the singing training scenario by using language-independent information -- onset and prior coarse duration. We propose a two-step method. In the first step, we jointly calculate the syllable and phoneme onset detection functions (ODFs) using a convolutional neural network (CNN). In the second step, the syllable and phoneme boundaries and labels are inferred hierarchically by using a duration-informed hidden Markov model (HMM). To achieve the inference, we incorporate the a priori duration model as the transition probabilities and the ODFs as the emission probabilities into the HMM. The proposed method is designed in a language-independent way such that no phoneme class labels are used. For the model training and algorithm evaluation, we collect a new jingju (also known as Beijing or Peking opera) solo singing voice dataset and manually annotate the boundaries and labels at phrase, syllable and phoneme levels. The dataset is publicly available. The proposed method is compared with a baseline method based on hidden semi-Markov model (HSMM) forced alignment. The evaluation results show that the proposed method outperforms the baseline by a large margin regarding both segmentation and onset detection tasks.

          Related collections

          Most cited references8

          • Record: found
          • Abstract: not found
          • Article: not found

          Exploring the state sequence space for hidden Markov and semi-Markov chains

            Bookmark
            • Record: found
            • Abstract: not found
            • Article: not found

            LyricSynchronizer: Automatic Synchronization System Between Musical Audio Signals and Lyrics

              Bookmark
              • Record: found
              • Abstract: not found
              • Article: not found

              Integrating Additional Chord Information Into HMM-Based Lyrics-to-Audio Alignment

                Bookmark

                Author and article information

                Journal
                05 June 2018
                Article
                1806.01665
                1e72259f-1dd0-4bfc-8928-5fbdee3f263a

                http://arxiv.org/licenses/nonexclusive-distrib/1.0/

                History
                Custom metadata
                Interspeech 2018
                cs.SD cs.IR eess.AS

                Information & Library science,Electrical engineering,Graphics & Multimedia design

                Comments

                Comment on this article