Wednesday, 19 June 2019

Building KENLM Language Model for Mozilla Baidu DeepSpeech

Most automatic speech recognition (ASR) application requires language model to run properly. Why? Because the ASR engine turn the voice sequences in to most likely phonemes and then form words . From this stage the engine pay no attention to language context. Here the language modelling comes in to play. Right after the words are generated, the words or sentence are fed to the language model, let mention KENLM for Deepspeech. And voila the language-aware sentence are produced.

Deepspeech uses KENLM under the hood as currently it’s the fastest language model library in the wild. From KENLM we need 2 files, i.e. language model binary (lm.binary) and the trie. Let’s see how they are built.

Steps

0. Cloning and making the KENLM

Clone the code from repository https://github.com/kpu/kenlm.

git clone https://github.com/kpu/kenlm

Then build it,

cd kenlm
mkdir -p build
cd build
cmake ..
make -j 4

1. Providing the corpus

Let’s assume we want to build Javanese language model. We are using openslr.org. The data, either the audio and utterances, are available for public in https://www.openslr.org/35/. Specifically for the utterances, we need to get the utt_spk_text.tsv. The structure of file are like following :

00004fe6aa	a4815	Kanthong semar minangka tanduran kang mrambat lan njulur
00005fb7fb	ede87	Rudy uga misuwur amerga inovasi anggoné nggawé bumbu dhasar
0000e5df79	ffe12	Banjur saluran bakal mbelok ning lateral tengen
0001418491	ce6f9	Ana telu seksi tradhisional ya iku Tripolitania Fezzan lan Cyrenaica
0001bbbc2e	0a834	Iwak Mas iki bisa kanggo praktik cuba cuba ana leb

Get only the sentence on third column and lowercase the text by :

 cat utt_spk_text.tsv | cut -d $'\t' -f3 | tr A-Z a-z > corpus.txt

Let’s assume the letters used in corpus.txt comprises of :

a
b
c
d
e
é
f
g
h
i
j
k
l
m
n
o
p
q
r
s
t
u
v
w
x
y
z

2. Creating arpa file

./bin/lmplz --text ${CORPUS_DIR}/corpus.txt --arpa ${CORPUS_DIR}/words.arpa --o 3

3. Building language model binary

./build_binary -T -s ${CORPUS_DIR}/words.arpa ${CORPUS_DIR}/lm.binary

4. Building the trie

./tensorflow/bazel-bin/native_client/generate_trie ${CORPUS_DIR}/alphabet.txt ${CORPUS_DIR}/lm.binary ${CORPUS_DIR}/corpus.txt ${CORPUS_DIR}/trie

Monday, 25 March 2019

Indonesia Presidential Debate Transcription Service in A Nutshell

I was writing this article when I was a software engineer at Bahasa Kita and Indonesia was going to have president election in 2019. As a way to educate the voters about their choice in the election, KPU (General Election Commissions) held debates. The debates were held five times with different topics discussed. The topics were as follows :

  1. Law, Human Right, Corruption, and Terrorism (17 January 2019 at Hotel Bidakara)
  2. Energy and Food, Natural Resources and the Environment, and Infrastructure (17 February 2019 at Hotel Sultan)
  3. Education, Health, Employment and Social and Culture (17 March 2019 at Hotel Sultan)
  4. Ideology, Government, Defense and Security and International Relations (30 March 2019)
  5. Social Economy and Welfare, Finance and Investment and Trade and Industry

As machine learning startup focused on voice, we at Bahasa Kita wanted to contribute to the election through the voice technology. In the time we look for the ideas, we remembered that our founder once had made use the voice technology for the deaf. Thus we decide to do the same thing upon the presidential debate. This was still relevant since we also found many comments from the disabled people hoping there was such technology to help them at Youtube video which has no caption available. Surprisingly, the idea is also needed by normal people when they miss a certain part while the debate is live. They want to review and dig deeper on the candidate’s speech. In short, we propose to transcribe the debate using our automatic speech recognition service and present the transcription quickly while the debate is still ongoing.

We had no long time to implement the idea because we had talked about how to contribute just 8 hours before the first debate. However, during the discussion, other thought appeared. It said that it feels incomplete when we just show the debate verbatim. Then, we searched references from the previous election and found an analytics company that do analytics on the debate verbatim. After analyzing the content, we realized that this analytics can be used as it is neutral. This neutrality is important as we were aware that anything in the political year can be assumed to be not neutral. The interesting thing was just hours or days after our website (debatcapres.bahasakita.co.id) went online, there was one of the candidate’s party contacted us to claim the website for their needs. It is surely a huge amount of money before our eyes. But we chose to decline the offers and keep providing the presidential debate independently for free.

Actually, it is too much if we call them by analytics, so I prefer to call them as a summary. The summary only counts the words statistic without prior knowledge about political, economics, or other needed. The summary comes initially in three forms: the total word count, the topic-specific word count, and word cloud. Word count is just a counting upon words uttered by each candidate. While topic specific word count must relate to the topics. The topics appeared to have other words association. We use these associations to count the topic related to word count. The last, word cloud shows a self-explaining visualization resembling a cloud.

If we take a look at Natural Language Processing (NLP) and Natural Language Understanding (NLU) state-of-the-art, there are still so many kinds of analysis that we can do and we are able to implement. For example, we can define whether utterances tend to polarize to a particular opinion or another way. Or we can display fact based on the news on the internet from the utterance told by the candidates. The most extreme one is we can show the personality and its description based on the given speech. All of them are done using artificial intelligence that we work on every day. But once again, it is too risky to release such information to the public as it is too subjective to judge the quality of analysis. Then we remain on the summary that I just mentioned and give further interpretation to the reader.

Later we were informed that our website is used by journalists to research and find the truth about each statement surfaced during the debate. We were quite happy that our website is useful to others.

Wednesday, 28 November 2018

CMU Dictionary Adaptation to Bahasa Indonesia Lexicon Building

While creating lexicon in voice recognition in Bahasa Indonesia, we need to define the phoneme set by ourselves since there are not such a widely used standard in Bahasa Indonesia. Instead of creating new definition, there is an idea to adapt from existing phoneme set. Obtained from Kaldi resources, we can adapt the phoneme set from English issued by Carnegie Mellon University (CMU Dictionary) which contains 134,000 words.

Bahasa Indonesia is quite simplelook here also as in major case the pronunciation and written letter are the same compared to English. Thus, it is not a tedious work to start building lexicon based on CMU Dictionary although we need to add new phonemes and leave some phonemes.

Table 1 CMU Dictionary Phoneme Set

Phoneme IPA Symbol Indonesian Words Example English Words Example
AA ɑ ternak, gembala odd, balm
AE æ - at, bat
AH ʌ ambil hut, butt
AO ɔ bakpao cow, story
AW saung ought, bout
AY kait, senarai hide, bite
B b lebah, beli bee, buy
CH cuka, ceri cheese, china
D d diam, duduk dump, did
DH ð ridho the, thy
EH ɛ enak, sepak education, bet
ER ɝ bageur*, reueus* hurt
EY - ate, bait
F ɾ faedah, fana fee, forest
G g gerbang, guna green, gate
HH h hampar, unggah he, hair
IH ɪ singgah, ikatan it, implication
IY i - eat, sheep
JH adiraja, keganjilan genuine, jimmy
K k kenangan, batuk key, camp
L l lingkaran, betul luck, love
M m minum, temaram mama, mine
N n naik, menikah knee, nice
NG ŋ yang, ngengat bank, sink
OW bongkar, bogor oat, boat
OY ɔɪ - toy, boy
P p pulsa, peluh pulp, pen
R ɹ ranjau, rintangan right, row
S s sakit, sayang sea, sun
SH ʃ masyarakat, syaikh* shine, she
T t tikung, timpa tea, tone
TH θ rabiul tsani* thug, theta
UH ʊ - hood, book
UW u kuku, suku two, coup
V v viral, vas vee, vocal
W w wejangan, wayang we, wide
Y j yakin, yoga yam, yield
Z z zaman, zamrud zoo, zee
ZH ʒ jangkrik, jerapah seizure, pleasure

From table above, it’s clear that there are phonemes that are not (commonly) used in Bahasa Indonesia. Yet these phoneme set does not cover all Bahasa Indonesia lemma regarding to the root of Bahasa Indonesia which come from majorly Malay, Dutch, Arabic, Chinese, Javanese, and Sundanese. To create the things short, here is list of phonemes that needs to be added to CMU Dictionary for Bahasa Indonesia

Table 2 Addition Phoneme Set

Phoneme IPA Symbol Indonesian Words Example Notes
NY ɲ kenyang, nyamuk alveolo palatal
KH x kholifah voiceless velar fricative
Q ʔ qurban
KX ʕ sa’at voiced pharyngeal fricative
DL dhuhur voiced alveolar sibilant with pharyngealization
GH ɣ ghaib voiced velar fricative

sources :
[1] http://kaldi-asr.org/doc/examples.html
[2] http://www.speech.cs.cmu.edu/cgi-bin/cmudict
[3] https://en.wikipedia.org/wiki/ARPABET
[4] https://open-dict-data.github.io/ipa-lookup/ma/

Saturday, 30 December 2017

Kalender Puasa 2018

Kalender Puasa 2018

Dari Abu Hurairah ra., bahwa Nabi shallallahu alaihi wasallam bersabda:

Demi Dzat yang jiwa Muhammad berada di tangan-Nya, sesungguhnya bau mulut orang yang berpuasa lebih harum di sisi Allah pada hari kiamat daripada bau misk atau kasturi. Dan bagi orang yang berpuasa ada dua kegembiraan, ketika berbuka mereka bergembira dengan bukanya dan ketika bertemu Allah mereka bergembira karena puasanya.”

(HR. Bukhari dan Muslim)

Maka saya hadiahkan, terutama untuk diri sendiri, saudara dan saudariku sebagai pengingat Kalender Puasa 2018 yang dapat diunduh di

Versi bahasa Indonesia

PNG : https://goo.gl/T7Nw6C

PDF : https://goo.gl/EDbgmk

English Version

PNG : https://goo.gl/UZuKsW

PDF : https://goo.gl/qbtQA6

Semoga Allah memberikan kemudahan dan menerima semua amal ibadah kita. Tiada daya dan upaya melainkan dari Allah. Silakan untuk disebarkan kepada keluarga dan kerabat.

Jazakumullahu khairan katsiran.

NB:

1. Kalender ini hanya sebagai perkiraan. Jika di kemudian hari terdapat selisih tanggal, harap menyesuaikan dengan penanggalan yang sesuai.

2. Korespondensi lebih lanjut, mohon melalui www.github.com/linerocks/FastingCalendar2018 di kolom issues

Muslim Fasting Calendar 2018

Muslim Fasting Calendar 2018

The Prophet Muhammad , peace be upon him, said :

By Him in Whose Hands my soul is, the smell coming out from the mouth of fasting person is better in sight of Allah than the smell of musk. (allah says about the fasting person), “He has left his food, drink, and desires for My Sake. The fast is for Me. So I will reward (the fasting person) for it and the reward of good deeds is multiplied ten times.”

^reference : Sahih al-Bukhari 1894^

So , I present to, especially myself, my brother, and sister as a reminder this Muslim Fasting Calendar 2018 which can be obtained in this repository or simply download at :

English Version

PNG : https://goo.gl/UZuKsW

PDF : https://goo.gl/qbtQA6

Bahasa Version

PNG : https://goo.gl/T7Nw6C

PDF : https://goo.gl/EDbgmk

May Allah grant us ease to do good deeds and accept our good deeds. There is no power and no strength except with Allah. Feel free to share to relatives and family.

Jazakumullahu khairan katsiran.

NB:

1. This calendar is only a prediction. If you find any unmatch date, please adjust with the correct one.

2. If you have issues related to this, please kindly contact me via www.github.com/linerocks/FastingCalendar2018 on tab issues

Wednesday, 20 December 2017

Matrix Construction in C++ using Armadillo

Matrix Construction in C++ using Armadillo

As a follow-through of vector, we go on matrix construction. Starting from the definition, matrix is a rectangular array of numbers, symbols, or expressions, arranged in rows and columns. The individual items in an matrix often denoted by or so-called elements, where is row index and is column index.

Because of there is no container for matrix in C/C++ standard library, let’s make our way construction matrix. In C we could use pointer of array or pointer of pointer as follows :

int m = 4, n = 3, i;

// Using pointer of array
int *B[m];
for (i = 0; i < m; i++) {
    A[i] = (int *) calloc(n, sizeof(int));
}

// Using pointer of pointer
int **B = (int **) calloc(m, sizeof(int *));
for (i = 0; i < m; i++) {
    B[i] = (int *) calloc(n, sizeof(int)); 
}

Vector Construction in C++ using Armadillo

Vector Construction in C++ using Armadillo

A collection of data in a row (or column) which may be summed together, multiplied by number, and having the same type (for instance real number, integer, complex number, float) is called vector. In programming, we tend to be familiar with array. Either in mathematics or programming, vector is accessed using index. For example, assume we have a vector that comprise , , , and so forth. Alternatively,

Monday, 14 August 2017

FFTW3 : C++ Implementation on MATLAB/OCTAVE Perspective

FFTW3 : C++ Implementaion on MATLAB/OCTAVE Perspective

I used to code using MATLAB and OCTAVE for my signal processing research. But, when it comes to real implementation and performance, I always stop and wonder how to make my concept coded in C/C++. Moreover, my MATLAB license is expired. All I need is Fourier Transform because it is the basic operation for signal processing. Then I made research on how Fourier Transform could be done in C/C++. I ended up on Fast Fourier Transform in the West (FFTW). It is considered the best and fastest FFT implementation in C/C++.

In this article, I would like to demonstrate basic forward and inverse transform of monotone signal. Because of we use audo signal in assumption, so this is 1D FFT problem. Well I think this simple example could be an entrance gate for MATLAB and OCTAVE programmers to implement their concepts on C/C++. I hope this article can be helpful for us.

Tuesday, 7 February 2017

Solving XOR problem using tiny-dnn

Solving XOR problem using tiny-dnn

REVISED August 13th 2017

Almost all mature deep neural network (DNN) libraries e.g. Tensor Flow, Theano, Caffe, and etc are written in python, not in C/C++. We will hardly find library for DNN written in C/C++. Even if we find one, it requires heavy resources. Fortunately, we now have tiny-dnn. tiny-dnn is a C++11 implementation of deep learning. Nothing needs to be compiled, header only. It is suitable for deep learning on limited computational resource, embedded systems and IoT devices.

For new tiny-dnn user, it may hard to get used with the environment because the examples provided are directly designated to solve MNIST or CIFAR problem. On this post, I try to give example to solve simple problem (XOR) using tiny-dnn. It may sound excessive to use DNN framework only to solve XOR problem,. But for the sake of better understanding of framework structure, I think it’s okay to do so.

Sunday, 15 January 2017

Indonesian Word to English Equal Phoneme Routine

Indonesian Word to English Equal Phoneme Routine

There are no special rules on how word in Indonesia should be converted into phoneme. Mostly , each alphabet in word is phoneme. If we have word cinta, it simply put c+i+n+t+a (Indonesian phoneme) or ch+ih+n+t+aa (English equal phoneme), as phonemes. The exception only applies on diphthong and nasal. Although we have exception, it remains simple. For example word sayang which has nasal ng, could simply be converted into s+a+y+a+ng (Indonesian phoneme) or s+aa+y+aa+ng (English equal phoneme). Or word aura which has diphthong au, could be converted into au+r+a (Indonesian phoneme) or aw+r+aa (English equal phoneme).

Here I write C++ routine to do such job. This routine maybe not the most effective one, but it works though. I design the routine to be able clean the non-necessary characters. Some part of the routine may seem useless. It is because I originally design for many task, but, in the end of the day I left the task to shell script. Number tokenization is not implemented yet. The output of the routine is English equal phoneme.

Persamaan Diferensial : Sebuah Pendahuluan

Persamaan Diferensial : Sebuah Pendahuluan

Persamaan diferensial (Differential equationDE) dapat dimanfaatkan untuk menjelaskan hampir semua fenomena yang kita temui dalam kehidupan sehari-hari. Sebagai contoh, telepon genggam yang kita pakai. Sinyal telepon genggam yang merupakan media transmisi kita dalam mengirim dan menerima informasi, berawal dari persamaan diferensial. Menjelaskan bagaimana planet di tata surya beredar di orbitnya mengelilingi matahari. Atau mengetahui bagaimana berita hoax menyebar di jejaring sosial. Bahkan untuk memahami tingkat penyerapan vitamin C di tubuh kita untuk membantu menangkal dari penyakit atau memahami cepatnya virus influenza penyakit menyebar. Itulah beberapa contoh manfaat persamaan diferensial. Saintis dan insinyur melihat dunia melalui persamaan diferensial. Karena fenomena ilmiah harus terukur dan dijelaskan.

Wednesday, 11 January 2017

Indonesian Phonemes Relation to English International Phonetics Alphabet

Indonesian Phonemes Relation to English International Phonetics Alphabet

Indonesian phonetic system has a simple pattern compared to English phonetization rule. Most of alphabet Indonesia directly stand to phonetic symbol. The addition for Indonesia phonetic symbol includes nasal (ng, ny), diphthong (ay, aw, ey, and oy), fricative (kh and sy), and so forth. To make a better phonetic system comparison between Indonesia and English, it’s important to have a relation table. Based on [1] and [2] (with some minor change from me in E vowel), the relation table between Indonesia and English is shown in following table 1.

Friday, 6 January 2017

Menulis Cantik dengan Stackedit

Menulis Cantik dengan Stackedit
StackEdit-logo

Menjadi seorang penulis, terutama penulis artikel online, sudah selayaknya kita fokus pada konten yang ingin kita sampaikan kepada pembaca. Namun, kadung asyiknya membuat konten, terkadang kita lupa membuat sebuah tulisan yang enak dibaca secara visual. Apalagi kita harus berurusan dengan formatting html atau command dari editor WYSIWYG yang kita gunakan. Sebagai seorang penggemar (hanya pengguna math display sebenarnya) adalah hal yang wajar menginginkan kecantikan dapat dibawa serta dalam hal online-publishing.

Salah satu solusinya adalah menggunakan Markdown sebuah markup language dengan format syntax yang mudah. Terdapat berbagai macam editor untuk Markdown, namun satu yang menarik adalah StackEdit. Dalam artikel ini akan saya ulas fitur dari Markdown pada StackEdit lengkap dengan contohnya. Berikut daftar konten yang akan kita bahas.

Daftar Isi
1. Tag Header 7. Backslash Escapes
2. Emphasis 8. Tautan
3. Quote 9. Blok Kode
4. List 10. Tabel
5. Gambar 11. Ekspresi Matematika LaTeX
6. Ikon 12. Diagram

Thursday, 5 January 2017

Installing IDLAK Deep Neural Network Text to Speech Synthesizer

Installing IDLAK Deep Neural Network Text to Speech Synthesizer

I was on duty to create text-to-speech (TTS) engine. After researching about the latest technology, the state-of-art of TTS had come to Deep Neural Network (DNN) scheme, instead of insisting on Hidden Markov Model (HMM). The main reason of migrating the scheme in to DNN was just the powerfullness of it.

One of the framework of doing DNN TTS is IDLAK. It was fresh from the oven. Their team had just presented their paper on Interspeech 2016. IDLAK is branch of well-known automatic-speech-recognizer (ASR) engine KALDI.

Here I am documenting the whole steps building the engine.