Speech Processing project notes

Speech Processing project notes **From cs231n: **

“Discarding pooling layers has also been found to be important in training good generative models, such as variational autoencoders (VAEs) or generative adversarial networks (GANs). It seems likely that future architectures will feature very few to no pooling layers.”

—> Try removing pooling layers

Prefer a stack of small filter CONV to one large receptive field CONV layer.

 —> If our goal is to isolate style from content, how might we shape the convolution stack?

Instead of rolling your own architecture for a problem, you should look at whatever architecture currently works best on ImageNet, download a pretrained model and finetune it on your data. You should rarely ever have to train a ConvNet from scratch or design one from scratch.

A neural algorithm of artistic style **Gatys, L. A., Ecker, A. S., & Bethge, M. (2015). A neural algorithm of artistic style. arXiv preprint arXiv:1508.06576. ** **For image synthesis we found that replacing the max-pooling operation by average pooling improves the gradient flow and one obtains slightly more appealing results, which is why the images shown were generated with average pooling. ** Dmitry Ulyanov: Audio texture synthesis and style transfer **https://dmitryulyanov.github.io/audio-texture-synthesis-and-style-transfer/ ** COMMENT: Can you apply this to transfer the ‘style’ of an original speaker’s voice to some other voice ? That would be the ultimate application in a speech to speech language translation system… ULYANOV: Speech style transfer won’t work with this simple pipeline, unfortunately. **COMMENT: Do you know of resources that reference the state of the art research in this area for voice to voice style transfer for now? Thanks! ** **Style transfer is able to effectively capture style. Just not content.  ** From Deep Speech **Hannun, A., Case, C., Casper, J., Catanzaro, B., Diamos, G., Elsen, E., … & Ng, A. Y. (2014). Deep speech: Scaling up end-to-end speech recognition. arXiv preprint arXiv:1412.5567. ** I have been thinking that we may need a different architecture to recognize speech content. Look at Deep Speech for example. **A recurrent GAN is going to detect that there’s no semantic content in ours. Or it should! ** Uses spectrogram features **Lexicon-Free Conversational Speech Recognition with Neural Networks ** **Maas, A. L., Xie, Z., Jurafsky, D., & Ng, A. Y. (2015). Lexicon-Free Conversational Speech Recognition with Neural Networks. In HLT-NAACL (pp. 345-354). ** CONVOLUTIONAL, LONG SHORT-TERM MEMORY, FULLY CONNECTED DEEP NEURAL NETWORKS **Sainath, T. N., Vinyals, O., Senior, A., & Sak, H. (2015, April). Convolutional, long short-term memory, fully connected deep neural networks. In Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on (pp. 4580-4584). IEEE. ** ****Very deep multilingual convolutional neural networks for LVCSR


**Sercu, T., Puhrsch, C., Kingsbury, B., & LeCun, Y. (2016, March). Very deep multilingual convolutional neural networks for LVCSR. In Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on (pp. 4955-4959). IEEE. ** Very similar architecture to ours (based on VGG ConvNet) From E2E Speech Recognition **Graves, A., & Jaitly, N. (2014). Towards End-To-End Speech Recognition with Recurrent Neural Networks. In ICML (Vol. 14, pp. 1764-1772). ** Uses spectrogram features Ideas: If we wanted to train our own conv net, we could use a corpus of sentences and a corpus of speakers, so that we have variation in style and content in two different dimensions. Does our corpus work like that now? **If a LSTM neural net is the best for end-to-end speech recognition, perhaps we could layer this on top of a conv stack? That way, the conv would encode the style and the LSTM would encode the speech content. Sounds complicated. **