This paper describes the Toshiba Mandarin Text-to-Speech (TTS) system that was submitted to the Blizzard Challenge 2008. The front-end of the system uses machine-learning approaches such as generalized linear models (GLM) and Quantification Method Type 1 (QMT1) to predict pause, duration and F0 contour. According to the predicted prosody information, the back-end of the system uses Toshiba's own "plural unit selection and fusion" method to create fused speech units which contain pitch-cycle waveforms. The pitch-cycle waveforms are then aligned along the predicted pitch marks and are overlapped with each other to generate the final speech waveforms. This paper also addresses the methods used to prepare the speech corpus and tune the performance of the back-end. The evaluation results showed that our Mandarin TTS was in the leading position among the 12 participating TTS systems.