Friday, 30 January 2009

DIRAC Time Stretching & Pitch Shifting

Whilst the majority of the recent work has involved a lot of planing and structuring of the tool application and its code base, I have also spent some time investigating the time stretching technique I intend to implement. Over the past few days it has become clear that I am unlikely to succeed in producing a time stretching algorithm capable of the tasks I would deligate to it. There are so many existing algorithms out there, each fairly complicated, that produce results of varying quality depending on the audio signal fed into them. For example the Rabiner and Schafer method does not work so well for polyphonic sounds but the phase vocoder method can introduce an audible after-effect known as "phase smearing".

However, I have come across a C/C++ library specially dedicated to pitch shifting and time stretching known as DIRAC. It comes in various flavours but their completly free DIRAC LE version interested me the most. This still provides the quality that you get with the studio and pro levels but puts some limits on things like supported sample rates and number of channels which don't really affect me too much.

DIRAC was developed by Stephan Bernsee, the founder of Prosoniq, and it can be seen in many professional level products such as Steinberg's Wavelab and Nuendo. It is an incredibly small and simple library containing a total of 7 functions and 3 enumerations, however the quality is astounding for it being a free product ( let's pretend we didn't see the 9870 EURO price tag for the full version ).

After a couple of days of tinkering I've managed to take a sound that's been loaded into FMOD and stretch/squash it using DIRAC. My test program is still a bit flaky though as I'm still not 100% sure how the sound buffers work. Doesn't help when FMOD works in bytes and DIRAC works in samples, but I'll get there.

Friday, 23 January 2009

Updating ...

It has been a while since the last post, partly due to needing a little time to relax but primarily because of the updates needed to be made to my team's Dare To Be Digital 2008 entry - a winning entry if you can believe that. Extra work was needed to be done to it in preparation for the BAFTAs in March, however, now that is out the way I can get down to business again with this project.

With the final submission of my project proposal I narrowed down two main features to focus on that I felt were most neglected in current game audio tools. These features are time stretching and sample accurate looping.

Time Stretching

This is the method of lengthening or shortening an audio sample while preserving its pitch. During my research I have not found a single game specific application that supports this feature meaning that if any samples were to be layered together as part of one track they would have to be edited in another package first before they could be used together if their tempo's differed. In a review of the tool Wwise, Michael Henein stated that he would like to see time stretching implemented in these kind of tools to allow for more automated syncing with musical cues. I agree with his statement as tools are supposed to aid the developer - it's no use if you have one tool but are quite restricted in how you use it unless you have access to others.

In this project the time stretching will be implemented in a similar fashion to how Sony Acid has used it. With Acid, when you load in a sample you can make it fit with the rest of the project your working on by making its tempo match the projects tempo by either stretching or shortening it. This feature in a game specific tool would allow the audio designer/composer to just dump any suitable samples into their application without having to worry about their pacing as the tool will be able to make them all fit together.

Time stretching can have some adverse affects on certain types and lengths of samples such as the creation of artifacts ( clicks and pops ) and just making it sound unplesant. This is where the judgment of the user will be required to assess if more drastic measures are needed such as recording a new sample or changing the layout of the track.

Sample Accurate Looping

Unfortunatly with the audio APIs I have looked into I discovered ( through experience ) that when you set a sound to loop there is a slight delay from when it ends to when it loops round to the start. While this delay is inaudable it does pose an issue where multi-layered, different length sounds are concerned. The reasoning behind this was explained this previous post.

Why not just use samples that are the same length when layering them I hear you ask? Well, if you have ever taken a look at a complicated musical composition in any sequencer based program ( Cubase, Logic, Reason, Acid etc.) you'll notice that little variations, short phrases and contained loops can be found all over the place, seperated by blank spaces. If you were to create a single loop with samples of the same length then each sample would have to be as long as the longest sample. Not only would this take up huge amounts of unnecessary memory/CPU overhead in playing silent passages it also restricts how dynamic the composition can be. Individual phrases or loops could never be randomly played at any time unless an extra layer was added just for those bits - a layer that would again have to be the same length as the longest sample.

With using sample accurate looping, sections that have a lot of silence would simply not be played, reducing memory footprint as well as CPU overhead. Compositions could also be made more dynamic by allowing the composer to define sections that can have variations played in them, of which a certain variation could be picked in real time depending on the game state.

Wednesday, 19 November 2008

Proposal Presenation

This week saw the entire year presenting their individual topics to a panel of lecturers in order to get their projects approved. Out of the few presentations I witnessed there were certainly some interesting, well thought out topics. This post is to sum up my proposal presentation in order to clarify my project's area of focus, as it has changed subtly throughout this blog so far.
Here is an e-rendering of my presentation based on my notes written before hand. It went pretty much as it is written:


" Hi, I'm Jonathon and my project is on developing adaptive music for video games "


"Why developing adaptive music? Well, I have experience in both the technical and artistic elements involved in developing audio for computer games. More notably, I have successfully implemented a basic adaptive music system for the recent Dare To Be Digital competition. This was based on mixing different layered tracks of a song in real-time depending on the main characters status."

"The issues surrounding this topic stem from a constant, growing demand placed on sound designers. This growing demand requires excellent, well established tools to give the content creator more freedom and control. Thus increasing their productivity."


"The aim of the project is to determine where current audio middle-ware has not advanced far enough to meet the demands of the content creator and, if possible, provide new alternatives and improvements. Audio is a cast discipline that contains many elements out with the scope of this project so the focus is purely on the music."


*recite research question*

"To address this question I intend to: Analyse existing audio tools, including applications not specific to video game audio. Determine what the sound designer/composer needs in such tools.
Develop an underlying framework that can support these requirements. Produce a suitable user interface to allow non-programmers to manipulate the framework effectively."


Any Questions? (there was another slide but didn't feel the need to upload the image)

...

Luckily the only major question asked was from my technical supervisor. Unfortunately I did not answer it as elegantly as I had hoped. The question was along the lines of "what do you mean by data-driven?". I cannot begin to replicate the ramblings of my original answer but for those who are interested here is a better thought out response.

"What do I mean by data-driven?"

"In essence the term data-driven is referring to how the framework is to be used. Once integrated into the game engine the framework will no longer require the use of a programmer to define functionality. All the rules and data will be gathered together by the sound designer/composer in the tool which will output information to control the framework. In other words, the "data" gathered in the tool is used to "drive" the framework."

So, to summarise, my project is about finding out as much as I can regarding current audio middle-ware that has some bearing on developing adaptive music. Once enough information has been gathered then it should be easy to ascertain where the game specific tools lack functionality found in the more mature audio applications. I will attempt to implement a few of these aspects in a simple tool of my own to show how they can make developing adaptive music easier.