Thursday, 28 August 2014

Profiling for the media

I was recently called upon by Associated Press, the BBC News Channel and Reading's Jack FM to comment on the spoken features of the jihadist from the video footage of the beheading of US journalist James Foley, as it has been suggested that he is British.  You can see me in the Associated Press clip on French television below.


I'm always ready to be called upon to comment on linguistic issues for the media, but this situation was particularly difficult owing to the content of the video.  I hope that the phonetics community is able to assist the authorities. I wish to make it clear that am not carrying out the analysis myself.

The speaker displays many of the features of a British accent known widely as Multicultural London English (MLE), such as producing vowels in e.g. FACE and PRICE as monophthongs, not dropping /h/ sounds (/h/-dropping is common in London accent Cockney) and pronouncing voiced dental fricatives in e.g. the as a [d].  There are glottal stops, which are less common in Afro-Caribbean or African English accents, and /l/-vocalisation. The speaker also has a more syllable-timed speech rhythm; instead of pronouncing the phrase from all walks of life as /frəm ɔːl wɔːks əv laɪf/ it sounds more like [frɒm ɔː wɔːks ɒv lɐːf], with a full vowel in each syllable.

See HERE for information on the features of MLE (yes, it's Wikipedia, but a good summary).

The accent was identified chiefly by Professor Paul Kerswill and colleagues; Paul was at Reading, but is now at York via Lancaster.

It's impossible to say exactly how many speakers there are of this accent, but it is common among younger working-class speakers in the London area, and features of the accent have also been observed in other urban areas of the UK.  It is not at all exclusive to speakers from an Afro-Caribbean background but is also spoken extensively but e.g. white and Asian speakers wishing to identify with a certain demographic / social group.

British impressionist and actor Alistair McGowan did a nice piece on MLE for the BBC's One Show, which you can view below.


I would say that the speaker in the clip is probably a UK or fully bilingual speaker of English rather than a second language learner or someone with an indigenised variety of English (e.g., Nigerian English). The speaker probably grew up in or near inner London and has probably been educated in the UK system.  I would be surprised if he was from outside the greater London area, but this is an accent which is socioculturally attractive and so he may be from further afield. I would also suggest that he is lower middle-class rather than working-class as he sounds educated.

It should be noted, however, that we cannot actually see him speak in the film. Most of his face including his mouth is covered.  It could, therefore, be a voice-over.

When I appeared on the BBC News Channel (I'm so sorry I don't have a clip of this to share) I was asked about forensic phonetic analysis of this speaker's voice. What we would need to be able to do this is a reference sample of a known speaker in order to make comparisons between that and the Foley video.  As one of my colleagues, Martin Barry, points out, unless this speaker has spoken into a police microphone it will be almost impossible to carry out forensic speaker comparison successfully.

My picture from the AP session also appeared in the Los Angeles Times.  You can view the online article HERE, which has comment from Martin Barry.

Friday, 6 June 2014

Phonetic vs phonemic inventories

In my first year "Sounds of Language" class, one of the things we do is look at phonetic vs phonemic inventories. I've just had a question about this on the discussion board for the module so I thought I may as well post my response, in case anyone is interested.
Sounds pattern differently in different languages. Speakers of two languages may produce exactly the same set of speech sounds - or phones (phonetics) - when talking, but the languages may use those sounds differently to create meaning. Once we're talking about meaning, we are considering phonemes.
Say, for example, there are two languages whose speakers produce the consonant sounds [p] and [b] and have one vowel, [a]. In both cases, the phonetic inventory contains [p], [b] and [a]. We put the sounds in [] brackets to indicate we are just talking about how the sound is produced at the moment. 
The only thing which is different between [p] and [b] is voicing; [p] is voiceless and [b] is voiced. Otherwise, they are both bilabial plosives.  
In language A, [bapa] and [baba] mean different things - [bapa] means "red" and [baba] means "yellow". [bapa] and [baba] constitute a MINIMAL PAIR, as only one sound differs between the two and it changes the meaning of the word. We can therefore say that /p/ and /b/ are phonemes - meaning units - because of this change in meaning, and we now put them in // brackets.  There are TWO consonant phonemes.  Thus, the phonemic inventory is /p/, /b/ and /a/.
In language B, however, [baba] and [bapa] both mean the same thing - they both mean "car". This means that it doesn't matter whether the consonant is voiced in language B. As there is no change in meaning when one substitutes [p] for [b], they are NOT different phonemes but belong to the same single phoneme. 
What we have to do for language B is decide which sound represents the phoneme, and we often choose the one which occurs in most environments. As we don't have a lot of data here, let's go with the phoneme being /b/ (as there are more of them). That phoneme contains the two sounds [p] and [b], which are ALLOPHONES (phonetic variants) of /b/. Thus, the phonemic inventory is /b/ and /a/. 
This is a very limited set of data, however!

Tuesday, 15 April 2014

The Guess List: a study in /t/ elision

The BBC is airing a new game show on Saturday nights hosted by the wonderful Rob Brydon and amusingly entitled The Guess List. You can read the Independent newspaper's less than glowing review of it here.

Why amusing?

This plays on a phrase, the guest list, which is a list of people invited to an event, i.e., a list of guests. No surprise there.

What amuses me is that it is an example of how the process of alveolar plosive elision can result in homophones in English - in this instance, a homophonic phrase.

When one produces the phrase the guest list in rapid speech, it is normal to leave out the /t/ sound at the end of /ɡest/.  This process is called /t/ (or /d/) elision.  The rules for when this can take place are as follows:
  1. The /t/ or /d/ must be in the syllable coda;
  2. It must be surrounded by other consonants;
  3. The consonant preceding the alveolar plosive must agree in voicing with it - so if the plosive is a /t/ it must be preceded by a voiceless consonant and, if it's /d/, it must be preceded by a voiced one.
  4. The consonant following cannot be /h/.
So, in guest list, which can be transcribed phonemically as /ɡest lɪst/, we can elide the /t/ at the end of /ɡest/ because it meets the requirements listed above.  This results in /ɡes lɪst/, which means guest list and guess list are homophonous.

There is, as far as I know, no such thing as guess list as a phrase in English. If one types it into Google, for example, it redirects you to guest list.

Other notable examples of homophones resulting from connected speech processes include handbag /hændbæɡ/ becoming homophonous with ham bag /hæmbæɡ/. There are two processes going on here: /d/ elision and assimilation.

Assimilation is a process by which sounds at word boundaries - often alveolar consonants - become more similar to each other in rapid speech. Here's a diagram showing consonants at word boundaries:

_ _ Cf | Ci _ _

Cf = final consonant; Ci = initial consonant

In English, we tend to get regressive assimilation, which means the initial consonant (Ci) at the beginning of the next word has a backwards effect on the final consonant (Cf) of the preceding word. As I mentioned above, this tends to affect alveolar consonants, and more often than not it will affect the place of articulation of Cf, i.e., it will not be produced as an alveolar consonant but will have the same place of articulation as the Ci of the next word.

In handbag /hændbaɡ/, the alveolar plosive /d/ is elided and the alveolar nasal /n/ is produced as a bilabial consonant because the following word - bag - begins with a bilabial consonant, /b/.  This results in the production /hæmbæɡ/, which is homophonous with ham bag. But of course, a lady wouldn't normally take a bag of ham out with her when she went shopping, and we can usually retrieve the real meaning from the context.

I should add a caveat in that this is a very brief overview of the theory of these two processes. In very rapid speech, all sorts of sounds get elided and / or assimilated, so analysing spontaneous speech can be a real challenge.

Another issue which arises from connected speech processes such as elision and assimilation is that they can make speech less intelligible or the message more difficult to understand.  Emilio's comment below led me to this example, spoken by Elizabeth Taylor and Richard Burton in Who's Afraid of Virginia Woolf? In it, there is /t/ elision in the word guests, which is perfectly legal. You can see how Burton's character doesn't understand Taylor's character until she repeats the word guests with the /t/ in it - although what "We've got guess" (as opposed to "We've got guests") might mean is difficult to ascertain.  There are obvious issues for speech intelligibility here in English as an international lingua franca.

Watch from 04:04 right near the end.  And thank you, Emilio!



I'd recommend the following books by way of introduction if you are interested in connected speech processes in English:

Lecumberri, M L G & Maidment, J. 2000. English Transcription Course. London: Arnold.

Roach, P. 2009. English Phonetics and Phonology: A practical course. Cambridge: Cambridge University Press.



Thursday, 19 December 2013

Syllable structure matters

You know, I can't remember who I was talking to about this recently or why we got on to the topic, but I have always done my best to educate language teachers (and learners, for that matter) about the importance of syllable structure and phonotactics in learning how to pronounce a new language.

I am, of course, usually coming at it from the angle of someone pronouncing English.

As you doubtless know, accents such as Standard Southern British English (SSBE) have up to three consonants at the start of a syllable (onset consonants) and four consonants at the end (coda consonants), and a syllable usually has a vowel as its peak. The structure of the basic SSBE syllable can therefore be described as follows:

(CCC)V(CCCC)

I'd recommend reading Roach (2009) chapter 8 or Cruttenden (2008) section 5.5 for a full description of what clusters are possible in syllable onsets and codas in SSBE. There's also a nice description on the Macquarie Linguistics pages. The main point I want to make here is that not all languages have syllables which are as complex as English (and English does not have the monopoly on complexity), and this is what can lead to problems with pronunciation as much as not being able to produce a sound.

The thing which always surprises me - and perhaps it shouldn't - is that teachers of English from other language backgrounds often know nothing about the phonology of their own language, and so do not understand that a learner's problem with pronouncing a sound in a particular position in the syllable is unlikely to be about not being able to produce the sound per se but that the learner's language does not permit certain sounds in certain positions in the syllable. If, for example, a learner is from a Chinese language background and that language only permits a zero-coda (i.e., no consonants at the end of syllables) or only a nasal of some description in the coda, pronouncing any other consonant at the end of a syllable may be difficult, and pronouncing clusters is going to be an extreme challenge.

In addition, learners have different strategies for dealing with clusters. Some learners (e.g., Japanese) will insert vowels between consonants in a cluster - this process is known as vowel epenthesis - in order to preserve as many consonants as possible. By comparison, Chinese speakers will often elide consonants in order to be more similar to Chinese syllable structure and number. Here's a favourite comparison of mine: In Japan, MacDonald's, which is /məkˈdɒnəldz/ in SSBE, is known as "ma-ku-do-na-ru-do", but in Hong Kong it is known as "mak-do-nau", with a strongly glottalised and unreleased [k] in the first syllable. Japanese tends to preserve the consonants but Cantonese preserves the number of syllables.

In World Englishes, we often see patterns of syllable structure influenced by a speaker's L1 or the indigenous language(s) of the region in which English has been adopted. This may be why speakers of many varieties of English around the world drop third-person singular "-s"; the meaning of it is retrievable from the context, and it's a rather superfluous inflection which is likely to be dropped anyway in complex codas in many L2 Englishes. 

Does it matter that clusters are simplified? Yes, it does, if intelligibility and therefore meaning is compromised. One is unlikely to be misunderstood if leaving off third-person singular "-s", but it becomes more of a problem in other contexts; my understanding of a Hong Kong English pronunciation of MacDonald's (I'd asked what the student's favourite things were) was that the speaker had said Madonna, thanks largely to the lack of consonants at the end of the word.

What can teachers do about this? First, one needs to be aware of the syllable constraints of the L1 of the learners you are going to be teaching, so you have an idea of whether they are used to complex syllable onsets and codas to start with. If not, chances are learners will be able to produce singleton consonants and some clusters in onset position with little difficulty, but codas are always more problematic.

One strategy, if coda consonants are a problem, is to try to "slide" the coda consonants into the next word; this doesn't always work, but it can also help learners with listening if they can understand that speech is a stream rather than a string of discrete words, and so it may well sound like coda consonants belong to the next word. For example, in a phrase such as "MacDonald's is my favourite", one could slide the final /z/ of MacDonald's into the start of the word "is" and it would then be a little more straightforward for the listener to retrieve.

References:

Cruttenden, A. (ed). (2008). Gimson's Pronunciation of English (7th ed.). London: Hodder Education.

Roach, P. (2009). English Phonetics and Phonology (4th ed.). Cambridge: Cambridge University Press.

Wednesday, 11 December 2013

Flipping phonetics

I am so sorry I've not posted for a while; it's been a hectic term!

One of the reasons it's been hectic is because I've been trying a different method of delivering some of my English phonetics and phonology classes and that - as always - entails preparation which takes TIME ... but time well spent which has been worth it.

I first heard about the flipped classroom from my friend and colleague Dr Patricia Ashby who is now an Emeritus Fellow of the University of Westminster. You may know Patricia from her excellent books Speech Sounds and Understanding Phonetics. She presented "Flipping Phonetics" at the Phonetics Teaching and Learning Conference at UCL in 2011; you can read the paper by clicking HERE, and will notice that the results for the topics Patricia "flipped" were very impressive.

The flipped classroom basically involves presenting what would normally be lecture content via vodcasts which the students watch ahead of the class, thus allowing more time in the actual class itself for practical work. This approach works well in the sciences where a lot of practical work is needed for students to progress, and Patricia had noticed how it was also suitable for phonetics, which also requires a lot of rehearsal of skills and time for class discussion of issues.

I had wanted to try this for a while as I have been becoming increasingly concerned that the growing number of students I have in my class meant that I had less time to spend with each of them and that it was difficult to support individual student needs. Thanks to a small grant from the University of Reading's "Partnerships in Learning and Teaching" (PLanT) pilot scheme, I was able to buy some software to do video capture of my desktop which enables me to record video and audio of me narrating my way through my lecture slides. I then post these on our virtual learning environment, Blackboard, for the students to view ahead of class.

The PLanT scheme also enabled me to work with students to produce materials for the post-exams period at Reading to scaffold first year students' learning in preparation for the English phonetics and phonology module in Year 2. You can read about this HERE and HERE (see p. 79).

One set of the vodcasts is on YouTube and I've posted them below if you'd like to take a look.  We follow Peter Roach's English Phonetics and Phonology on this course, and this class presents material from chapters 15-17. In it you will see some embedded YouTube clips and also the excellent programme RT pitch which is available for download from UCL's wonderful phonetics and speech resources.

I would value feedback on these videos (aside from the fact that I say "so" a lot!) either at the end of this message or on the YouTube pages themselves via my channel (be warned: also contains some videos of one of my bands, Crimson Sky).





Students' responses so far have been very positive. They mention how appreciative they are to have more time in the class to work on practical skills. They also indicate how presenting material this way aids independent learning and allows students to take notes at their own pace, and they can of course return to these videos when it comes to exam revision; our exams are in May/June so there is a lot of time to forget the content. One student has written a blog post of her own about this (lots of other good stuff on there!), and another comments: "Your delivery and humour makes them very interesting and engaging." They have asked for other tutors to adopt this method and I hope some of my staff will consider it.

Last year I was disappointed to see that the average overall marks among undergraduates for the dictation test we do at the end of term had dropped by around 11 percentage points. I usually expect the average to be in the mid-to-low 60s; the previous year's average had been around 66%.  I'm just about to mark the transcription tests this year and will report back on whether they have improved with an update to this post.

UPDATE #1: I've marked 20 out of 59 scripts and can report that the current average is **over 15 percentage points higher***. Watch this space ...

UPDATE #2: Having finished the marking, I can confirm that the average is up over 10 percentage points on last year's dictation test scores. Although not quite as impressive as 15%, it's still pretty darned good!

Friday, 13 September 2013

Intonation and train announcements

This is a post about the intonation of announcements on trains in the UK. 

I didn't actually think it was worth posting on this topic until one of my students on the UCL Summer Course in English Phonetics mentioned that he thought the intonation was odd - and he was talking particularly about a feature which I had noticed and thought amusing.

I was returning to Reading on the South West Trains London Waterloo service one evening when I noticed two things that interested me: first, the company who made the in-train announcements had chosen a falling-rising tone rather than a rising tone for certain functions; and second that the falling-rising tone was used in some unexpected places.

Excuse me while I explain a few things about intonation in standard British English, southern accent. This is taken from a chapter I wrote a while ago (Setter 2005) on Discourse Intonation and adopts that framework (see e.g. Brazil et al. 1980; Brazil 1997). There are four central elements to discourse intonation: tone, key, the tone unit and tonicity. I'm going to focus on tone here; readers with interest in the subject should seek the publications mentioned for a fuller introduction.

Tone refers to the major pitch movement(s) in an utterance. Brazil et al. (1980: 13) distinguish between five tones: falling, rising, falling-rising, rising-falling and level. The falling and falling-rising tones “embody the basic meaningful distinction carried by tone”, whereas the other three “can usefully be seen as marked options, understood and meaningful in contrast” (Brazil et al. 1980: 13).

The following example is given to show the contrast between two utterances using the two basic tones, falling and falling-rising (Brazil et al. 1980: 13); I have used ↘ to indicate falling and ↘↗ to indicate a falling-rising, and // indicates a tone unit boundary (or intonational phrase boundary):

(1)         // when I’ve finished ↘↗Middlemarch // I shall read Adam ↘Bede //

(2)         // when I’ve finished ↘Middlemarch // I shall read Adam ↘↗Bede //

Other meanings notwithstanding, by using the falling-rising tone on Middlemarch and the falling tone on Bede in example (1), the speaker is showing that he/she believes the listener already knows the speaker is reading Middlemarch, but does not know the next book the speaker intends to read is Adam Bede. By contrast, in example (2), it is believed that the intention to read Adam Bede is known, indicated by the falling-rising tone, but not the fact that the speaker is reading Middlemarch at the moment, indicated by the falling tone. The use of specific tones therefore indicates what the speaker believes either to be common ground in any utterance, be it general knowledge of the world or information mentioned earlier in the same piece of discourse or some other context, or unknown – whether information is given or new.

Brazil et al. state that “all interaction proceeds, and can only proceed, on the basis of the existence of a great deal of common ground between participants” (1980: 15). Given information, or common ground, is indicated by what Brazil et al. call “referring” tones, and new information is indicated by “proclaiming” tones (1980: 15). The falling tone is therefore the default proclaiming tone, and is given the symbol p, which is placed at the beginning of the tone unit. The falling-rising tone is default referring tone, and is indicated by the symbol r. The nucleus, referred to as the tonic syllable, is capitalised and underlined; the two examples above could therefore be represented as follows:

(1a)       r when i’ve finished MIDdlemarch // p i shall read adam BEDE //

(2a)       p when i’ve finished MIDdlemarch // r i shall read adam BEDE //

The choice of tone used by a speaker is, then, dependent on the speaker’s evaluation of “the relationship between the message and the audience” (Brazil 1980: 18) – whether the speaker believes information in the message to be given or new. From this point of view, the speaker might be assuming common ground which does not exist, and therefore erroneously using referring tones, or using proclaiming tones where the information is, in fact, already part of the common ground.

The other tones mentioned are rising, rising-falling and level. The rising and rising-falling tones are variants of the r and p tones respectively; the symbol for the rising tone is r+, and for the rising-falling tone, p+. The level tone is symbolised with an o.

An r+ tone is used to reactivate background material. Brazil et al. give the following example (1980: 53):

(3)         Where’s the typewriter?

(3a)       r in the CUPboard // (where it always is)

(3b)       r+ in the CUPboard // (why don’t you ever remember …)

In (3a), the fact of the typewriter being in the cupboard is deemed by the speaker to be “vividly present in the background”, whereas in (3b) the speaker is indicating that the listener has to be reminded of what should be common knowledge.

Use of the r or r+ tone can, therefore, show the relationship between speakers in a conversation. The r+ tone is used by a speaker who is assuming some kind of dominant role in the conversation, and is commonly used by teachers in teacher-student interactions, doctors in doctor-patient interactions, or those giving directions or instructions to someone who (it is assumed) has no prior knowledge. As Brazil et al. point out, a patient who starts a doctor-patient interaction with an r+ tone will sound rather aggressive (4); an r tone is usually used (5) (examples from Brazil et al. 1980: 16 & 54). 

(4)         r+ i’ve COME to SEE you // p with the RASH // r+ i’ve GOT on my CHIN //

(5)         r  i’ve COME to SEE you // p with the RASH // r i’ve GOT on my CHIN //

(Where there are other stressed syllables preceding the tonic syllable, these are capitalised but not underlined in this system.)

The p+ tone (rising-falling), like the p tone, is used to indicate that the information is new, but with the additional meaning of being surprising, disappointing or horrifying to the speaker also – the speaker is adding “to his own store of knowledge” (Brazil et al. 1980: 56). It is noted that the p+ tone tends to be used by a dominant speaker.

The level tone, symbolised o and referred to as the “oblique” tone, is used to indicate that the speaker considers he/she has not arrived at the potential completion point of an utterance (Brazil et al. 1980: 88), but it can also show that the speaker is not very involved in, e.g., reading a passage.

OK, that's the end of the section from Setter (2005). Are you still with me?

On South West Trains, some of the announcements are something like the following:

(6)         This is the South West Trains service from London Waterloo to Reading, calling at Vauxhall, Clapham Junction, Putney, Richmond, Twickenham, Hounslow, Feltham, Ashford, Staines, Egham, Virginia Water, Longcross, Sunningdale, Ascot, Martin's Heron, Bracknell, Wokingham, Winnersh, Winnersh Triangle, Earley and Reading.

(7)         The next station is Sunningdale.

(8)         This station is Sunningdale.  The next station is Ascot.

These announcements are clearly made up of "slots" - e.g.:

(6a)       This is the (slot 1) service from (slot 2) to (slot 3), calling at (slots 4, 5, 6 ...) .... and (slot 7).

(7a)       The next station is (slot 1).

(8a)       This station is (slot 1).  The next station is (slot 2).

In order to do this, the company producing the announcements has to have some kind of idea of how intonation works. Among other things, we are dealing with a list in (6) and (6a), so some slots will have an intonation pattern which indicates the speaker has not got to the end of the list, requiring a referring (r) tone of some kind - i.e., slots 2, 4, 5 and 6 in (6a) - and a pattern which indicates the end of a list, requiring a proclaiming (p) tone of some kind - i.e., slots 3 and 7 in (6a). The company has therefore recorded two versions of each town/city at which the train stops, one with an r tone and one with a p tone. In (7a), the p tone is used in slot 1 as this is a statement.

The intonation patterns are as follows:

(6b)         r this is the SOUTH west TRAINS service // r from LONdon waterLOO // p to READing // r calling at VAUXhall // r CLAPham JUNCtion // r PUTney // r RICHmond // r TWICKenham // r HOUNSlow // r FELtham // r ASHford // r STAINES // r EGham // r virGINia WATer // r LONGcross //  r SUNningdale // r AScot // r  MARtin's HERon // r BRACKnell // r WOkingham // r WINnersh // r WINnersh TRIangle // r EARley // p and READing //

(7b)         r the NEXT station // p is SUNningdale //

(8b)         r THIS station // r is SUNningdale //  r the NEXT station // p is AScot //

Can you spot the things which amuse me?

First, as the announcement being made is authoritative, I would expect the referring tone in the list to be a rising tone (r+) rather than falling-rising (r). This was what the very observant non-native English-speaking student asked me about in class this year. One could argue that commuters take this train every day and so the information is already "vividly present" in the background somehow ("this train always stops at these stations and you know it" - see 3a above), but I'm going to dismiss that.

Second, in (8b), the company who selects which spoken version of the town/city goes into which slot has chosen an r tone for "Sunningdale" - i.e., (8a) slot 1. I assume this is because there is another town/city about to be mentioned later in the announcement (slot 2 - this correctly has a p tone on it) and so the company sees it a type of list. However, whenever I hear this it makes me laugh, because using an r tone here makes it sounds like a surprise that one has arrived in e.g. Sunningdale.

(9)         This station is Sunningdale  ..??
              ... What??? I was expecting Longcross! I must have fallen asleep! Blast!!

What should it be on "Sunningdale" in slot 1 (8a)? A falling tone (p), of course ("this is definitely Sunningdale and you don't have to be in any doubt about it").

So, next time you are on trains in the UK, listen to see what intonation patterns are being used. Has the automated system of putting things in slots in announcements worked? Do let me know!

(Don't get me started on nucleus placement on prepositions ...)

References

Brazil, D. (1997). The Communicative Value of Intonation in English. Cambridge: Cambridge University Press.

Brazil, D., Coulthard, M., and Johns, C. (1980). Discourse Intonation and Language Teaching. Harlow: Longman.

Setter, J. (2005). Communicative patterns of intonation in L2 English teaching and learning: the impact of discourse approaches. In K. Dziubalska-Kołaczyk & J. Przedlacka (Eds.), English Pronunciation Models: a changing scene, Bern: Peter Lang, pp. 367-389.

Thursday, 29 August 2013

The International Phonetic Alphabet

I was recently asked to contribute some history and other information on the IPA chart for Babel Magazine's fourth issue. The text below is the unedited version of what appears, with references. I have been given permission to post this to my blog.
--
The International Phonetic Alphabet represents the sounds of all the world’s documented languages. The first published version of this chart can be found in Passy (1888) and appeared in a journal called The Phonetic Teacher (or Dhi Fonètik Tîtcer). This speaks a lot to its origins, as the International Phonetic Association (IPA) itself was inaugurated as the Phonetic Teacher’s Association in 1886 which was mainly involved with teaching English (IPA 1949, p. 2 of cover). During the first two years of the association, and as more than one script was in use, it was decided by its members to establish a single alphabet which could be applied to the description of all languages. Since Paul Passy’s publication of the first chart in 1888, the association has worked tirelessly to improve the alphabet and there have been several published revisions. The alphabet itself is “on romanic basis” (IPA 1949, p. 1), meaning it uses a script which derives from Roman characters rather than e.g. Cyrillic (Russian), Arabic or other written traditions, and is presented by the Association as “a consistent way of representing the sounds of language in a written form” (IPA 1999, p. 3).
The International Phonetic Alphabet (2005 revision)*
The current revision of the IPA chart (above) starts with a large table showing consonant sounds, or phones, made on a pulmonic egressive airstream (i.e., with air from the lungs). Place of articulation (POA) is indicated by which column a symbol is located in. The passive articulator is usually indicated, i.e., the part of the oral cavity which remains in place while the active articulator – often the tongue – moves towards it; e.g., if a sound is labelled “alveolar” it means the tongue moves towards the alveolar ridge. Manner of articulation (MOA) is indicated by row. Where voiceless and voiced pairs of consonants are given, the one on the left is voiceless. The usual way of describing a consonant is to use a VPM label, where VPM stands for “voice place manner” – so [t] is a voiceless alveolar plosive.
This table is followed by non-pulmonic sounds – clicks, implosives and ejectives – below and to the left, with the vowel chart to the right. The vowel chart represents cardinal values for vowels which can be used as a reference to describe vowel sounds in languages. Symbols are placed on a trapezium which represents the vowel space; this space is in fact very small, with “front” vowels being articulated with the front of the tongue raised to various degrees in proximity to the hard palate, and “back” vowels involving the back of the tongue being raised in proximity to the velum. It is usual to describe vowels in terms of height, backness/frontness and lip rounding – so e.g. [i] is a close front unrounded vowel.
Beneath the non-pulmonic sounds are “other”consonantal symbols; these are ones which do not appear in the main chart because they have two POAs, two MOAs, or cannot be otherwise accommodated.  E.g., [w] has both lip rounding and a movement of the back of the tongue towards the soft palate, so is labial-velar and therefore classed as a double articulation; [t͜s] is an affricate, which involves a plosive followed by a fricative both produced at the same POA (i.e., homorganic) with the same voicing; alveolo-palatal fricatives [ɕ] and [ʑ] are not on the main consonant chart because that region simply cannot accommodate any further symbols.
Beneath this list is a table of diacritics which allow the sounds on the chart to be modified further. For example, the symbol [t] can be modified to represent a voiceless dental plosive by adding the dental diacritic to give [ t̪ ].
Under the vowel chart is a list of symbols for representing suprasegmental information, i.e., phonetic features above the level of individual speech sounds.  Here we can find stress marks, length marks, syllable-division marks and suprasegmental boundary markers and, below these, symbols for tones and word accents.
Although the alphabet largely achieves what it sets out to do, there are a number of issues which arise. One is simply to do with what people understand this chart to mean.
Note that square brackets are placed around each of these symbols: [ ]. This indicates that the symbol is representing the production/articulation of a given sound and not that it is a linguistic unit or a phoneme belonging to any particular language. In transcribing languages phonemically, one uses (or “borrows”) a small subset of the symbols on the chart to represent distinctive linguistic units in a language. To show the symbols are used as phonemes in a given language and not as phones or allophones, which are representations of the articulation of a sound, we use slash brackets: / /. In training people to use the chart, it has to be understood that producing a phone such as [p] is not going to be the same as, e.g., an English /p/, which is usually aspirated when at the start of a syllable, even though we have used the same symbol. If your phonetics teacher wants to hear a typical English /p/ sound, it will be represented [pʰ] in phonetic transcription.
Another issue is to do with processes for updating or revising the alphabet. There was general rejoicing in the phonetics community when, in 2005, the symbol for the voiced labiodental flap [] was added to the main consonants chart; this was the first revision since 1996. But it is not a matter of someone simply proposing a new sound; evidence must be given for the existence of the sound as a distinctive unit in a given language. More recently, there has been discussion about whether the vowel chart should have a symbol for an unrounded open central vowel as found in German, proposed in Barry and Trouvain (2008), which would be represented with a small capital A [a]. This proposal had in fact been first discussed by the association in 1989 and rejected at that time; Barry and Trouvain (2008) provided a convincing rationale for reconsidering this position, including the fact that there are central vowels at every cardinal tongue height except for in open position, and that there are languages which have this vowel quality in a stable enough way to support representing it on the chart with a single symbol not requiring diacritics (it can, for example, be represented as [ɐ̞], which is the symbol for the most open unrounded central vowel currently on the chart plus the “lowered” diacritic). In December 2011 the matter was finally decided by the IPA council voting against adopting [a] (IPA 2012, p. 245).
Something else the basic alphabet fails to cater for is many of the sounds produced by atypical speakers – that is, people with speech deficits – or children in the developmental stages of sound production. For example, some speakers may produce sounds classified as labioalveolar, which means the speaker’s bottom lip touches the alveolar ridge behind the teeth (amongst typical speakers, the furthest back the bottom lip travels is to meet the upper teeth in labiodental sounds such as [f] and [v]). In order to deal with this, an additional chart known as “the Ext-IPA symbols for disordered speech” was devised, where “Ext” is abbreviated from “extensions” (Duckworth, Allen, Hardcastle & Ball 1990; PRDS Group 1983). This currently exists in its 1997 revision in the IPA handbook (1999, p. 193). For information, labioalveolar sounds are represented by a double-underline diacritic combined with a bilabial or labiodental symbol, so a voiced labioalveolar nasal is [m͇] and a voiceless labioalveolar fricative is [ f͇ ].
And yes, it is possible to describe every sound produced with human vocal apparatus with phonetic terminology. Did you know, for example, that when you “blow a raspberry”, you are performing a voiceless linguolabial trill?  Or that a “gee-up” noise to encourage a horse is a voiceless alveolar lateral click?  Well, now you do.

*IPA Chart, http://www.internationalphoneticassociation.org/content/ipa-chart, available under a Creative Commons Attribution-Sharealike 3.0 Unported License. Copyright © 2015 International Phonetic Association.

References
Barry, W. J. & Trouvain, J. (2008). Do we need a symbol for a central open vowel? Journal of the International Phonetic Association 38(3), pp. 349-357.

Duckworth, M., Allen, G., Hardcastle, W. & Ball, M. J. (1990). Extensions to the International Phonetic Alphabet for the transcription of atypical speech. Clinical Linguistics and Phonetics 4, pp. 273-280.

International Phonetic Association, The. (1949). The principles of the International Phonetic Association. London: The International Phonetic Association.

International Phonetic Association, The. (1999). The handbook of the International Phonetic Association. Cambridge: Cambridge University Press.

International Phonetic Association, The. (2012). IPA council votes against new IPA symbol. Journal of the International Phonetic Association 42(2), p. 245.

Passy, Paul. (1888). Our revised alphabet. The Phonetic Teacher, pp. 57-60.

PDRS Group (1983). The phonetic representation of disordered speech: Final report. London: The King’s Fund.