Wednesday, August 22, 2007

Defenders of the pronoun

Recently, for whatever reason, I've been in situations where the question of the status of pronouns has come up. Inevitably, there are people who vehemently reject the idea that pronouns are just a special kind of noun. One correspondent writes,
"Why insist on pronouns as a 'special case of nouns' when current handbooks from Hacker to Troyka and Ready Reference unswervingly place pronouns in a different category from nouns?"
Well, because they generally:
  1. signify the same range of concepts
  2. both are subject to distinctions of case
  3. both are subject to distinctions of gender
  4. both are subject to distinctions of number
  5. share almost all the same functions (e.g., subject, object, determiner)
  6. share the same set of modifiers

Actually, 6. is a bit of a stretch. Typically pronouns don't license determinatives or adjectives but they sometimes can in a pinch (e.g., the new you). Then again, proper nouns don't usually license determinatives or adjectives either and nobody wants to set them off on their own.

These defenders of the pronoun inevitably argue that if you look up "parts of speech" in any reference, you will be told that pronoun is one. It has always been so, they say pointing to the etymology (and falling foul of the etymological fallacy). But they offer nothing beyond tradition to explain why pronouns should get a class all to themselves. Nor can they explain why, if pronouns should, auxiliary verbs, for example, shouldn't.

These people are typically unsurprised that the physics, biology, and chemistry they studied in high school is no longer up to date. But they get positively defensive when somebody suggests that grammatical description has moved on. Why is that?

Saturday, August 18, 2007

Grammar by gosh and by golly

I have been assigned to teach a freshman writing course using Writing by Choice, by Eric Henderson. I've read most of it now and do appreciate the main theme about choice. I was, however, rather surprised when I came to the grammar section. I understand the need to keep things simple, but it seems to me that, perhaps in his efforts to do so, he has developed a basic framework which has rather too many inconsistencies and outright errors.

We begin with his discussion of "substantives". Though tradition is to recognise pronouns and nouns as separate "parts of speech" there appears to be no good reason for doing so. It seems more parsimonious to simply note that pronouns are a special case of nouns. Henderson follows tradition.

Regardless of what you think of the above point, the definition he provides for each remains problematic because of his reliance on semantic properties to the exclusion of morphology and syntax. To wit,
  • "Noun (nomen: 'name'): name of a person, place, or thing."
  • "Pronoun (pro + nomen: 'in place of the noun'): a word that takes the place of a noun in a sentence."

A punch, for example, is in no natural way a thing. It is an action (and I'm sure you can guess what Henderson's definition of verb is), yet punch can be either noun or verb. He goes on to say that the noun that is replaced by the pronoun is its antecedent and that indefinite pronouns have no antecedent. In other words, they don't take the place of a noun. This means they don't meet the defining criteria for pronouns.

The same thing is true of the "demonstrative pronouns" (which modern grammar deals with much more effectively under the category determiner--or determinative if you take determiner to be a function.) What noun does yonder, for instance, replace? And on p. 371, he notes that all pronouns must match the person of the antecedent, and writes, "all nouns are third-person." If we take the definition to be true, this negates the possibility of first and second person pronouns.

The simple fact that nouns (pronouns included) change for number and in forming possessives can easily be incorporated into a definition. Similarly, verbs conjugate for past tense (must excepted). Other properties can also be brought to bear to increase the level of precision (but at the expense of simplicity).

Another problem lies in Henderson's blurry distinction between parts of speech and functions. Nouns, he writes, have five functions: subject, object, object of a preposition, subjective complement, and appositive. Yet, when a noun modifies another noun, he says it is functioning as an adjective. Why is this not in the list of functions? And when did adjective become a function? In this book it is described as a part of speech, a category, that has the function of modifier. So would it not be more consistent to say that nouns functions as modifiers?

The same can be said of the case on p. 320 where phrases are described as acting as nouns. Rephrasing this as "phrases can function as subjects or objects" makes for a much more coherent grammar.

And, before leaving nouns, as far as I can tell, his system completely overlooks their function in cases like night in "I met her last night" or day in "six days old".

The next issue that caught my attention was the treatment of subject. Although the definition on p. 349 ("The subject of a sentence or clause is the noun or pronoun that performs the action of the verb, or that exists in the state or condition expressed by the subjective complement.") is better than the one on p. 307 ("The subject noun is the doer or performer of the action."), neither of them sufficiently handles stative transitive verbs such as 'have', which denote no action and have no "subjective complement". This is similar to the issues raise above regarding the definition of nouns.

But a more serious problem is that the definitions completely ignore passive sentences. In such sentences, the subject is never, as far as I can discern, the "doer or performer of the action", regardless of how forgivingly you interpret action. (In fact, insofar as I can see, the book ignores passive sentences altogether; at least, there is no index entry for passive or voice and I haven't come across any mention in the text.) And while we're on the subject of subjects, he writes, "prepositions cannot ever be the subject of a clause". This denies the existence of sentences such as "After nine is good for me."

Shifting now to verbs, we have the same kind of defining problems again, but those aside, I was shocked to see that he has conflated intransitive verbs and "linking verbs". While there are good reasons for including linking verbs as a special case of intransitives, I can see no basis at all for calling all intransitives linking verbs. Yet, that is what he's done. I've yet to see another grammar that sanctions this grouping. His choice to ignore other valencies (ditransitives or complex transitives) can be dismissed as an issue of scope, but this choice is much harder to defend.

Then he completely loses it with this bit of analysis. Considering the sentence "He acted splendidly as Hamlet in Shakespeare's play", he writes, "acted is used as a transitive verb--there is a direct object of this activity: 'as [in the role of] Hamlet.'" Direct object is not among the functions he lists for prepositions (and rightfully so, I believe). Indeed, Merriam-Webster's online dictionary gives the following example for INtransitive act: "trees acting as a windbreak".

Seemingly by way of explanation, during this discussion of intransitive/linking verbs, he makes the assertion that "in some verbs, the traditional use of to be has been dropped in speech and informal prose," giving the example of seem (e.g., "She seems [to be] well.") As far as I can tell, seem has been used freely without to be as far back as Chaucer's middle English and likely all the way back to Old Norse. For example, from The Canterbury Tales

The vapour, which that fro the erthe glood,
Made the sonne to seme rody and brood;
But natheless, it was so fair a sighte
That it made alle hir hertes for to lighte,
What for the sesoun and the morwenynge,
And for the foweles that she herde synge;

There are other problems, but I've gone on too long already. By the way, this book was first published last year. What justified it, I have no idea.

Thursday, August 16, 2007

big size(d) question

Over on the ETJ list, Joshua Myerson wrote to ask about the differences between pairs such as short sleeve(d) and large size(d).

It's an interesting question and I'm not sure exactly how to approach it. Here are some other examples:

  • long leg(ged)
  • large size(d)
  • oval shape(d)
  • broad shoulder(ed)
  • pencil neck(ed)
  • open neck(ed)
  • small frame(d)
  • small size(d)
  • flat bottom(ed)
  • low ceiling(ed)
  • different colour(ed)
  • middle age(d)
  • double strand(ed)
  • white stripe(d)
  • high power(ed)
  • light colour(ed)
  • good size(d)
  • dark hair(ed)

Note that these are often interchangeable (e.g., low-ceiling(ed) house) with no difference in meaning. But consider the difference between a dark hair gene and a dark haired gene.

While it's commonly thought that only adjectives modify nouns, nouns can also modify nouns (e.g., faculty office), Thus, we can look at the group in which the second constituent is a noun as noun phrases (NPs) that modify other NPs. This isn't problematic.

The -ed group, though is rather harder to deal with. The -ed word may be an adjective or a verb. Either way, it's being modified by an adjective which is something I wasn't aware could happen. In something like long legged, we can use pronunciation to help us decide that leg' ed (two syllables) is an adjective where legged (one syllable) is a verb. This approach is rather limited though.

Another thing to noticed is that while some would be fine without the adjective (e.g., _ power(ed) tools) others make little sense at all without the adjective (e.g. *a _ bottom(ed) boat), though the -ed forms tend to work better here (e.g., a _ sized shirt vs. *a _ size shirt).

Hmmm...

Tuesday, August 14, 2007

The Elements of Typographical Style

Almost a year ago I posted about The Solid Form of Language by Robert Bringhurst. The Elements of Typographical Style is by far the most famous of his books, and I had long been interested in reading it. My wife is now taking courses in graphic design, which gave me a perfect excuse to buy and read the book. After all, what dedicated husband wouldn't want to find out more about his wife's fields of interest?

Bringhust is a renowned Canadian poet, and it shows in his prose. Though the title portends little but fussiness, pedantry, detail and drudgery, the book is actually a delight to read. Here are a few samples (not necessarily the best):
  • on typography "Like oratory, music, dance, calligraphy - like anything that lends its grace to language - typography is an art that can be deliberately misused. It is a craft by which the meanings of a text (or its absence of meaning) can be clarified, honored and shared, or knowingly disguised.
    "In a world rife with unsolicited messages, typography must often draw attention to itself before it will be read. Yet in order to be read, it must relinquish the attention it has drawn."

  • on case "The union of uppercase and lowercase roman letters - in which the upper case has seniority but the lower case has the power - has held firm for twelve centuries. This constitutional monarchy of the alphabet is one of the most durable of European cultural institutions."

  • on notes "Relegating notes to the foot of the page or the end of the book is a mirror of Victorian social and domestic practice, in which the kitchen was kept out of sight and the servants were kept below stairs."

  • on proportions "The proportions of a page are like an interval in music. In a given context, some are consonant, others dissonant. Some are familiar; some are also inescapable, because of their presence in the structures of the of natural as well as the man-made world."

Monday, August 13, 2007

Pilobolus

Over on Small Things Considered, there's a wonderful description of the fruiting body of the fungus pilobolus.

(Via The Loom.)

Friday, August 10, 2007

White noise

A few weeks ago, I doubted a NewScientist interview in which linguist Annie Mollard-Desfour makes the claim that to a Japanese person the brightness of a colour is more important than its hue and that the Japanese language has a large number of words for white, "from the dullest to the most brilliant". In the August 4 issue, Mollard-Desfour responds.

"The importance accorded in Japanese culture to matte-gloss and brightness distinctions is mentioned in numerous linguistics papers dealing with the cultural aspects of language and of naming - as are these features of Inuit language. True, some linguists currently propose that we need to distinguish terms for an abstract "true white", that does not refer to any particular instance of "whiteness", from those that refer to materials such as snow. Japanese certainly has terms for "white in general" and others linked to particular instances of whiteness: www.edicojaponais.com and www.dictionnaire-japonais.com.

It remains the case that the lexicon of colours is difficult to understand and to translate, because the parameters used may be fundamentally different. Hence the controversies: what words translate the French blanc - and are the whites of snow or other bearers of whiteness true colour terms?"

The dictionaries to which she links return the following results when you search for blanc.

  1. *白 white (noun)
  2. *白い white (adj)
  3. *ホワイト white (Japonification of the English word white, used in brands etc.)
  4. *ブランク white (Japonification of the French word blanc, used in brands etc.)
  5. 修正液 white out correcting fluid
  6. 卵の白身 the white of an egg (as distinct from the yoke)
  7. 白目 the white of an eye (as distinct from the iris & pupil)
  8. 笹身 white meat (of a chicken)
  9. 白樺 white birch
  10. 白馬 white horse
  11. 白黒 black and white
  12. 白血球 white blood cell
  13. 白鷺 white heron
  14. 白熊 polar bear
  15. 白紙 white (i.e., blank) paper
  16. 白米 white rice
  17. *真っ白 pure white (opposite of pitch black)
  18. 白髪 gray hair
  19. *灰白 gray (ash white)

The items with an asterisk are the only ones that are actually colour terms, the others merely denoting white things. If you then search for 白 and include only the colour terms, you can add the following to the list:

  1. 青白 pale; green (said of a person who is feeling queasy or shocked; literally blue white)

So there you are. If we stretch it, we can find 7 words for white. And they don't exactly describe a continuum from the dullest to the most brilliant.

Now we can start looking for other words that signify white. Parchment, for example is 灰味黄 (ashy yellow; literally ash flavour yellow), and pearl comes out as 真珠色 (pearl colour), ivory is 象牙色 (elephant tusk colour), etc. If you're really keen, you can have a look at a list of Japanese colour names by kana order here (in Japanese).

I put the issue to Language Logger Bill Poser. He writes.

"That is very curious. My reaction is the same as yours...

I wonder if Mollard-Desfour has got hold of a warped idea about the classical color terms, the ones like imayauiro "red" found in beautiful charts in the endpapers of classical dictionaries, used, as far as I know, basically for describing the colors of kimono in Genji Monogatari? Some of those distinctions might be described in terms of brilliance, but the system still doesn't have multiple types of white."

By they way, I can't find anything on Google scholar about "the importance accorded in Japanese culture to matte-gloss and brightness distinctions" but maybe I'm not looking in the right place or perhaps it's all in French.

[I just got my print-edition of NewScientist, and despite the tag "From issue 2615 of New Scientist magazine, 04 August 2007, page 21" on the web, my letter and Mollard-Desfour's response are not printed]

Thursday, August 09, 2007

Defining words (or not)

Another new journal has cropped up: ELR Journal. ELR stands for empirical language research.

In the inaugural issue, Yasunori Nishima at the University of Birmingham (Google has apparently never heard of him) has a paper entitled "A Corpus-Driven Approach to Genre Analysis: The Reinvestigation of Academic, Newspaper and Literary Texts". The paper really seems more like that of a student learning to use the tools of the trade rather than a real contribution to the field. It's mostly a rehash of existing work with different corpora, none of it done in a way that is particularly novel, and none of it really challenging any existing results or even testing out anything questionable.

Much of the paper is built around word counts and frequencies. Surprisingly though, Nishima doesn't find it useful to define for us what he means by word. When he discusses the most frequent words, does he count run (e.g., I went for a run) and run (e.g., You run well) as two instances of one word or as two distinct words? What about if we add in runner, running, ran, runs, etc.? Who knows? Nishima says he got his frequency information from Adam Kilgarriff (though there is no proper citation). Presumably, he means he got them here, but did he use the lemmatised list or the unlemmatised one? He doesn't say. At least we can guess he is either counting unique word forms or lemmas.

But then he seems to conflate two senses of word when he compares his frequency data to arguments made by Paul Nation. As I have discussed before, Nation feels that we should consider only the most frequent 2000 word families to be high-frequency. Note, however, that Nation is explicitly talking about word families, while as far as I can tell Nishima isn't.

Though Nishima is merely the most recent linguist to ignore this terminological conundrum, his is a rather flagrant and troubling oversight, largely because the paper doesn't even show an awareness of the issue. How could you be doing research with words and not give a second thought to what a word is?

Not an auspicious start for the journal.

Wednesday, August 08, 2007

Quaint, so quiet

I was reading The Mouse and the Motorcycle, by Beverly Cleary to my kid when I came across the following sentence.
“Matt, who had seen guests come and go for many years, knew there were two kinds—those who thought the hotel was a dreadful old barn of a place and those who thought it charming and quaint, so quiet and restful.”

I'm assuming here that so is not used in its sense as an intensifier. This threw me for a bit of a loop because I didn't think that so could be used to connect anything below the level of a clause. In fact, that was one of the properties that I thought distinguished coordinators from conjunctive adverbs, the topic with which this blog started over a year ago.

It strikes me as a little odd, but the more I look at it, the less objectionable it seems. What do you think?

Tuesday, August 07, 2007

10,000th visitor

According to sitemeter.com, somewhere in the wee hours of August 6, English, Jack had its 10,000th visitor. Yeah!

Word drive successful

The word drive at the Simple English Wiktionary that I announced a month ago has been successful. We reached our 2000-word target a few days before the Aug. 4 deadline. Then an editor deleted a number of spurious entries which dropped us down below the 2000 mark again. By August 4th, however, we were back above it.

This seems to have created some momentum which I hope will continue. I would encourage anyone else with any interest to contribute. Also, please point it out to ESL learners you may know.

Sunday, August 05, 2007

Word spurt or gradual acceleration

Bob McMurray was kind enough to respond to my post the other day. The relevant part of his mail is reproduced below.

I've seen Paul Bloom's book, as well as a number of book chapters containing similar arguments. I don't disagree with him at all. I think his points are two-fold. First, there's nothing sudden or stage-like about the vocabulary explosion--rather, it represents smooth, continuous acceleration. This was elegantly demonstrated empirically by Ganger & Brent (2004, Developmental Psychology). Second, the major acceleration may be occurring late. But the [smaller] gains made by children in their second year are particularly noticeable given that they are starting from nothing.

That said, he doesn't offer an explanation for why we see acceleration at all. Moreover, these two points are all perfectly consistent with my model. The model doesn't make strong predictions about when the acceleration occurs. In fact, if you examine its rate of acquisition after the first 50 words (analogous to an 18 m.o.) it's a lot lower than it is after 2000 words (maybe a 3 year old? I'm not sure). What it does show is that acceleration is a guaranteed result in any parallel-learning situation. I think any system in which growth is the integral of a Gaussian distribution of difficulty will actually show faster learning much later than late infancy. I'm certainly not arguing that the acceleration we see at 18 m.o. (and that is apparent in your own numbers) is the top speed of the learning-system.

I think what's important here is that the model offers an explanation for acceleration at all. It simply shows that two commonly held assumptions (parallelism and variation in difficulty), when implemented, can have surprising results. They may be all you need to account for acceleration.

Actually, Bloom does suggest a variety of possible explanations including neurological changes, accumulation of adequate phonological knowledge, increases in memory, increased understanding of kinds and individuals, emergence of theory of mind, increased use of syntax, and exposure to an increasing number of words as children begin to read. Perhaps I'll ask him for his thoughts on McMurray's paper.

Murray doesn't nod

I recently finished The Meaning of Everything: The Story of the Oxford English Dictionary by Simon Winchester and quite enjoyed it. The story and characters are wonderfully quirky and heroic. Winchester does go slightly hyperbolic in his praise, especially of the English language and the English people of the time. And he has the odd habit of using an interesting or unusual word or turn of phrase and then recycling it a few chapters later, but this doesn't detract much from the enjoyment.

What did annoy somewhat is the unneeded sic on p. 200. There is a quote from a letter that James Murry wrote to Dr. William Chester Minor, one of the most significant volunteer contributors to the OED and a man with serious psychological problems which propelled him to murder.
"The supreme position ... is certainly held by Dr. W. C. Minor of Broadmoor, who during the past two years has sine in no less [sic] than 12,000 quots."

Winchester includes the footnote "Even Home nods". The suggestion is that Murry should have used fewer rather than less. Obviously, Winchester didn't check the words in the OED. According to the Merriam Webster Dictionary of English Usage,

"the OED shows that less has been used of countables since the time of King Alfred the Great -- he used it that way in one of his own translations from Latin -- more than a thousand years ago (in about 888). So essentially less has been used of countables in English for just about as long as there has been a written English language."

Saturday, August 04, 2007

Vocab spurt explained?

Bob McMurray has an article in Science that has been picked up in the popular press, "Defusing the childhood vocabulary explosion." He's also put up his own explanation here.

For years, psychologists have argued that since the speed of vocabulary learning increases dramatically at a certain age (somewhere around 18 months), it must mean that there is a fundamental change in the learning process, a shift in strategy perhaps.

I haven't read the Science article, but on his web page and on the various media accounts, this change in learning is said to be accounted for by differences in the input rather than differences in the processing. McMurray shows that word difficulty/frequency can account for the change in learning speed. Basically, there are only a few very easy words and once you get past them, there are more and more words at that level of difficulty/frequency. This is sort of the upside of what I described here.

It's certainly an interesting result. But there's one assumption that remains unquestioned. Do children really have a learning spurt at this age? Paul Bloom thinks not. Citing various studies in his book How Children Learn the Meanings of Words, he produces the following table on p. 44 (you can search inside on Amazon.com):
12 months to 16 months: 0.3 words per day
16 months to 23 months: 0.8 words per day
23 months to 30 months: 1.6 words per day
30 months to 6 years: 3.6 words per day
6 years to 8 years: 6.6 words per day
8 years to 10 years: 12.1 words per day

In other words, children are gradually increasing their word-learning rate at least until the age of 10. It seems likely that the early changes are quite visible to us because we can keep track of which words they know and we readily notice new words. As the stock of words grows, it becomes much harder to do this.

That's not to say that McMurray's results are wrong. Not having read the paper, I can't really say. But it is another reason to believe that there is no particular change in the learning process that happens early on.

Monday, July 23, 2007

Funding Funding Funding

, writing in The Toronto Star points out the difference between the comparably well-funded federal LINC (Language Instruction for Newcomers to Canada) program and provincially funded ESL programs in Ontario.

Eric Bakovic over at Language Log is also talking up the issue, but in the U.S.

ESL speakers too lazy to learn a third language

The Economist has a bit about the consequences of the dominance of English in Europe. One point I'd never considered before was that people are not learning other languages.
"The rush to learn English can sometimes hurt business by making it harder to find any staff who are willing to master less glamorous European languages.

English is all very well for globe-spanning deals, suggests Hugo Baetens Beardsmore, a Belgian academic and adviser on language policy to the European Commission. But across much of the continent, firms do the bulk of their business with their neighbours. Dutch firms need delivery drivers who can speak German to customers, and vice versa. Belgium itself is a country divided between people who speak Dutch (Flemish) and French. A local plumber needs both to find the cheapest suppliers, or to land jobs in nearby France and the Netherlands."

But what do you do to avoid this problem? Apparently, there is research "by the European Commission suggesting that this risk can be avoided if school pupils are taught English as a third tongue after something else." Given the spectacular failure of most high schools worldwide to teach a second language, I wonder at the practicality of this solution.

Sunday, July 22, 2007

My spammy blog

Yesterday, when I tried to post, I got the following message: "Blogger's spam-prevention robots have detected that your blog has characteristics of a spam blog."

Hmm....

Judy Sierra

Yesterday, my mother was reading Thelonius Monster's sky-high fly pie: a revolting rhyme to my kids while I washed the dishes. Something clicked as she read,
"THELONIUS urgently
e-mailed a spider.
He wanted advice from a savvy insider.
"

"Who wrote that?" I asked. Sure enough, it was Judy Sierra, winner of the 2005 E. B. White Read Aloud Award for Wild About Books. It's somewhat astounding how one simple line can be so characteristic that you immediately know who has written it.

Saturday, July 21, 2007

Ontario: more ESL regs; no teeth, no new funding

It appears that the province will require schools to improve orientation, testing, and reporting for ESL students and their parents. No new funding accompanies the announcement. Nor will their be any requirements that schools actually spend ESL funding on ESL. Previous discussion of the issue is here.

Friday, July 20, 2007

Learning the Language

I just discovered a language-learning-related blog attached to the Education Week website. It's called Learning the Language. The author, Mary Ann Zehr,
"is an assistant editor at Education Week. She has written about the schooling of English-language learners for more than seven years and understands through her own experience of studying Spanish that it takes a long time to learn another language well. Her blog will tackle difficult policy questions, explore learning innovations, and share stories about different cultural groups on her beat."

The blog has existed since February. It's fairly US-centric and is focussed mostly on policy issues though the posts do everything from introducing new materials to visiting individual classrooms.

Thursday, July 19, 2007

Even where there are ESL teachers...

Samuel Freedman reports in the New York Times on a frustrating situation in which the few ESL teachers are pulled away from teaching ESL by paperwork and other tasks. Many of these teachers

"were responsible for completing more than a dozen different forms, evaluations, assessments and reports that came variously from the levels of district, city, state and federal government, and grading standardized tests.

Teachers like Ms. Rabenau were also repeatedly conscripted within their schools to substitute for absent colleagues, to proctor exams in other classes and to chaperon field trips."

Wednesday, July 18, 2007

Idioms: interpreting the frequencies

I suppose this isn't really specific to idioms, it would apply to any vocabulary item.
As I wrote before, one response to my explanation about idioms was,
"None of the correspondents have suggested that they have any difficulty recognising or understanding 'hit the jackpot', yet the low level of occurrence of the expression in corpora suggests that it should be so unfamiliar as to cause difficulty even to native speakers."
I'm afraid this doesn't show up a problem with the corpora themselves, but it might go some way to explaining why language teachers seem to be so loath to use corpus data: they don't understand what it tells them.
It wouldn't be unusual for a native speaker of English to encounter language that occurs with the frequency of "hit the jackpot" a number of times per month. That's because native speakers of English tend to encounter millions of words each month. The recent Mehl paper in Science suggests that we speak on average something like 16,000 words per day. Presumably, we're doing much of that in conversation with others, often more than one person, so let's put our conversational word count at 40,000 per day spoken and heard.
Then there's TV. I don't have average numbers, but after looking at a few transcripts, it looks like 7,000 words per hour might be a reasonable estimate. According to Neilson, the average American spends 4.5 hours per day watching TV, so we can add another 30,000 words or so to our count, which now totals 70,000.
I have no data on how much people write, but I suspect it's very little. In terms of reading, I can find no adult data, but 5th-grade children read about 5,300 words per day, bringing our total daily word exposure to roughly 75,300 or 2,290,000 words per month. There are likely other sources of input that I have omitted, but this should be sufficient to make the point.
At the previously established rate of 0.18 to 2.0 occurrences paw, we could expect to see "hit the jackpot" about one to four times a month. If you're about my age, you've probably heard it about 900 times in your life. So, contrary to the above writer's conclusion, it's not at all surprising that we know it. But would you be surprised hear that my six-year-old son doesn't? (I just asked him [update: May 25, 2009. He's almost eight and he still says he doesn't know. update 2: Oct 9, 2011: 10 and still unfamiliar.]).
In contrast to native speakers, our learners don't get anything like 2.3 million words a month input. And what input they do get is degraded by the fact that they don't understand much of it. Thus, what seems very common to us, is quite rare for learners. Somehow, though, it's hard to get many language teachers to accept this. They refuse to believe that idioms are not common, but as we saw recently, anything below about 30 occurrences pmw should be considered low frequency.
There are many factors that can skew our perception of a word's commonality. Psychologists have taken this issue much more seriously than have language teachers/applied linguists and have evolved a number of measures. These include:
  • number of letters/phonemes/syllables
  • written/spoken frequency
  • range/keyness/burstiness
  • subjective familiarity rating
  • concreteness rating
  • imagability rating
  • meaningfulness
  • average age of aquisition
  • word category (noun, verb, adj, etc.)
  • affixation
  • status (colloquial/dialect/alien etc)
  • semantic grouping
It would obviously be too onerous to consider all of these constantly in our teaching, but it might not be a bad thing to know about and understand each measure.
Earlier posts this series: Idioms, Differences between the corpora, & Where's the cutoff

Tuesday, July 17, 2007

NY Times: unbalanced but honest

Cornelia Dean, writing in the New York Times today included the following statement in an article about a creationist book
"In fact, there is no credible scientific challenge to the theory of evolution as an explanation for the complexity and diversity of life on earth."

Every article, in every newspaper, that discusses creationism should include such a clear statement. Too often you get so-called balanced reporting where creationism and belief in evolution are both put forward as viable alternatives, as in this Canadian Press story in the Globe and Mail.

Monday, July 16, 2007

Idioms: Where's the cutoff?

Like most people, English teachers can find it tedious to address the same basic grammar points and vocabulary items week after week. It's repetitive and hardly sexy. What many teachers really want is to get into the subtle points of language, the nuances of sophisticated use. Teaching idioms and collocations can help us feel like we're giving our students some value for their money, something they might not be able to get in a standard dictionary. But this is a rather selfish way to go about teaching a language.

As we saw the other day, learning a language is a long slow process. The fact is that few learners move beyond the rudimentary levels. Most students arrive in my classes not knowing the most common 2,000 words of English.

Paul Nation has argued that the top 2,000 to 3,000 word families should be considered high-frequency vocabulary. For one thing, there is broad agreement from one list to the next as to what these words are. This gives us confidence that they are not merely an artifact of a particular corpus construction. This list of words will also give enough coverage that learners will be able to begin to function independently. Furthermore, the additional coverage gained by learning words beyond this level is minuscule. The next 1,000 words only increases your coverage by about 2% and the payback is smaller and smaller as you go up. That is not to say that students should stop at 2,000 words, but merely that it is at this level that they should really take over from the teachers. Finally, from a pragmatic viewpoint, 2,000 to 3,000 words is about all that one can realistically expect to deal with in a course of study, be it the six years of jr. & sr. high school language instruction that is common around the world or a one-year intensive English language program.

If you look at how common these words are, you find that the lower frequency words occur roughly 30 times per million words. So teaching anything less common (and idioms are almost always far less common) really requires some extraordinary justification.

I don't know any teacher who would address say the subjunctive before teaching the progressive aspect, yet when it comes to idioms and vocabulary, a different standard seems to be applied. Perhaps the problem is that teachers have very little sense of what 30 times per million words or 0.3 times per million words tells them. More on this to follow.

In this series: Idioms, Differences between the corpora, & Interpreting the frequencies

Saturday, July 14, 2007

Language Learner Literature Award deadline

The Extensive Reading Foundation's Language Learner Literature Awards (previously mentioned here) will be announced August 31. Readers still have a chance to submit their votes, but the deadline is July 20th.

According to this article, one library in New Zealand has had a great deal of success with displays of nominated books. Not a bad idea.

Friday, July 13, 2007

Japanese words for white

In a NewScientist interview, linguist Annie Mollard-Desfour claims that, to a Japanese person, the brightness of a colour is more important than its hue and that the Japanese language has a large number of words for white, "from the dullest to the most brilliant".

I'm afraid that despite spending ten years in Japan I had never noticed any of this, so I put it to my Japanese wife. She was as perplexed as me. We had a look in the Kenkyusha New English-Japanese Dictionary, 5th ed., and could find only a single translation for the colour white, well 2 actually: the adjective 白い and the noun 白. Oh, there were words like クリーム(cream) and compounds like 雪白 (snow white), and even metaphorical uses meaning pure, snowy, Caucasian and what have you, but only one word for the colour white.

This, of course, brings to mind the great Eskimo Vocabulary Hoax and the endless snowclones that people love to rehearse in exoticising a language or a people.

Yet, since Mollard-Desfour is a linguist and a lexicographer, I'm sure she's not simply making this stuff up. I, therefore, sent a letter to NewScientist requesting clarification. I look forward to seeing the examples or citations that my wife and I must have overlooked.

[See the follow up here]

Thursday, July 12, 2007

Idioms: differences between corpora

One of the arguments that came up yesterday was that idioms are simply not reflected in the frequency counts because the corpora don't reflect the type of language use in which these idioms would usually occur. Michael Stout suggests something similar in his comment.

It is certainly true that particular words, expressions, and forms vary quite considerably in their distribution and frequency on a scale that is often seen as ranging from spoken/informal language to written/academic language. (In fact this is probably better seen as multidimensional space, but that's another issue.) Indeed, if we look at fuck, we can see that it is strikingly common in spoken conversation, occurring 136 times pmw in the BNC vs. 1.63 times pmw in the academic subcorpus.

The following look at "hit the jackpot" should show the range in frequency of these types of idioms:
  • From yesterday, we have 0.32 occurrences per million words in the BNC. Looking at the subcorpora, we have a high of 1.88 pmw in News.
  • In the Time corpus, we have 0.8 pmw with a high of 2.0 pmw in 1940s.
  • The MICASE corpus that Michael mentioned in his comment yesterday has zero instances of "hit the jackpot" in 1,848,364 words. MICASE is unscripted 'merican speech at universities, mostly in lectures and academic discussions.
  • The Enron e-mail corpus has 18 occurrences of "hit the jackpot" in 96.3 million words (about 0.2 pmw; but a number of them are duplications)
  • The first release of the ANC is 11 million words. "Jackpot" occurs 3 times, but "hit the jackpot" does not occur.
  • The million-word Brown corpus, which is an early US written-text corpus has no instances of "jackpot".
  • The Corpus of Spoken, Professional American-English does not have any instances of "jackpot" in its sample of 42,739 words.
  • Nor does the 2-million word US talk TV corpus at Lextutor.
  • I don't have access to the The Wellington Corpus of Spoken New Zealand English, but if anyone else does, please let me know the results. I'll bet that they are all in the same range. [update: July 23| Bernadette Vine, who manages the corpus, was kind enough to do a search for me. She reports that there are no instances of 'hit the jackpot' in the one-million-word corpus.]
    [Update2: Oct 9, 2011. Some new corpora have become available, so I've added them below.]
  • In the Corpus of Current American English, we have a high of 0.61 in magazines and a low of 0.05 in academic writing.
  • And here's the frequency in the Google Books corpus throughout the 20th century, which maxes out at 0.06 pmw.

So, overall, we nothing goes above 2.0 and most are much lower. I would be very surprised, then, to find a situation in which "hit the jackpot" occurs at, say, about 30 pmw or more, at least not one that is going to be relevant to many learners of English.
Tomorrow, more about what exactly "common" would mean for a learner (hint: see the previous paragraph.)
In this series: Idioms, Where's the cutoff, & Interpreting the frequencies

Wednesday, July 11, 2007

Idioms

Language teachers are overly enamoured of idioms. My pointing out, over on the TESL-L list, that a series of "common gambling idioms" is not at all common began a raveled thread of comments questioning the value of corpus data and supporting the teaching of idioms.

Here's what I posted. The numbers are the occurrences per million words, first in the British National Corpus, and second in the Time corpus (in peak decade).
  • hit the jackpot: 0.32 (2.0 in 1940s)
  • on a roll: 0.30 (2.21 in 1990s)
  • ace in the hole: 0.04 (0.08 in 1940s)
  • Bingo!: 0.17 (0.64 in 1990s)
  • play(s/ed/ing): [somebody's] cards close to [somebody's] chest 0.07 (0.06 in 1960s)
  • wild card: 0.54 (1.38 in 1990s)
  • shoot the works: 0 (0.80 in 1930s)
  • put(s/ting) * money down: 0.05 (0.11 in 1990s)
  • beginner's luck: 0.04 (0.32 in 1960s)

To give you some context anathema, which is about the 23,800th most common word in the British National Corpus, occurs 1.42 times per million words. In other words, unless a learner of English has a huge vocabulary, there are lots and lots and lots and lots and lots of more useful things teachers can be teaching them than gambling idioms (or almost any other idiom, for that matter).

The following is a sample of the responses. Over the next few days I'll try to untangle some of them.

  • "So I think we can safely say that there are times that word frequency lists can be misleading."

  • " I have noted, at first with some dismay, the rabid attacks on any form of linguistic sophistication. Apparently, our foreign students have far better things to do than learn the subtleties of the language they are studying."

  • "Comments have been made on teaching idioms. Idioms are of utmost necessity in using and understanding English. "

  • "I have never said the word anathema, partly cos I'm not sure how to say it. But the gambling idiomatic terms turn up frequently - maybe once a month for each, in colloquial speech in NZ, so they should/could be taught."

  • "None of the correspondents have suggested that they have any difficulty recognising or understanding “hit the jackpot”, yet the low level of occurrence of the expression in corpora suggests that it should be so unfamiliar as to cause difficulty even to native speakers (I could imagine that there are many native speakers of English who would have problems with “anathema”). If the expression is so uncommon, how do we all know it?

    "One possible explanation is that the corpora that we have available seriously misrepresent the language we encounter in our daily lives."

Followups: Differences between the corpora, Where's the cutoff, & Interpreting the frequencies

Tuesday, July 10, 2007

A tale of two numbers

  • $980: funding that schools in Ontario received per ESL student in 2004/2005
  • $245: average personal spending on English-language tutors in South Korea in 2006

Details:

According to a report by the Auditor General of Ontario, schools get $225 million (all figures in Canadian dollars) in funding for ESL students. That works out to about $980 per student. In contrast, Korea with a population of about 72 million people spent $17.7 billion on English-language tutors last year. That works out to $245 for every Korean man woman and child. For tutors.

Sunday, July 08, 2007

Things to re-meme-ber me by

The Ridger FCDE at The Greenbelt has tagged English, Jack in an internet meme. Here's a recent history of this particular one:
  1. The Greenbelt
  2. Thoughts in a Haystack
  3. Evolving Thoughts
  4. On Evolution
  5. Scientia Natura: Evolution and Rationality
  6. The Flying Trilobite
  7. Pharyngula
  8. Ironicus Maximus
These are the rules:
  1. We have to post these rules before we give you the facts.

  2. Players start with eight random facts/habits about themselves.

  3. People who are tagged need to write in their own blog about their eight things and include these rules in the post.

  4. At the end of your post, you need to choose eight people to get tagged and list their names.

  5. Don't forget to leave them a comment telling them they're tagged, and to read your blog.

And, these are the facts:

  1. My earliest memory (true or created, I'm not sure) is of hanging on a fence in my Grandpa's backyard in Pipestone, Manitoba. A barb was impaled in my left palm, but the memory is simply a placid image, nothing more.

  2. When I was living in Chiang Mai, thinking a bicycle would be be a great way to get around, I bought a mountain bike. It was, indeed, a wonderful means of transportation and afforded me a good deal of exercise as well. One day, I decided to ride the bike to Wat Phrathat Doi Suthep the top of Suthep mountain. I made my preparations and set out the next morning, early before the sun was too high. I rode out past Chiang Mai University and started up the mountain. Soon, however, I had run out of water and food and the temperature was well up into the 30s. The climb was much harder than I had anticipated and I was just about done in. I hadn't passed anywhere to get food for a long distance and had no idea if I could make it back. Just then, I came around a bend in the road and over to my left saw an open-air restaurant. Salvation! I coasted down to the edge and then, dripping with sweat, red in the face, rubber legged, dressed in cycling shorts and feeling very foolish, I walked over to the buffet. I looked around but could see no wait staff. Eventually a woman came and asked me what I wanted. I pointed it out and reached for my money. "Not to pay," she said. "Is wedding."

  3. In grade 3, I won the school speech contest with a speech about goldfish. The first line was "Blub, blub, blub, swish, swish, swish. Ladies and gentlemen, my speech is on fish, goldfish that is."

  4. When I was teaching jr. high school English in Japan, I was dealing with a very rowdy bunch of first year girls. One girl, in particular, who was very weak academically was fooling around not paying attention. After using various tactics, including explicit warnings, most of the class had settled down. Then I saw the girl writing a note on some cutesy Hello Kitty stationary. I blew up. I grabbed the note from her and berated her harshly. I had no more problems for the rest of the class, but at the end, when I looked at what she had written, I saw it was not a note, but notes related to the lesson. I stopped everyone from leaving and apologised to her. Since then I've done my best never to make assumptions about my students intentions.

  5. One fall, I made about ₤20 per hour busking on the Queensway, just north of Hyde Park in London.

  6. My first bicycle was a light-blue CCM with 16-inch wheels and a movable crossbar.

  7. After racing to Ambon on the Summer Wind II out of Australia, we sailed back south towards Bali. The race had been marked by a lack wind and many of the boats had turned on their engines. After five days at sea, the captain, a 50-something Aussie who wanted to sail around the world, had promised us a leisurely trip with lots of stops. One, however, he discovered that it was hard to get steak and potatoes and that the locals didn't understand him E V E N W H E N H E S P O K E S L O W L Y A N D L O U D L Y, he rescinded that offer and headed straight for Bali. Things came to a head at Komodo where I was put off.

  8. My second toe is longer than my big toe.

Now, as for the 8 people I tag, I'm ignoring rule 5. I'll list the blogs and if these fine folks find this and respond, wonderful. And if they don't, that's just swell too.

  1. From A to Zimmer
  2. Mishka Jaeger
  3. Career Limiting Moves
  4. Stoutfellow
  5. Separated by a common language
  6. Brashaw of the future
  7. Designers who blog
  8. David Crystal's Blog

Saturday, July 07, 2007

Perception, meet reality

A recent survey in Utah seems to explain why so many Americans think that immigrants aren't trying to learn English reports The Salt Lake City Tribune.
"One of the biggest surprises from the survey, community leaders said, is the time employers think it takes to learn English. Almost half of employers said it should take six to 12 months to learn English, the survey said."

The US Department of State classifies various languages by difficulty; it all depends on your first language. But let's look at category III, which for English speakers would include Russian and Persian. According to the National Foreign Language Center, the DoS estimates that

"44 weeks of intensive language training in U.S. government language schools (five days per week, six hours per day) are required to achieve minimum working proficiency... Similar results are achieved after five years of typical college language courses, especially if students spend at least one semester abroad learning the language, in addition to their language courses in the US...

"What is 'minimal working proficiency?' Someone able to function on their own, able to talk about familiar topics and daily life."

And how many of these immigrants have the wherewithal to attend high-quality full-time English-language courses?

By the way, the survey also found that more than 80% of immigrants and refugees say they have formally tried to learn English.

Friday, July 06, 2007

Word drive on

Over at the Simple English Wiktionary, we're trying to reach the goal of having 2,000 entries by August 4th. We won't interrupt regular programming with endless boring discussions about funding issue while bugging you to pledge, but if everyone who visits here would just contribute one word, it would be very much appreciated.

Words per day

Over at Language Log, Mark Liberman has spent a good deal of time addressing the baseless claims in Louann Brizendine's book and her subsequent media appearances, in particular the idea that women speak more than men do. Today, a paper came out in Science supporting Liberman's arguments that men and women use roughly equal numbers of words. Unfortunately, Brizendine's claims are not dead; they are the undead and will continue to wander stupidly among us leaving ignorance in their wake.

The NYT has a nice graphic showing the distribution. The article it goes with isn't so hot though.


Differences between the sexes aside, the numbers are interesting. In the Mehl et al paper, women and men both spoke about 16,000 words per day. The Longman Grammar of Spoken and Written English though says that "on average speakers produce around 7,000 words per hour in the conversational texts of the LSWE Corpus, or a little under 120 words per minute. Based on this speech rate, a one-million word corpus corresponds to 140-150 hours of conversational interaction." (p. 27) Presumably the Longman folks were only considering the time when people were actively engaged in conversation. Given these numbers, we can extrapolate that your average university student spends just over 2 hours speaking per day and is silent for almost 22 hours. If you allow that most of that speaking is part of a balanced conversation, it would seem that these students spend about 4.5 hours per day involved in conversation.

In the written mode, Anderson, Wilson, & Fielding estimated that, outside of school, the median fifth-grade student reads about 600,000 words per year (about 1,650 words per day), while in school, according to Nagy & Anderson, they read about 1.3 million words per year (about 3,650 words per day) for a combined reading exposure of about 5,300 words per day.

It would be interesting to know what our total daily word count is, including everything we write, read, hear and speak. [update: July 19, 2007. I've done a quick guesstimate here which suggests an average of about 75,300 words per day (with huge individual and daily variations).]

Thursday, July 05, 2007

Super Mario Vocabulary

Kahori Sakane has an article in the Daily Yomiuri about schools giving students Nintendo DS handheld game consoles with specially designed vocabulary study software and providing them with time in class to use them. Even better, the schools seem to have done some research into the effectiveness (not well controlled, admittedly) and the students' reactions. It turns out they like it.

Actually this isn't all that surprising. We had similar success with a Mac program called Vocab (which is still free, but seems to have been forgotten by its developers) and its companion Vocab Scheduler back in the late 90s at the Tokyo high school where I taught. The benefit of this new program is that the consoles are a whole lot cheaper than computers, teachers don't need any computer savvy to run them, and they can be moved around from class to class.

Of course, you need to be studying the right words, good content (definitions, translations, example sentences, etc.) and well-established review regime, but I think this is certainly a useful intervention. Even better would be to deploy a version that works with learners' cell phones.

More about vocabulary here, here, and here.

Link courtesy of David Paul at ETJ.

Tuesday, July 03, 2007

Error collections

From time to time it is useful to have a detailed look at the kind of errors that learners of English make. The International Corpus of Learner English from Université catholique de Louvain (Belgium) has an error-tagged corpus of written text produced by learners of English. More information here. Unfortunately, it is not freely available online, but the price is modest. You can also gain access to it by sending them your students' texts (after, I assume, receiving permission from your students.)

Erors can be amusing as well as edifying, such as this one. Anders Henrikson has been collecting students' mistakes for many years. Here's a collection that has been around since at least the early 1990s. He's also got a book of them.

Finally, last week the Language Loggers posted a link to a video of Taylor Mali changing all the usual typos into speakos.

Right to bargain

English-language teachers in Ontario Colleges (and elsewhere) are disproportionately non-full time workers. In our EAP department, for instance, fewer than 2/3 of the classes are taught by full-time faculty.

In Ontario, most part-time and contract college faculty have been denied the right to join a union. These workers are supported by the CAAT division of OPSEU and have set up their own non-bargaining group, OPSECAAT, but have never had the right to bargain collectively. It looks as though change may be coming.

Back in June, the Supreme Court of Canada ruled that,

"The right to bargain collectively with an employer enhances the human dignity, liberty and autonomy of workers by giving them the opportunity to influence the establishment of workplace rules and thereby gain some control over a major aspect of their lives, namely their work."

However, the Ontario government will not reconvene until after the October election, so there will be no legislative changes until then. As far as I can tell, none of the parties, not even the labour-friendly NDP, has made any mention of this ruling in their platform. Finally, the Supremes have given governments a year to address the issue, so we shouldn't expect anything to change soon.

Friday, June 29, 2007

Simon Li stresses me out

Today marks one week since Simon Li's last day as Friday host of CBC's "The Current". You can read and listen to parts of the show on the CBC website (e.g., here). Since his first appearance on June 1st, I've been asking myself what I think of him and the answer keeps coming back that he bugs me.

As the blurb on June 1 says, "while he's new to the CBC, he's not new to the radio-waves. Simon Li is the long-time host of 'Power Politics' on Toronto First Radio, a Chinese-language current affairs program." It does appear that Li was host back in February as well. Li is an MA student in History at Queen's University.

It seems that Li is originally from Hong Kong, which leads to what bugs me: his accent.

I find this hard to accept and that's why it's taken so long for this post to see the light of day, but it seems I can't escape it. I'm not sure why this should be. I listen to learners of English from all over the world every day and don't have any problems with it. I hear many experts interviewed on the radio and TV who have much worse English than Li and yet I focus on what they're saying, not their accents. I can listen to announcers with Australian accents, South African accents, Jamaican accents, Cockney accents and French-Canadian accents and none of them trouble me at all. So why do Li's broadcasts cause me such discomfort?

I think there are a few factors:
  • He seems to be working very hard to be crisp in his pronunciation, yet he swallows a lot of consonant clusters and finals that are clearly pronounced in standard English. (e.g., McJob comes out as /mdʒɒ/).
  • He makes me work just to understand him. There are times when I have to think to recover words that I've missed.
  • His voice quality doesn't do it for me aesthetically. He's got one of these androgynous voices.

What really kept needling me though was his habit of stressing the final noun where we would normally stress the modifier. For example, we would usually say soccer ball, but Li would call it a soccer ball. Here are some examples from his broadcast:

  • Li was host of the Friday edition, but he kept saying the Friday edition
  • the Ipperwash inquiry
  • our Toronto studio
  • specific findings
  • Dudley George
  • Fighting words

This seems to jive with general findings that it tends to be the suprasegmental aspects of accent that bother people more than individual vowels and consonants.

Monday, June 25, 2007

Grammar for New South Wales teachers

It appears that to get a teachers license in New South Wales, you will need to study grammar. Do you suppose that Rodney Huddleston, being right nearby in Queensland, might be able to influence WHAT they study about grammar? One can hope...

Thursday, June 21, 2007

On lexicons, gaps, Zipf and exponents

While I won't spend more time dwelling on the numerical problems that I brought up in my previous comment on Brown's lexicon article in the Globe, I would like to note an interesting feature which does deal with numbers.

Part I:
"If Little Princess is an average child, she'll know 6,000 root-word meanings by the end of Grade 2. That's okay, but nothing special: At that point, the top 25 per cent of children already know twice as many words as the lowest 25 per cent, and the gap grows exponentially."

I don't think that Brown means exponentially in the literal sense. Such a literal meaning would be described by the formula:

G_y = G_b^y

where Gy is the gap in year y, y is the number of years that has elapsed, and Gb is the gap in the base year. This kind of exponential growth would soon have adults knowing hundreds of thousands of words, far beyond what any reasonable estimate claims.

Part II:

"But by (graduation) the foundation of her so-called mind has hardened. Limited by early lexical laxity, the average North American adult knows only 30,000 to 60,000 words, out of a potential "working vocabulary" of 700,000. If only Little Princess had learned more words earlier! If only you were a better parent!"

Brown is right about something: People learn fewer and fewer words as they grow older. But how much has it to do with calcification of the synapses and how much with the frequency of words? This brings us to our second equation of the day.

Empirical research has found that in English, the frequencies of the approximately 1000 most-frequently-used words are approximately proportional to 1/ns where s seems to be just slightly more than 1.

This second equation describes a distribution that is knows as Zipf's Law after George Zipf (not to be confused with the ill-fated George Zipp). In other words, the second most common word is about 1/2 as frequent as the first (which happens to be the), the third most common is only about 1/3 as frequent as the, and the fourth most common is only about 1/5 as frequent.

Now this isn't exponential change either, but it sure means that the frequency of words drops off pretty darn fast. Consider then a child's chances of meeting and learning their 3000th word compared to an adult's chance of meeting and learning their 60,000th word. At less than 1/60,000 the frequency of the, you could easily be well into middle age before you even encounter the word for the first time. If you're lucky enough to encounter it again before you retire, you're unlikely to recall that first meeting when you do. In other words, it's mainly the inherent distribution of words that makes it so hard to learn new words as you get older.

Wednesday, June 20, 2007

Mark Davies & the Time corpus

Mark Davies has a new free online interface. The following blurb is from the site:
"This website allows you to quickly and easily search more than 100 million words of text of American English from 1923 to the present, as found in TIME magazine. You can see how words and phrases have increased and decreased in usage and see how words have changed meaning over time."
The interface will be familiar to anyone who has used his excellent VIEW interface to the BNC BNC corpus interface (now available here).

Monday, June 18, 2007

Losing our lexicons?

Ian Brown, writing in Saturday's Globe & Mail, takes a hell-in-a-handbasket tour of society's love-hate relationship with vocabulary. (As far as I can tell, the Globe conspires to make it impossible for me to link directly to an article without making it appear as though you have to pay for it. So I'm hoping that by linking through Google I can to bypass this barrier.)

I'll start by getting my quibbles with the first line out of the way. "The last days of long words! The sunset of syntactical surplusage!" I can understand the lure of easy alliteration, but in a feature about vocabulary, the choice of syntactical here is sad. Semantic would have been a smidgen closer, and there's still that juicy /s/ at the outset. The loss of lexical lavishness would have been just right.

Yes, well, on to the article proper: In one corner, quoth Brown, we have the logophiles like Conrad Black: those poor, misunderstood folk who simply love words and can't understand what anyone could have against dropping "tricoteuses (knitters of yarn, used to describe reporters and gossips, augmented by the adjective "braying"), planturous (fleshy), poltroon (a coward, a.k.a. former Quebec premier Robert Bourassa), spavined (lame), dubiety (doubt: Mr. Black rarely uses a simple word where a splashy lemma will do), gasconading (blustering) and velleities (distant hopes)" into everyday chit chat.

In the other corner (because setting up false dichotomies makes for a juicier read) we have
"the linguists, who have the upper hand at the moment, (and) are very much of a type. They tend to be acolytes of American scholar William Labov, who developed the concept of code-switching. Standard vocabulary doesn't need to be taught, the Labovites claim, because there's no such thing as a standard vocabulary... Teaching a standard vocabulary today isn't just ineffective: According to the linguists, it's undemocratic and limiting."

Some of the most militant linguists are Canadian. Clive Beck, a professor of education at the Ontario Institute for Studies in Education, relishes the collapse of the standard Western vocabulary. "I think it's partly a democratization, of getting teachers to have a closer relationship with their students, and being able to talk on the same level. I love correctness in speech and in writing. But I think to some extent I have to go with the change."

Here I fight my innate desire to smother hyperbole with sarcasm and simply note that the phrase, "I think to some extent I have to" doesn't exactly reek of militancy or relish. Indeed, there is little truth to be found in this quoted section (though the lack doesn't stop there).

  • I have yet to meet a linguist, in any sense of the word, who is not also a logophile.
  • To be effective at code switching, you need a diverse vocabulary, not a restricted one.
  • Of course there is a standard English vocabulary, but, by definition, it doesn't include words that most people don't know.
  • As to whether you should teach vocabulary or not, that all depends on what the learners already know and what their purpose is, (more here).

All in all, Brown misses the point that it's simply impractical to teach someone enough words to have a large vocabulary (problems with counting words aside, his numbers are wonky; high school graduates know 6,000 + 35,000 = 41,000 words, but the average adult knows only 30,000?) A large vocabulary is merely a symptom of an educated person. Yes, some direct vocabulary teaching is probably useful with young children and ESL students, and having a good basic vocabulary will bootstrap other learning, but simply stuffing a bunch of words into your head is a trivial pursuit. The benefits that come with a broad vocabulary are those that stem from having a variety of interest, reading widely, and discussing concepts in some depth with other interested parties.

Which leads us to one more of Brown's forced choices: deploy grandiloquent vocabulary indiscriminately or use only short simple words. This is like saying that your wardrobe should be limited to white t-shirts & jeans or formal-wear. It should be obvious that you match your vocabulary to the topic and audience at hand. Occasionally using a rare but fitting word in a context that will make it clear is fine. Weighing down your speech with word after word that your audience is unlikely to be familiar with is like wearing a tux to your child's soccer practice. What purpose could the speaker have other than obfuscation or pretentiousness? I suppose there is one other likely explanation: an abiding lack of empathy. Either way, you're going to arouse suspicion. And that's why Black's lawyer kept him from testifying.

Sunday, June 17, 2007

The perils of grammatical celebrity

For better or for words, Grammar Girl is one of the most popular educational podcasts, and Technorati ranks it as about the 4000th most popular blog (roughly 2 orders of magnitude ahead of English Jack). Mignon Fogarty has done a great job publicising her site, but with publicity, you end up getting misrepresented in the media. Now I'm not really sure what all went on with this article, but I would bet that Fogarty's points are rather decontextualised. The premise involves comparing a grammarian's responses to various song lyrics with a DJ's responses. Grammar Girl comes out looking ridiculous. Here's an example:

Song: “If You Love Somebody Set Them Free,” by Sting.

Offending line: “If you love somebody, if you love somebody, if you love someone, set them free.”

Grammar Girl says: “Someone” is singular, so technically he should have sung, “If you love someone, set her free,” or “set him free,” or “set him or her free.”

DJ Steve says: That’s the way we talk as normal human beings. Imagine the alternative. If he were to say “his,” then everyone would be sitting around talking about it and we’d all be wasting our time talking about this otherwise likable but forgettable Sting song.

To make it worse, elsewhere they quote Fogarty referring to the "subjunctive case". Now GG doesn't strike me as being a deep grammatical thinker. Her whole gig involves rehashing the same old issues covered in any style guide. If there's creativity, its entirely in the presentation, not in the analysis. Still, GG does know the difference between a mood and a case, at least terminologically speaking. She has at least 3 posts referring to the subjunctive mood and none referring to "subjunctive case".

Saturday, June 16, 2007

More ESL funding frustration

Two article from separate areas, one from here in Toronto and one from the U.S., document the ongoing funding problems for ESL students. The first reports that only half the ESL funds in the Toronto District School Board are used for ESL. The second catalogues the consequences of underfunding.
"The results of national testing conducted in 2005 shows that nearly half (46%) of 4th grade students in the English language learner (ELL) category scored "below basic" in mathematics in 2005--the lowest level possible. Nearly three quarters (73%) scored below basic in reading. In middle school achievement in mathematics was lower still, with more than two-thirds (71%) of 8th grade ELL students scoring below basic. Meanwhile, the same share (71%) of 8th grade ELL students scored below basic in reading."

Thursday, June 14, 2007

Over delaytion

Nothing for weeks and then twice in the same day. Anyhow, driving in I was listening to "The age of persuasion" on CBC. I missed the start, so I'm not sure what the theme was, but then again, I still wasn't really sure at the end. It was something of a grab-bag of anecdotes and observations related to advertising and just like ads, seemed to jump from topic to topic rather promiscuously, although with some hint of an underlying cohesive idea. Anyhow, it's a scripted/edited show, so I was a bit surprised to hear host Bill O'Rielly's explanation for why a new technology is ready but not yet available: "So what's holding up the delay?" he asks rhetorically. "The human factor."

Holding up the delay? This seems to be related to overnegation, but neither hold up nor delay are syntactically negative.

Google only knows of two other instances of "holding up the delay": Denise Quaid saying, "You're always waiting around for something and then when you get there, you find out that what's holding up the delay is that something doesn't work," and this example, in which the speaker self-corrects. "You wanted Mr Burkett to tell Vince Alessandrino that it was the cock-up of the Department of Family and Children Services that was holding up the delay - - that was delaying the settlement?---No, I don't think that's correct. "

All this talk of negation reminds me of the study site for a chapter we've been looking at in my level 6 reading class. It includes review questions, a sadistic number of which are negative. Take the following example:
According to the United Nation's Human Development Index (HDI), China has a better score than all EXCEPT which of the following countries?

How about this instead: "Which of the following countries has a higher HDI score than China?" Can't you avoid writing questions that aren't less difficult than this?

Accent quiz

Sorry I've been rather derelict. Things wrap up at school in a few more weeks.

In the mean time, here's an interesting quiz to mollify you. It pegged me as Canadian, though the difference between Canadian (as if there were only one Canadian accent!) and midland seems slight (perhaps just the /a/ in pasta.)

Personality Test Results

Saturday, June 02, 2007

Grammatically speaking

Richard Firsten has published his June "Grammatically Speaking" column. I've commented on these before here, here, here, here, and here.

This time there's a mixture of good and bad.

In the first question about the difference between I wish they didn’t do that and I wish they wouldn’t do that, I take issue with the idea of future tense and subjunctive, but it's mainly a label issue. I generally agree with his analysis.

I quite like his discussion about the difference between participles and gerunds. Again, I would have used different labels, but the reasoning is good, and that's what matters. As Firsten explains, "if you can place an article and/or an adjective before it, it is a gerund". More grammar books should keep this in mind.

I also thought the 'much' question was handled well. Where I completely disagree is with his rejection of "graduated college". This answer is everything that is wrong with prescriptivism. I've actually addressed the issue before when Russell Smith complained about it.

Merriam Webster's has an entry for it: (transitive) "1b. to be graduated from" and says that though it is often condemned, it is "standard".
The American Heritage Dictionary: (transitive) "1b. Usage Problem To receive an academic degree from." (25% of their usage panel accept this)
Random House Unabridged: "Informal. to receive a degree or diploma from: She graduated college in 1950."

Last time, I said that it wasn't common in edited prose, but I think I may have been wrong, at least as far as the US and Canada is concerned:
  • "The average person that has graduated college changes careers seven to eight times," said Christopher Hopey, vice president and dean of the School of..." (Boston Herald)
  • "'Hill' will pick up next season four years into the future, after the characters have graduated college." (Washington Post)
  • "For students who have just graduated university, scoring a job interview after graduating is an important step on the way to getting their first full-time..." (CTV.ca)
  • "The killing of the 22-year-old Kentucky native, who recently graduated university with honors, in a tough neighborhood in Boston's Dorchester district..." (Reuters)
  • "Nisbet was a star student when he graduated high school and was picked as his school's athlete of the year" (Edmonton Journal)
  • "Before the Iraq war began, the percentage of Army recruits who graduated high school surpassed 90 percent." (Boston Globe)
  • "Elizabeth ''Betty'' Reily went through life proud that she graduated high school with A's and B's." (Miami Herald)
  • “Today three quarters of boys and half of girls have had sex by the time they graduate high school” (Newsweek).
  • “‘I have a reading disorder,’ Leschuk says, yet he struggles to think of any friends who graduated college who are doing as well” (USA Today).
This transitive sense is no less clear than the intransitive form, and it seems arbitrary and capricious to simply say it's wrong, not to mention that such treatment fails to promote any kind of understanding of the matter.