Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Just curious, how accurate was the rest of the translation?


Well, for a translation about finance it is (very!) surprising that it seems to have just transformed "kronor" (="crowns", the everyday name for the SEK, the Swedish currency) into "dollars". It's not as they two currencies are 1:1, or even close to it. Very strange.

Also there was a Swedish word that was just copied ("anullerades", meaning more or less "cancelled", "voided" or something along those lines) into the English text.

But, overall, it's still rather impressive.


I can confirm the translating-names-of-currencies issue, but something even worse is that the names of languages can also be similarly converted.

For example, the word "Eesti" will often get Google-Translated to "English" rather than "Estonian".

This means that a film at my local cinema that my web browser assures me is in "English" will in fact be in Estonian. And an interview with a Russian saying that he doesn't speak Estonian gets translated so that he appears to say that he doesn't speak English.

The product designers special-cased language names, doing extra work to produce what will almost always be the wrong result.

(And what they can do to place names is often patently ridiculous. For example, "Peterburi tee" should either be left alone or maybe translated to "St Petersburg Road" but actually somehow becomes "Hertford Road". And the ZIP + City name "13415 Tallinn" becomes "thirteen thousand four hundred and fifteen Tallinn".)


> The product designers special-cased language names, doing extra work to produce what will almost always be the wrong result.

They absolutely did not do this. It's an artefact of statistical translation. In the corpus there are a lot of English documents saying "This document is in English", whose translated versions in Afrikaans (because I know Afrikaans) say "Hierdie dokument is in Afrikaans". Thus the translator learns the "hierdie" is Afrikaans for "this", "dokument" is "document, ..., and "English" is "Afrikaans".

The street name issue probably comes from an organisation whose Estonian office is in Peterburi tee and whose English office is in Hertford Road.


> The product designers special-cased language names, doing extra work to produce what will almost always be the wrong result.

Are you sure about that? It would sound quite likely to me that simply, the word for "English" tends to appear in the same context (N-gram etc.) as the word for "Estonian". For example the sentence "I speak English" would be common in English, while "I speak Estonian" would be common in Estonian, so it might associate the words together.


The machine learning behind Google's translations uses a corpus that consists of texts that have already been translated into many different languages. But if some of those documents start off with, say, "this is the <language> version of this press release" that's exactly the sort of thing that would confuse an algorithm that can't distinguish between translation and localization. Pure conjecture of course. I'm just not sure if anything is being special-cased.


One of the things Google Translator is horrible about is translating stuff like kroner into dollars, SEK into USD, etc. Granted, there are cases where translating kroner into dollars is ok, but I'm guessing this is the wrong translation most of the time.


Since it detects the value being qualified by the currency and localizes the unit in front of the value, I would understand that behavior if it would also do some cash value translation at the current rate (which would arguably be somehow less wrong), but it does not [0]. Also, weirdness comes when it changes it to DKK for 100 [1], and $ for 1000 [0]. It just makes no sense.

[0]: http://translate.google.com/#auto/en/1000%20kroners

[1]: http://translate.google.com/#auto/en/100%20kroners


It makes no sense to the casual user, but you can understand why it does it when you learn that Google Translate uses statistical methods to learn. Google feed Translate with articles and pages that have been already been translated by a human, and the program learns the translations of words and sentences from that. The problem arises when the two documents it's taught with differer slightly. With translations, this often happens with currencies, country names (lots of examples of translate screwing those up too) and numbers.


Are there common expressions that are equivalent in english and swedish with the respective currencies? E.g. stuff like "dollar to dollar", "bet someone dollars to doughnuts" ?

Anyway, plenty of text would probably match word-wise ("X bought Y for $AMOUNT $UNIT") in financial news, so the mapping of "dollar" over "kronor" seems a reasonable error.

For probably the same reason, google also translates hungarian "1000 forint" to english "1000 HUF" going from the full word to a quasi-acronym for "HUngarian Forint".


I would guess "annulled" to be the most direct translation of "anullerades" ("annulled" carries the same meaning as "voided" or "invalidated" but tends to be used more in legal contexts).


...making "anulled" a relatively rare word in the English training corpus.


> it seems to have just transformed "kronor" (...) into "dollars".

I saw some article once, where the names of a prime minister or some such was "translated" to the name of the US president. Weird.


Just an FYI, there is no way to tell how much of that article is machine translated since Google Translate allows anyone with a Google account to "Contribute to the translation".

Though I imagine it is mostly machine translated and it requires several agreeing "contributions" before it accepts them as accurate.


Does the contribution method actually change the document being edited or just add the parallel texts to the training data?


Alright, but it has its faults - it missed the word "error": "Instead, it is about a parsing [error] incurred in exchange system due to a technical error"


I was profoundly impressed. At first, I didn't realise it was a translation. Apart from a few odd passages, the text flows really well.


Well, that's obviously because English is a Scandinavian language... http://www.apollon.uio.no/english/articles/2012/4-english-sc...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: