Re: Unicode support
От
Greg Stark
Тема
Re: Unicode support
Дата
Msg-id
4136ffa0904140849h36bdb5adl8b4e765b1906c4ed@mail.gmail.com
Ответ на
Re: Unicode support (Peter Eisentraut)
Список
Дерево обсуждения
Unicode support - - <crossroads0000@googlemail.com>
Re: Unicode support Alvaro Herrera <alvherre@commandprompt.com>
Re: Unicode support "Kevin Grittner" <Kevin.Grittner@wicourts.gov>
Re: Unicode support Andrew Dunstan <andrew@dunslane.net>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support Greg Stark <stark@enterprisedb.com>
Re: Unicode support Tom Lane <tgl@sss.pgh.pa.us>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support "David E. Wheeler" <david@kineticode.com>
Re: Unicode support Andrew Dunstan <andrew@dunslane.net>
Re: Unicode support Tom Lane <tgl@sss.pgh.pa.us>
Re: Unicode support "David E. Wheeler" <david@kineticode.com>
Re: Unicode support Martijn van Oosterhout <kleptog@svana.org>
Re: Unicode support - - <crossroads0000@googlemail.com>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support "Kevin Grittner" <Kevin.Grittner@wicourts.gov>
Re: Unicode support Andrew Dunstan <andrew@dunslane.net>
Re: Unicode support Tom Lane <tgl@sss.pgh.pa.us>
Re: Unicode support Greg Stark <stark@enterprisedb.com>
Re: Unicode support Tom Lane <tgl@sss.pgh.pa.us>
Re: Unicode support - - <crossroads0000@googlemail.com>
Re: Unicode support Gregory Stark <stark@enterprisedb.com>
Re: Unicode support Andrew Gierth <andrew@tao11.riddles.org.uk>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support Andrew Gierth <andrew@tao11.riddles.org.uk>
Re: Unicode support Andrew Dunstan <andrew@dunslane.net>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
Re: Unicode support Tom Lane <tgl@sss.pgh.pa.us>
Re: Unicode support Peter Eisentraut <peter_e@gmx.net>
On Tue, Apr 14, 2009 at 1:32 PM, Peter Eisentraut wrote: > On Monday 13 April 2009 22:39:58 Andrew Dunstan wrote: >> Umm, but isn't that because your encoding is using one code point? >> >> See the OP's explanation w.r.t. canonical equivalence. >> >> This isn't about the number of bytes, but about whether or not we should >> count characters encoded as two or more combined code points as a single >> char or not. > > Here is a test case that shows the problem (if your terminal can display > combining characters (xterm appears to work)): > > SELECT U&'\00E9', char_length(U&'\00E9'); > ?column? | char_length > ----------+------------- > é | 1 > (1 row) > > SELECT U&'\0065\0301', char_length(U&'\0065\0301'); > ?column? | char_length > ----------+------------- > é | 2 > (1 row) What's really at issue is "what is a string?". That is, it a sequence of characters or a sequence of code points. If it's the former then we would also have to prohibit certain strings such as U&'\0301' entirely. And we have to make substr() pick out the right number of code points, etc. -- greg
В списке pgsql-hackers по дате отправления