Re: [xliff-inline] Type of codes

From
Dr. David Filip <>
Date
2011-12-07T19:40:48+00:00
ID
CANw5LKkv=
Thread
Re: [xliff-inline] Type of codes
Hi all, I would go for A, (and C as a second option).

The simple reason is that both A and C contain pc that is XML friendly. X(ML)LIFF should not be XML hostile IMHO.

Primarily, I am for expressivity. As for simplicity, even A is a big simplification (in good sense) compared to TMX 1.4b and XLIFF 1.2.

I agree with Yves that less inline elements does not necessarily mean simplicity, as then the semantics must be carried by attributes anyway..

I consider A the minimum reasonably expressive set. C would be a compromise, largely harmless, still unnecessary IMHO.

Rgds

dF
Dr. David Filip
=======================

LRC | CNGL | LT-Web | CSIS

University of Limerick, Ireland

telephone: +353-6120-2781

mobile: +353-86-049-34-68

facsimile: +353-6120-2734

mailto: 

On Mon, Nov 14, 2011 at 16:31, Yves Savourel <> wrote:

Hi everyone,

One last summary on the type of codes discussion.

I've listed below the four options we have:

A) ph, pc, sc/ec

This is the current draft. We allow both pc and sc/ec, the second being used when the first cannot.

The main advantage is that it offers the pc cleanness to users who need it.

It has the drawback of making thing a bit more complex for the tools.

B) ph, sc/ec

This would remove pc, marking all span-like original codes with sc/ec.

It simplifies somehow the writing of the documents.

Its main drawback is that not having a pc notation is not very friendly for XML-like original formats.

C) ph, pc

This would change ph to handle also the sc/ec cases.

This would add several attributes to ph. Essentially it would transfer the syntax of two element names to new attributes in ph.

The pc element would stay to provide a clean markup for the well-formed span-like codes.

D) ph

This would code all cases with ph: One code to rules them all.

The ph element would have many attributes to handle all the cases.

This option is especially attractive for tool extracting only placeholder codes.

Note:

- Regardless of the option, we would still have a distinction between placeholder and span-like codes. In other words, even with option D (<ph> only) a filter should be able to extract an HTML <BR/> and a <B>...</B> making a distinction, for example: <ph id='1'>&lt;BR/></ph> and <ph id='2' kind='start'>&lt;B></ph>...<ph rid='2' kind='end'>&lt;/B></ph>. At some point we'll have to have a discussion about whether extraction tools MUST or only SHOULD do the distinction, but that's independent of the representation.

- Keep also <mrk> in mind. The presence of <pc> would make the use of <mrk> a bit more complex: (e.g. <mrk> overlapping <pc>). But not that much because <mrk> can overlap <mrk> too so the use case exist even if <pc> is not there.

My personal opinion:

I think have distinct elements for original codes that are placeholder vs span-like is useful, even important. So I would not put everything into ph.

I would tend to think having the span-like codes handled only with sc/ec is fine. But looking at how v1.2 is used, the requirements we have for 2.0 and the feedback of the past months, I can see there are arguments for also having pc, and since the drawbacks of keeping pc are not that big I would. We just need to make sure it can be mapped transparently to sc/ec and conversely.

So I would pick option A.

What would you pick and, as importantly, why? (so all can understand the rationale of the choice)

Cheers,

-yves

---------------------------------------------------------------------

To unsubscribe, e-mail: 

For additional commands, e-mail: