Theory
Read a stranger's XML
The fest's sound vendor sends their equipment list as XML. Above the actual data sit 3 mysterious lines: an <?xml ...?>, a comment naming their software, and an odd <?xml-stylesheet ...?> you have never met.
Where does the preamble end and the data begin? Which lines could you delete?
XML's anatomy has an official answer. Every document, theirs and yours, splits into exactly 2 sections, and knowing the boundary is a standard 5-mark question.
Theory
Cover page and thesis
A project report has a cover page: title, author, formatting notes for the binder. Useful, sometimes skippable, but never the content: no examiner grades the cover.
Then the thesis itself: every chapter, every mark-earning word, inside one binding.
An XML document is bound the same way: a prolog (the cover: declarations and notes) and the document element section (the thesis: the root and ALL the data).
Theory
Section 1: the prolog
The prolog is everything before the root element. It may contain:
- the XML declaration:
<?xml version="1.0" encoding="UTF-8"?>, first if present - comments:
<!-- maintained by the fest committee --> - processing instructions (PIs):
<?xml-stylesheet type="text/css" href="fest.css"?>, messages to specific software - optionally a DOCTYPE line referencing a DTD (a validation grammar, beyond this syllabus)
The whole prolog is optional, and it carries no data: delete it and the information survives.
Theory
Section 2: the document element section
The document element section is the root element and everything inside it: every element, attribute and text node. All the data lives here, which is why the root is called the document element: it IS the document, structurally.
Comments may appear inside too (annotating a tricky entry), with 2 rules everywhere they go: no -- inside a comment's text, and comments never nest.
Minimal legal document? Just <events></events>: no prolog at all, one empty root. Well-formed.
Practical
events.xml, dissected
<!-- ============ PROLOG SECTION ============ -->
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/css" href="fest.css"?>
<!-- events.xml: maintained by the fest committee -->
<!-- ====== DOCUMENT ELEMENT SECTION ====== -->
<events>
<!-- comments may sit inside the data too -->
<event id="1">
<name>Garba Night</name>
<venue>Main Ground</venue>
</event>
</events>
Quiz
Which of these belongs to the PROLOG section of an XML document?
- The root element and its attributes
- The XML declaration and any comments written before the root
- The text content of child elements
- Every comment in the file, wherever it appears
Show the answer
The XML declaration and any comments written before the root
The prolog is defined by position: whatever legally stands BEFORE the root: the declaration, pre-root comments, processing instructions, an optional DOCTYPE. The root and its attributes (option A) ARE the document element section's opening, and child text (option C) is its cargo. Option D is the subtle trap: comments are allowed in both sections, so a comment inside the root belongs to the document element section: the boundary is the root's start tag, not the kind of line.
Think first
Delete the prolog: what breaks?
Take the dissected listing and delete its entire prolog: declaration, PI, comment. Before tapping: is the file still well-formed, and what, concretely, is lost?
Show the answer
Still well-formed: the prolog is optional, and a bare root section is a complete document. What is lost is advice, not data: no declared encoding (parsers assume UTF-8, risky the day a Gujarati event name arrives), no stylesheet hint for browsers, no maintainer note for the next student. So the exam nuance: the prolog is dispensable to the PARSER but valuable to PEOPLE and tools. Data never lives there, which is exactly why deleting it cannot break the content.
Watch out
Comment law and a PI clarification
Comments: <!-- like this -->, never containing a bare -- inside, never nested. One comment swallowing another is a parse error, not a style issue.
PIs are not declarations: both wear <? ?>, but the XML declaration is a fixed first-line announcement while a PI like xml-stylesheet is an instruction addressed to particular software, placeable anywhere in the prolog. Calling every <? ?> line "the declaration" costs the distinction mark.
Theory
Anatomy makes error messages readable
Parser errors quote positions like "line 2, before root element": you now know that means PROLOG territory, so suspect the declaration or a stray character, not your data. Anatomy turns rejection slips into directions. One prolog citizen deserves its own lesson: the declaration and its picky placement rules close this XML unit next.
Summary
Key takeaways
- An XML document = prolog section + document element section, split at the root's start tag.
- Prolog (all optional, no data): XML declaration first, then comments, processing instructions, optional DOCTYPE.
- Document element section: the root and everything inside: all elements, attributes, text.
- Comments <!-- --> live in either section; no -- inside, no nesting.
- A bare root element with no prolog is still well-formed.
- PIs (like xml-stylesheet) are software-directed instructions, not the declaration.
- Memory hook: cover page optional, thesis compulsory.