Theory
एक अजनबी का XML पढ़िए
Fest का sound vendor अपनी equipment list XML की तरह भेजता है। असली data के ऊपर 3 mysterious lines बैठी हैं: एक <?xml ...?>, उनके software का नाम लेता एक comment, और एक अजीब <?xml-stylesheet ...?> जिससे आप कभी नहीं मिले।
Preamble कहाँ खत्म होता है और data कहाँ शुरू होता है? आप कौन सी lines delete कर सकते थे?
XML की anatomy का एक official answer है। हर document, उनका और आपका, exactly 2 sections में split होता है, और boundary जानना एक standard 5-mark question है।
Theory
Cover Page और Thesis
एक project report की एक cover page होती है: title, author, binder के लिए formatting notes। Useful, कभी-कभी skippable, पर कभी content नहीं: कोई examiner cover grade नहीं करता।
फिर thesis खुद: हर chapter, हर mark-earning word, एक binding के अंदर।
एक XML document उसी तरह bound है: एक prolog (cover: declarations और notes) और document element section (thesis: root और सारा data)।
Theory
Section 1: Prolog
Prolog root element से पहले सब कुछ है। इसमें हो सकते हैं:
- XML declaration:
<?xml version="1.0" encoding="UTF-8"?>, present होने पर पहला - Comments:
<!-- maintained by the fest committee --> - Processing instructions (PIs):
<?xml-stylesheet type="text/css" href="fest.css"?>, specific software के लिए messages - optionally एक DOCTYPE line जो एक DTD reference करती है (एक validation grammar, इस syllabus से आगे)
पूरा prolog optional है, और यह कोई data carry नहीं करता: इसे delete कीजिए और information survive करती है।
Theory
Section 2: Document Element Section
Document element section root element और इसके अंदर सब कुछ है: हर element, attribute और text node। सारा data यहीं रहता है, यही वजह है root को document element कहते हैं: यह structurally document IS है।
Comments अंदर भी appear हो सकते हैं (एक tricky entry annotate करते हुए), 2 rules हर जगह के साथ जहाँ ये जाते हैं: comment के text के अंदर कोई -- नहीं, और comments कभी nest नहीं होते।
Minimal legal document? सिर्फ़ <events></events>: कोई prolog नहीं, एक empty root। Well-formed।
Practical
events.xml, Dissected
<!-- ============ PROLOG SECTION ============ -->
<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/css" href="fest.css"?>
<!-- events.xml: maintained by the fest committee -->
<!-- ====== DOCUMENT ELEMENT SECTION ====== -->
<events>
<!-- comments may sit inside the data too -->
<event id="1">
<name>Garba Night</name>
<venue>Main Ground</venue>
</event>
</events>
Quiz
इनमें से कौन एक XML document के PROLOG section में belong करता है?
- Root element और इसके attributes
- XML declaration और root से पहले लिखे कोई भी comments
- Child elements का text content
- File में हर comment, चाहे यह कहीं भी appear हो
Show the answer
XML declaration और root से पहले लिखे कोई भी comments
Prolog position से define होता है: जो कुछ भी legally root से PEHLE खड़ा होता है: declaration, pre-root comments, processing instructions, एक optional DOCTYPE। Root और इसके attributes (option A) document element section की opening HAIN, और child text (option C) इसका cargo है। Option D subtle trap है: comments दोनों sections में allowed हैं, तो root के अंदर एक comment document element section का है: boundary root का start tag है, line की kind नहीं।
Think first
Prolog Delete कीजिए: क्या टूटता है?
Dissected listing लीजिए और इसका पूरा prolog delete कीजिए: declaration, PI, comment। Tap करने से पहले: क्या file अभी भी well-formed है, और concretely क्या lost होता है?
Show the answer
अभी भी well-formed: prolog optional है, और एक bare root section एक complete document है। जो lost होती है वह advice है, data नहीं: कोई declared encoding नहीं (parsers UTF-8 assume करते हैं, risky उस दिन जब एक Gujarati event name आए), browsers के लिए कोई stylesheet hint नहीं, अगले student के लिए कोई maintainer note नहीं। तो exam nuance: prolog PARSER के लिए dispensable है पर PEOPLE और tools के लिए valuable है। Data कभी वहाँ नहीं रहता, यही exactly वजह है इसे delete करना content नहीं तोड़ सकता।
Watch out
Comment Law और एक PI Clarification
Comments: <!-- like this -->, अंदर कभी bare -- नहीं, कभी nested नहीं। एक comment दूसरे को swallow करना style issue नहीं, parse error है।
PIs Declarations नहीं हैं: दोनों <? ?> पहनते हैं, पर XML declaration एक fixed first-line announcement है जबकि xml-stylesheet जैसा एक PI particular software को addressed एक instruction है, prolog में कहीं भी placeable। हर <? ?> line को "declaration" कहना distinction mark cost करता है।
Theory
Anatomy Error Messages को Readable बनाती है
Parser errors positions को "line 2, before root element" जैसा quote करते हैं: अब आप जानते हैं इसका मतलब है PROLOG territory, तो declaration या एक stray character पर shak कीजिए, अपने data पर नहीं। Anatomy rejection slips को directions में बदल देती है। एक prolog citizen अपना खुद का lesson deserve करता है: declaration और इसके picky placement rules इस XML unit को अगली बार बंद करते हैं।
Summary
Key takeaways
- एक XML document = prolog section + document element section, root के start tag पर split।
- Prolog (सब optional, कोई data नहीं): पहले XML declaration, फिर comments, processing instructions, optional DOCTYPE।
- Document element section: root और इसके अंदर सब कुछ: सारे elements, attributes, text।
- Comments <!-- --> दोनों sections में रहते हैं; अंदर कोई -- नहीं, कोई nesting नहीं।
- बिना prolog के एक bare root element अभी भी well-formed है।
- PIs (xml-stylesheet जैसे) software-directed instructions हैं, declaration नहीं।
- Memory hook: cover page optional, thesis compulsory।