Theory
Python List की Speed अड़चन
आपके Semester 2 BCA204 collections labs में, आपने सीखा कि Python lists बेहद flexible हैं क्योंकि वे mixed data types रखती हैं और dynamically फैलती हैं। पर, जब हज़ारों transaction amounts या laboratory observations सँभालते हैं, standard Python lists बहुत धीमी हो जाती हैं। क्योंकि list items individual type wrappers के साथ machine memory में बेतरतीब बिखरी होती हैं, उन्हें process करने के लिए धीमे element-दर-element iteration loops चाहिए। NumPy इस बिखरे storage layout को एक कठोर, single-type data block से पूरी तरह क्यों बदलता है जो mathematical calculations सैकड़ों गुना तेज़ चलाता है?
Theory
बिखरे Envelopes बनाम Concrete Egg Tray
एक standard Python list को एक घर के अलग कमरों में बिखरे individual postal envelopes के एक collection की तरह सोचिए; हर envelope में कागज़ की एक पर्ची होती है जो एक शब्द से लेकर एक decimal number तक कुछ भी रख सकती है। उन सबको पढ़ने के लिए, आपको हर जगह एक-एक करके जाना होगा। एक NumPy array एक भारी, ढले concrete egg tray की तरह है जहाँ हर एक pocket बिल्कुल एक ही आकार का है और केवल एक ख़ास प्रकार का object (जैसे integers) रखता है। क्योंकि pockets एक अकेले contiguous block में एक-दूसरे से बिल्कुल सटे बैठते हैं, एक machine पूरी row में एक high-speed operation में sweep कर सकती है।
Theory
NumPy Array Mechanics औपचारिक रूप से
NumPy (Numerical Python) library ndarray (N-dimensional array) पेश करती है, एक तेज़, space-efficient multidimensional sequence container। standard lists के उलट, एक NumPy array सख़्ती से homogeneous है, हर एक element को बिल्कुल एक ही data type (जैसे float64 या int32) साझा करना होगा। Arrays के पास एक fixed संरचनात्मक configuration होती है जो उनके shape attribute (हर dimension का आकार दर्शाता एक tuple) और उनके dtype (internal data type allocation wrapper) से परिभाषित होती है।
At a glance
Table 1: मुख्य NumPy संरचनात्मक initialization और shape transformation operations।
| Creation & Shape Tool | Functional Mathematical Action | PocketMoney Budget Blueprint |
|---|---|---|
| np.array(sequence) | standard sequences (lists/tuples) को एक तेज़ ndarray container में बदलता है। | np.array([30, 50, 120]) |
| np.arange(start, stop, step) | एक half-open range के अंदर समान दूरी पर values रखता एक array generate करता है। | np.arange(0, 100, 20) -> [0, 20, 40, 60, 80] |
| np.linspace(start, stop, num) | एक closed span पर समान दूरी वाली fractional values की एक निर्दिष्ट संख्या generate करता है। | np.linspace(0, 10, 5) -> [0. , 2.5, 5. , 7.5, 10.] |
| matrix.reshape(rows, cols) | internal sequential values बदले बिना dimension shapes को reconfigure करता है। | arr.reshape(2, 3) एक 6-item line को एक grid matrix में बदलता है। |
| matrix.flatten() | एक multidimensional grid matrix को वापस एक 1D sequence line में समेटता है। | grid.flatten() एक grid को वापस एक line sequence में बदलता है। |
Theory
Worked Example: Ledger Metrics को फिर से संरचित करना
आइए देखें कि हमारा चलता project, PocketMoney, financial logs को संरचित करने के लिए NumPy arrays कैसे इस्तेमाल करता है। हम tracking IDs की एक continuous range initialize करेंगे, संरचनात्मक grids calculate करेंगे, और overall financial habits का मूल्यांकन करने के लिए sum(), average(), min(), और max() जैसी high-speed aggregation methods लागू करेंगे।
Practical
PocketMoney Financial Array Matrix
import numpy as np
# Step 1: Initialize a list of transaction values into a fast homogeneous array
spends_arr = np.array([30, 45, 60, 25, 90, 50])
# Step 2: Restructure the 6-element line array into a 2-row, 3-column matrix grid
spend_matrix = spends_arr.reshape(2, 3)
# Step 3: Run fast structural statistical aggregations
total_outflow = np.sum(spend_matrix)
mean_expense = np.average(spend_matrix)
peak_purchase = np.max(spend_matrix)
print("Reshaped Matrix:\n", spend_matrix)
print("Total Sum of Outflow:", total_outflow)
print("Average Spend Metric:", mean_expense)
print("Maximum Single Spend:", peak_purchase)This example runs in Gri-Learn on the web, where you can edit it and see the output.
Think first
Grid Transformation को ट्रेस करें
NumPy array transformations का ध्यान से विश्लेषण कीजिए। .reshape(2, 3) call के दौरान elements rows और columns में कैसे बँटेंगे? calculated statistical values क्या हैं?
Show the answer
script output करेगी:
Reshaped Matrix:
[[30 45 60]
[25 90 50]]
Total Sum of Outflow: 300
Average Spend Metric: 50.0
Maximum Single Spend: 90
क्यों? numbers की 1D line row-दर-row एक 2x3 matrix grid structure में packed होती है। sum operation सारे elements जोड़ता है (30+45+60+25+90+50 = 300)। average calculation उस कुल sum को 6 कुल slots से भाग देता है, 50.0 return करते हुए। पूरे collection के अंदर maximum value स्पष्ट रूप से 90 के रूप में पहचानी जाती है।
Quiz
क्या होता है अगर आप बिल्कुल 8 elements रखते एक 1D NumPy array को declaration command matrix.reshape(3, 3) इस्तेमाल करके एक matrix structure में reshape करने की कोशिश करते हैं?
- यह एक 3x3 matrix बनाता है, अंतिम ग़ायब block slot को 0 से भरते हुए।
- यह 3x3 dimensions में साफ़-सुथरे fit होने के लिए अंतिम elements truncate करता है।
- यह एक ValueError फेंकता है क्योंकि target shape size को original array size से मेल खाना होगा।
- यह array को dynamically एक list collection में बदलता है।
Show the answer
यह एक ValueError फेंकता है क्योंकि target shape size को original array size से मेल खाना होगा।
.reshape() से array dimensions बदलते समय, नई rows और columns का कुल गुणा size original elements की count के बिल्कुल बराबर होना चाहिए। एक 3x3 layout को बिल्कुल 9 elements चाहिए। 8 elements को 9 slots में fit करने की कोशिश एक तुरंत runtime ValueError trigger करती है।
Quiz
इस generation expression पर विचार कीजिए: items = np.arange(5, 20, 5)। इस array object के अंदर कौन से elements generate होंगे?
- [5, 10, 15, 20]
- [5, 10, 15]
- [10, 15, 20]
- [5, 6, 7, 8, 9]
Show the answer
[5, 10, 15]
function np.arange() एक half-open mathematical boundary range इस्तेमाल करता है जहाँ upper stop limit हमेशा बाहर रखी जाती है। 5 से शुरू और 5 के एक step value से increment करना 5, 10, और 15 generate करता है, पर 20 boundary limit से टकराने से ठीक पहले रुकता है।
Watch out
Classic जाल: Text में Silent Conversion
university laboratory examinations में सबसे बार-बार marks-गँवाने वाली ग़लती mixed types वाली एक standard Python list (जैसे [10, 20, 'Samosa']) को सीधे np.array() में देना है। NumPy arrays प्रति cell अलग data types बनाए नहीं रख सकते। crash होने के बजाय, NumPy चुपचाप हर एक item को एक text string wrapper में बदल देता है, आपके sum() या average() जैसे mathematical tools को एक तुरंत crash trigger कराते हुए क्योंकि integers text के रूप में फिर से लिखे जाते हैं!
Theory
Matrices को Semester 3 से जोड़ना
High-performance vector matrices automated analytics की मूल mathematical layers बनाते हैं। Semester 3 (BCA302/BCA303) में, जब multidimensional data tables, images, या sensor feeds process करते हैं, आप अपने raw data records को NumPy arrays के अंदर लपेटेंगे। यह एक भी nested loop लिखे बिना पूरे dataset में एक साथ high-speed multi-row calculations सक्षम करता है।
Summary
Key takeaways
- NumPy ndarrays अत्यधिक processing speed के लिए numbers को computer memory के contiguous blocks में store करते हैं।
- Arrays सख़्ती से homogeneous हैं, यानी हर item को एक जैसा data primitive type साझा करना होगा।
- arange method निर्दिष्ट steps के साथ sequences बनाता है, जबकि linspace ranges को सटीक fractional slices में बाँटता है।
- reshape और flatten tools अंतर्निहित data elements reallocate किए बिना multidimensional सीमाएँ modify करते हैं।
- sum, average, min, और max जैसी statistical methods जटिल matrices में values तुरंत aggregate करती हैं।
- Memory Hook: Python lists बिखरे envelopes हैं; NumPy arrays ठोस concrete trays हैं!