अध्याय 16

टेस्ट डिज़ाइन और TDD — अच्छा टेस्ट कैसा होता है

टेस्ट लिखने से पहले की योजना, Arrange-Act-Assert, implementation नहीं बल्कि व्यवहार को टेस्ट करना, red-green-refactor, Hypothesis के साथ property-based टेस्टिंग, flaky टेस्ट के कारण और इलाज, और हर बग के लिए एक regression टेस्ट।

60 मिनटPython 3.12
  1. 1समस्या
  2. 2समझें
  3. 3उदाहरण
  4. 4अनुमान
  5. 5स्वयं करें
  6. 6चुनौती

वह समस्या जिसे हम हल कर रहे हैं

यह कोर्स जो-जो टूल सिखाने निकला था, वे सब अब आप जानते हैं: assert, raises, parametrize, fixtures, mocks, markers, configuration, coverage, plugins। और यह पूरी तरह मुमकिन है कि आप इन सबका इस्तेमाल करें और फिर भी ऐसे टेस्ट लिख बैठें जो मदद कम और परेशानी ज़्यादा करें।

यह रहा एक छोटा-सा शॉपिंग कार्ट, और उसके दो टेस्ट जो पास होते हैं:

python
class Cart:
    def __init__(self):
        self._items = []

    def add(self, name, price, quantity=1):
        self._items.append((name, price, quantity))

    def total(self):
        return sum(price * quantity for _, price, quantity in self._items)
python
from cart import Cart


def test_add():
    cart = Cart()
    cart.add("pen", 15.0, 2)
    assert cart._items == [("pen", 15.0, 2)]


def test_total():
    cart = Cart()
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)
    assert cart.total() == 880.0
text
..                                                                       [100%]
2 passed in 0.01s

अब कोई कार्ट को बेहतर बनाता है। एक ही आइटम को दो बार जोड़ने पर वह एक ही लाइन में मिल जाना चाहिए, इसलिए लिस्ट की जगह नाम को key बनाकर एक डिक्शनरी आ जाती है। कार्ट अब भी चीज़ें जोड़ता है और अब भी उनका टोटल सही निकालता है — ग्राहक को दिखने वाला कुछ भी नहीं बदला:

python
class Cart:
    def __init__(self):
        self._lines = {}

    def add(self, name, price, quantity=1):
        _, already = self._lines.get(name, (price, 0))
        self._lines[name] = (price, already + quantity)

    def total(self):
        return sum(price * quantity for price, quantity in self._lines.values())
text
F.                                                                       [100%]
=================================== FAILURES ===================================
___________________________________ test_add ___________________________________

    def test_add():
        cart = Cart()
        cart.add("pen", 15.0, 2)
>       assert cart._items == [("pen", 15.0, 2)]
               ^^^^^^^^^^^
E       AttributeError: 'Cart' object has no attribute '_items'

test_cart.py:7: AssertionError
=========================== short test summary info ============================
FAILED test_cart.py::test_add - AttributeError: 'Cart' object has no attribut...
1 failed, 1 passed in 0.01s

टेस्ट लाल है, लेकिन बग कोई नहीं। टेस्ट यह नहीं देख रहा था कि कार्ट करता क्या है; वह यह देख रहा था कि कार्ट बना कैसे है। ऐसा टेस्ट तब फ़ेल होता है जब आप कोड सुधारते हैं, और तब चुप रहता है जब आप उसे तोड़ देते हैं — यानी अपने काम का ठीक उल्टा।

यह आख़िरी अध्याय किसी नए टूल के बारे में नहीं है। यह विवेक के बारे में है: अच्छा टेस्ट कैसा दिखता है, टेस्ट्स को कोड की दिशा तय करने कैसे दें, कंप्यूटर से आपके लिए टेस्ट केस कैसे बनवाएँ, और टेस्ट्स को बेवजह फ़ेल होने से कैसे रोकें।

इस अध्याय के अंत में आप कर पाएंगे

  • टेस्ट लिखने से पहले उसकी योजना बनाना: contract, cases, और क्या छोड़ना है
  • किसी टेस्ट को Arrange, Act, Assert में बाँटना, और उसे एक वाक्य की तरह नाम देना
  • व्यवहार (behaviour) के टेस्ट और implementation के टेस्ट में फ़र्क़ पहचानना, और दूसरे को पहले में बदलना
  • red-green-refactor चक्रों में काम करना, जिसमें हर red कदम पर असली failure हो
  • Hypothesis के साथ property-based टेस्ट लिखना और shrink किया गया failing उदाहरण पढ़ना
  • flaky टेस्ट्स के चार आम कारण पहचानना और उन्हें ठीक करना
  • हर बग रिपोर्ट को एक regression टेस्ट में बदलना

ज़रूरी शर्तें: async टेस्ट, hooks और plugins।


टेस्ट लिखने से पहले

ज़्यादातर ख़राब टेस्ट बुरी तरह लिखे नहीं जाते। उनकी योजना बुरी होती है — वे तब लिख दिए जाते हैं जब किसी ने यह तय ही नहीं किया कि कोड वादा क्या करता है। इसलिए कोई भी टेस्ट कोड लिखने से पहले, चार सवाल।

इस पूरे अध्याय में हम एक छोटा-सा फ़ंक्शन test-first तरीक़े से बनाएँगे: slugify, जो किसी पोस्ट के शीर्षक, जैसे "Café au lait", को URL के हिस्से "cafe-au-lait" में बदलता है।

1. Contract क्या है? इसे एक-दो वाक्यों में कहिए, वैसे जैसे कॉल करने वाला इसे देखता है: एक शीर्षक (`str`) दिए जाने पर एक slug लौटाओ: lowercase ASCII अक्षर और अंक, शब्द एक-एक hyphen से जुड़े हुए, किसी भी सिरे पर hyphen नहीं। accent वाले अक्षर अपने सादे अक्षर बन जाते हैं। कोई side effect नहीं — यह न किसी फ़ाइल को छूता है, न घड़ी को, न नेटवर्क को। ध्यान दीजिए कि contract में किन चीज़ों का ज़िक्र नहीं है: regular expressions, Unicode normalisation, helper functions। ये बताते हैं कि काम कैसे होता है, और इन्हें बदलने की पूरी छूट है।

2. पहले से क्या तैयार होना चाहिए? एक virtual environment जिसमें टूल्स इंस्टॉल हों, और कोड ऐसा हो कि टेस्ट्स उसे import कर सकें:

text
$ python -m venv .venv
$ source .venv/bin/activate
$ pip install pytest hypothesis

फिर दो फ़ाइलें अगल-बगल — slugs.py और test_slugs.py — और pytest उसी फ़ोल्डर से चलाया जाए, ताकि from slugs import slugify काम करे। इस विषय के लिए न fixtures चाहिए, न अस्थायी फ़ाइलें, न environment variables। यह भी ध्यान देने लायक़ है: pure function दुनिया की सबसे आसानी से टेस्ट होने वाली चीज़ है, और यही logic को pure functions में डालने की एक अच्छी वजह है।

3. कौन-से cases? पहले happy path, फिर किनारे के cases, फिर invalid input, फिर boundaries:

| Case | Input | Expected | |---|---|---| | सामान्य शब्द | "Hello World" | "hello-world" | | विराम चिह्न | "Hello, World!" | "hello-world" | | accent वाले अक्षर | "Café au lait" | "cafe-au-lait" | | सिरों पर अतिरिक्त spaces | " Hello World " | "hello-world" | | अंक बने रहते हैं | "Top 10 Tips" | "top-10-tips" | | काम का कुछ भी नहीं | "!!!" | ? |

वह प्रश्नचिह्न इस टेबल का सबसे उपयोगी cell है। योजना लिखने से एक ऐसा फ़ैसला सामने आ गया है जो किसी ने लिया ही नहीं: जब शीर्षक में एक भी अक्षर न हो, तब क्या होना चाहिए? हम इसे जानबूझकर खुला छोड़ेंगे, और इसी अध्याय में आगे देखेंगे कि एक खुला सवाल कितना महँगा पड़ता है।

4. आप क्या टेस्ट नहीं करेंगे? पायथन के re मॉड्यूल, unicodedata और str.lower() के अपने टेस्ट पहले से हैं; उन्हें दोबारा टेस्ट करना बस आपकी रफ़्तार धीमी करता है। Private helpers (जो कुछ भी _ से शुरू होता है) तक slugify के ज़रिए पहुँचा जाता है, सीधे कभी नहीं। और आप हर Unicode character को हाथ से गिनाने की कोशिश नहीं करेंगे — वह काम property-based टेस्ट का है, जो इस अध्याय में आगे आएगा।

यह टेबल ही टेस्ट बन जाती है। नीचे का सब कुछ इसी योजना को लागू करता है।

Arrange, Act, Assert

हर अच्छे टेस्ट के एक जैसे तीन हिस्से होते हैं, एक ही क्रम में:

  • Arrange — वह दुनिया तैयार करें जिसकी टेस्ट को ज़रूरत है
  • Act — वह एक काम करें जिसे टेस्ट किया जा रहा है
  • Assert — जाँचें कि नतीजा क्या निकला

यह रहा कार्ट फिर से टेस्ट किया हुआ — इस बार उसके वादों के ज़रिए, इस बात से नहीं कि वह चीज़ें कैसे सहेजता है:

python
from cart import Cart


def test_an_empty_cart_totals_zero():
    cart = Cart()

    assert cart.total() == 0


def test_total_is_price_times_quantity_summed_over_items():
    # Arrange
    cart = Cart()
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)

    # Act
    total = cart.total()

    # Assert
    assert total == 880.0


def test_adding_the_same_item_twice_adds_up_the_quantity():
    cart = Cart()
    cart.add("pen", 15.0, 2)

    cart.add("pen", 15.0, 1)

    assert cart.total() == 45.0

लिस्ट वाले कार्ट पर, और फिर डिक्शनरी वाले कार्ट पर:

text
...                                                                      [100%]
3 passed in 0.01s
...                                                                      [100%]
3 passed in 0.01s

वही टेस्ट, दो implementations, और दोनों बार सब हरे। यही वह गुण है जो आप चाहते हैं: जो refactor व्यवहार को बनाए रखे, उसे टेस्ट्स को भी हरा बनाए रखना चाहिए। बीच वाले टेस्ट में comments सिर्फ़ ढाँचा दिखाने के लिए हैं; ज़्यादातर लोग इन्हें हटा देते हैं और तीनों हिस्सों को ख़ाली लाइनों से अलग करते हैं, जैसे तीसरे टेस्ट में।

Act का एक ही कदम होना क्यों मायने रखता है? क्योंकि जब टेस्ट फ़ेल हो, तो आप जानना चाहते हैं कि कौन-सी चीज़ टूटी। अगर Act में पाँच calls हैं, तो failure पाँच संदिग्धों की ओर इशारा करता है।

व्यवहार टेस्ट करें, implementation नहीं

मोटा नियम: टेस्ट सिर्फ़ उसी चीज़ का इस्तेमाल कर सकता है जिसका इस्तेमाल कॉल करने वाला कर सकता है। कार्ट के लिए वह है Cart(), .add() और .total()। न _items, न _lines। आगे लगा underscore पायथन का यह कहने का तरीक़ा है कि "यह मेरा है, और मैं इसे बदल सकता हूँ"।

ग़लत तरीक़ा टेस्ट को एक ऐसे फ़ैसले से बाँध देता है जो कभी contract का हिस्सा था ही नहीं — कि आइटम एक लिस्ट में रखे जाते हैं। सही तरीक़ा वही सवाल पूछता है जो कॉल करने वाला पूछेगा: अगर मैं ये चीज़ें जोड़ूँ, तो टोटल क्या होगा?

दो संकेत कि कोई टेस्ट implementation से बँधा हुआ है:

  • वह _ से शुरू होने वाले attributes पढ़ता है, या किसी ऐसे फ़ंक्शन को mock करता है जिसे कोड अंदर ही अंदर कॉल करता है
  • कोई ऐसा refactor, जिसे कोई user नोटिस नहीं कर सकता, उसे लाल कर देता है

यहाँ mocks पर एक बात कहनी ज़रूरी है। अध्याय ग्यारह के mocks किसी boundary पर सही टूल हैं — नेटवर्क, घड़ी, कोई payment service। अपने ही कोड के अंदर इस्तेमाल किए जाएँ ("assert करो कि _calculate इन arguments के साथ कॉल हुआ"), तो वे दूसरे नाम से implementation टेस्ट ही बन जाते हैं।

ऐसे नाम जो वाक्य की तरह पढ़े जाएँ, और फ़ेल होने की एक ही वजह

तीनों कार्ट टेस्ट -v के साथ चलाइए:

text
test_cart_behaviour.py::test_an_empty_cart_totals_zero PASSED            [ 33%]
test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items PASSED [ 66%]
test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity PASSED [100%]

सिर्फ़ नाम पढ़िए, और आपके पास कार्ट का specification है। test_add और test_total सिर्फ़ यह बताते थे कि कौन-सा मेथड छुआ गया। अच्छा नाम स्थिति और अपेक्षित नतीजा बताता है: ख़ाली कार्ट का टोटल शून्य होता है। नाम लंबा है; कोई बात नहीं। आप इसे कभी टाइप नहीं करते, और जब भी यह फ़ेल होता है, आप इसे पढ़ते हैं।

नियम का दूसरा हिस्सा है फ़ेल होने की एक ही वजह। यह रहा इसका उल्टा — एक ही टेस्ट जो सब कुछ जाँचता है, एक ऐसे कार्ट पर चलाया गया जिसमें बग है (वह quantity को अनदेखा करता है):

python
from cart import Cart


def test_cart():
    cart = Cart()
    assert cart.total() == 0
    cart.add("pen", 15.0, 2)
    cart.add("bag", 850.0)
    assert cart.total() == 880.0
    cart.add("pen", 15.0, 1)
    assert cart.total() == 895.0
text
$ pytest -q --tb=no
F                                                                        [100%]
=========================== short test summary info ============================
FAILED test_one_big.py::test_cart - assert 865.0 == 880.0
1 failed in 0.01s

(--tb=no tracebacks को छिपा देता है और सिर्फ़ summary छोड़ता है — वह हिस्सा जिसे आप सबसे पहले पढ़ते हैं।) test_cart फ़ेल हुआ — जिससे कुछ पता नहीं चलता — 865.0 == 880.0 के साथ, जिस पर अब आपको माथापच्ची करनी होगी। और तीसरा assertion कभी चला ही नहीं, इसलिए आप नहीं जानते कि वह पास होता या नहीं। अब वही बग तीनों focused टेस्ट्स पर:

text
$ pytest -q --tb=no
.FF                                                                      [100%]
=========================== short test summary info ============================
FAILED test_cart_behaviour.py::test_total_is_price_times_quantity_summed_over_items
FAILED test_cart_behaviour.py::test_adding_the_same_item_twice_adds_up_the_quantity
2 failed, 1 passed in 0.01s

एक भी traceback खोलने से पहले ही summary किसी निदान की तरह पढ़ी जाती है: ख़ाली कार्ट ठीक है, लेकिन quantities में गड़बड़ है। "फ़ेल होने की एक ही वजह" का मतलब एक ही assert लाइन नहीं है — एक ही नतीजे के दो पहलू जाँचने वाले दो asserts ठीक हैं। इसका मतलब है हर टेस्ट में एक ही व्यवहार।

टेस्ट पिरामिड

टेस्ट अलग-अलग आकार के होते हैं:

  • Unit टेस्ट एक फ़ंक्शन या class को जाँचते हैं, उसके आसपास कुछ भी असली नहीं होता। हर एक कुछ मिलीसेकंड का। ये आपके पास हज़ारों होते हैं।
  • Integration टेस्ट जाँचते हैं कि टुकड़े आपस में ठीक बैठते हैं: आपका कोड किसी असली database, असली file system, या test client के ज़रिए किसी असली HTTP app के साथ। धीमे; ये आपके पास दर्जनों से लेकर सैकड़ों तक होते हैं।
  • End-to-end टेस्ट पूरे सिस्टम को वैसे चलाते हैं जैसे कोई user चलाएगा — एक browser, एक deployed API। धीमे और नाज़ुक; ये आपके पास मुट्ठी भर होते हैं, जो उन रास्तों को कवर करते हैं जिनसे कमाई होती है।

गिनती के हिसाब से बनाइए, तो यह एक पिरामिड है: नीचे unit टेस्ट्स का चौड़ा आधार, ऊपर end-to-end टेस्ट्स की पतली चोटी। वजह है लागत। जब कोई unit टेस्ट फ़ेल होता है, तो वह कुछ लाइनों की ओर इशारा करता है। जब कोई end-to-end टेस्ट फ़ेल होता है, तो वह पूरे सिस्टम की ओर इशारा करता है। हर जाँच को पिरामिड में जितना नीचे ले जा सकें, ले जाइए, और चोटी को उन चीज़ों के लिए रखिए जिन्हें सिर्फ़ चोटी ही देख सकती है।

Red, green, refactor

Test-driven development (TDD) एक कसे हुए चक्र में चलता है:

  1. Red — ऐसे व्यवहार के लिए एक छोटा टेस्ट लिखिए जो अभी मौजूद नहीं है, और उसे फ़ेल होते देखिए
  2. Green — वह कम से कम कोड लिखिए जिससे वह पास हो जाए
  3. Refactor — कोड को व्यवस्थित कीजिए, और इस पूरे समय टेस्ट्स हरे रहें

उसे फ़ेल होते क्यों देखें? क्योंकि जिस टेस्ट को आपने कभी फ़ेल होते नहीं देखा, हो सकता है वह कुछ भी टेस्ट न कर रहा हो। red कदम साबित करता है कि टेस्ट feature की ग़ैरमौजूदगी को पकड़ सकता है।

हम टेबल से योजना लेते हैं, एक बार में एक row।

Red. पहली row, slugs.py के बनने से पहले:

python
from slugs import slugify


def test_lowercases_and_joins_words_with_hyphens():
    assert slugify("Hello World") == "hello-world"
text
==================================== ERRORS ====================================
________________________ ERROR collecting test_slugs.py ________________________
ImportError while importing test module '/home/you/blog/test_slugs.py'.
Hint: make sure your test modules/packages have valid Python names.
Traceback:
/usr/lib/python3.12/importlib/__init__.py:90: in import_module
    return _bootstrap._gcd_import(name[level:], package, level)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
test_slugs.py:1: in <module>
    from slugs import slugify
E   ModuleNotFoundError: No module named 'slugs'
=========================== short test summary info ============================
ERROR test_slugs.py
!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
1 error in 0.01s

यह एकदम ठीक red है। टेस्ट कुछ ऐसा माँग रहा है जो मौजूद नहीं है।

Green. पास कराने वाला कम से कम कोड — और यह सच में कम से कम है:

python
def slugify(title):
    return title.lower().replace(" ", "-")
text
.                                                                        [100%]
1 passed in 0.01s

यह धोखा जैसा लगता है। पर यह जानबूझकर है: कोड सिर्फ़ उतना ही करता है जितना अब तक किसी टेस्ट ने माँगा है। इससे ज़्यादा कुछ भी ऐसा कोड होगा जिसे कोई टेस्ट नहीं जाँच रहा।

Red. टेबल की दूसरी row:

python
def test_drops_punctuation():
    assert slugify("Hello, World!") == "hello-world"
text
.F                                                                       [100%]
=================================== FAILURES ===================================
____________________________ test_drops_punctuation ____________________________

    def test_drops_punctuation():
>       assert slugify("Hello, World!") == "hello-world"
E       AssertionError: assert 'hello,-world!' == 'hello-world'
E
E         - hello-world
E         + hello,-world!
E         ?      +      +

test_slugs.py:9: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_drops_punctuation - AssertionError: assert 'hello,...
1 failed, 1 passed in 0.01s

? वाली लाइन ठीक उन्हीं दो characters को चिह्नित करती है जो वहाँ नहीं होने चाहिए।

Green. अब एक space बदलने वाली तरकीब काफ़ी नहीं है। "अक्षर या अंक नहीं" वाले हर सिलसिले को एक hyphen से बदलिए, फिर सिरों से hyphens हटा दीजिए:

python
import re


def slugify(title):
    return re.sub(r"[^a-z0-9]+", "-", title.lower()).strip("-")
text
..                                                                       [100%]
2 passed in 0.01s

अतिरिक्त spaces और अंकों वाली rows इस कोड के साथ पहले से पास हो जाएँगी। फिर भी उन्हें जोड़िए — वे contract का हिस्सा हैं — लेकिन यह जान लीजिए कि जो टेस्ट पहली बार चलते ही हरा हो, उसने आपके अभी लिखे कोड के बारे में कुछ भी साबित नहीं किया। यह सिर्फ़ उस टेस्ट ने किया है जो पहले लाल और फिर हरा हुआ।

Red. accent वाले अक्षर:

python
def test_turns_accented_letters_into_plain_ones():
    assert slugify("Café au lait") == "cafe-au-lait"
text
..F                                                                      [100%]
=================================== FAILURES ===================================
_________________ test_turns_accented_letters_into_plain_ones __________________

    def test_turns_accented_letters_into_plain_ones():
>       assert slugify("Café au lait") == "cafe-au-lait"
E       AssertionError: assert 'caf-au-lait' == 'cafe-au-lait'
E
E         - cafe-au-lait
E         ?    -
E         + caf-au-lait

test_slugs.py:13: AssertionError
=========================== short test summary info ============================
FAILED test_slugs.py::test_turns_accented_letters_into_plain_ones - Assertion...
1 failed, 2 passed in 0.01s

é a-z में नहीं है, इसलिए वह विराम चिह्नों के साथ ही फेंक दिया गया।

Green. Unicode "NFKD" normalisation é को e और एक अलग accent चिह्न में तोड़ देता है; फिर "ignore" के साथ ASCII में encode करने पर चिह्न हट जाता है और e बचा रहता है:

python
import re
import unicodedata


def slugify(title):
    plain = unicodedata.normalize("NFKD", title).encode("ascii", "ignore").decode()
    return re.sub(r"[^a-z0-9]+", "-", plain.lower()).strip("-")
text
...                                                                      [100%]
3 passed in 0.01s

Refactor. यह काम करता है, लेकिन body एक ठसाठस भरी लाइन है। हर कदम को एक नाम दीजिए, pattern को एक बार compile कीजिए, और type hints जोड़िए — व्यवहार में कोई बदलाव किए बिना:

python
import re
import unicodedata

NOT_ALLOWED = re.compile(r"[^a-z0-9]+")


def _to_ascii(text: str) -> str:
    """Split accented letters into letter + accent, then drop the accents."""
    decomposed = unicodedata.normalize("NFKD", text)
    return decomposed.encode("ascii", "ignore").decode("ascii")


def slugify(title: str) -> str:
    words = NOT_ALLOWED.sub("-", _to_ascii(title).lower())
    return words.strip("-")
text
test_slugs.py::test_lowercases_and_joins_words_with_hyphens PASSED       [ 33%]
test_slugs.py::test_drops_punctuation PASSED                             [ 66%]
test_slugs.py::test_turns_accented_letters_into_plain_ones PASSED        [100%]

============================== 3 passed in 0.01s ===============================

यही वह कदम है जिससे TDD का फ़ायदा मिलता है। आप खुलकर ढाँचा बदल सके क्योंकि टेस्ट व्यवहार जाँचते हैं; अगर वे _to_ascii या regex के अंदर पहुँचते, तो refactor उन्हें तोड़ देता।

Hypothesis के साथ property-based testing

अब तक हर टेस्ट में एक गढ़ा हुआ उदाहरण है: "Hello World", "Café au lait"। उदाहरण उतने ही अच्छे होते हैं जितनी आपकी कल्पना, और आपकी कल्पना के blind spots वही होते हैं जो आपके अभी लिखे कोड के।

Property एक ऐसा कथन है जो हर input पर सच हो। Hypothesis inputs ख़ुद बनाता है — default रूप से हर टेस्ट के लिए सौ — और पूरी कोशिश करता है कि कोई ऐसा input मिल जाए जो कथन को तोड़ दे। @given बताता है कि inputs कहाँ से आएँगे; strategies (जिसे हमेशा st के रूप में import किया जाता है) उनका आकार बताता है:

python
from hypothesis import given
from hypothesis import strategies as st


@given(
    prices=st.lists(st.integers(min_value=0, max_value=10_000), max_size=20),
    discount=st.integers(min_value=0, max_value=100),
)
def test_a_discount_never_raises_the_total(prices, discount):
    total = sum(prices)

    discounted = total * (100 - discount) // 100

    assert 0 <= discounted <= total
text
.                                                                        [100%]
1 passed in 0.01s

एक बिंदु, लेकिन इससे सौ price lists और सौ discounts गुज़रे। दूसरी strategies जिनकी आपको ज़रूरत पड़ेगी: st.text(), st.floats(), st.booleans(), st.sampled_from([...]), st.dictionaries(...), st.builds(...)।

slugify की properties क्या हैं? आप यह नहीं बता सकते कि किसी random string का slug क्या है, लेकिन यह बता सकते हैं कि वह कैसा दिखना चाहिए, और यह कि किसी slug को slugify करने से कुछ नहीं बदलता:

python
import re

from hypothesis import given
from hypothesis import strategies as st

from slugs import slugify


@given(st.text())
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
    slug = slugify(title)

    assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*|", slug)


@given(st.text())
def test_slugifying_a_slug_changes_nothing(title):
    once = slugify(title)

    assert slugify(once) == once
text
..                                                                       [100%]
2 passed in 0.01s

st.text() emoji, चीनी अक्षर, control characters, ख़ाली strings बनाता है — ऐसे inputs जो आप योजना वाली टेबल में कभी टाइप न करते।

एक round trip असली बग पकड़ता है

सबसे उपयोगी property है round trip: अगर आप किसी चीज़ को encode करके फिर decode कर सकते हैं, तो encoding को decode करने पर मूल चीज़ वापस मिलनी चाहिए। यह रहा एक run-length encoder — "aaab" बन जाता है "3a1b" — दो उदाहरण टेस्ट्स के साथ जो पास होते हैं:

python
import re


def encode(text: str) -> str:
    """'aaab' -> '3a1b': each run of a character becomes count + character."""
    out = []
    for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
        run = match.group(0)
        out.append(f"{len(run)}{run[0]}")
    return "".join(out)


def decode(encoded: str) -> str:
    return "".join(char * int(count) for count, char in re.findall(r"(\d+)(\D)", encoded))
python
from hypothesis import given
from hypothesis import strategies as st

from rle import decode, encode


def test_encode_counts_each_run():
    assert encode("aaab") == "3a1b"


def test_decode_expands_each_run():
    assert decode("3a1b") == "aaab"


@given(st.text())
def test_decoding_an_encoding_gives_back_the_original(text):
    assert decode(encode(text)) == text
text
..F                                                                      [100%]
=================================== FAILURES ===================================
______________ test_decoding_an_encoding_gives_back_the_original _______________

    @given(st.text())
>   def test_decoding_an_encoding_gives_back_the_original(text):
                   ^^^

test_rle.py:16:
_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _

text = '0'

    @given(st.text())
    def test_decoding_an_encoding_gives_back_the_original(text):
>       assert decode(encode(text)) == text
E       AssertionError: assert '' == '0'
E
E         - 0
E       Failing test case: test_decoding_an_encoding_gives_back_the_original(
E           text='0',
E       )

test_rle.py:17: AssertionError
=========================== short test summary info ============================
FAILED test_rle.py::test_decoding_an_encoding_gives_back_the_original - Asser...
1 failed, 2 passed in 0.01s

टेक्स्ट "0" encode होकर "10" बनता है — "एक शून्य" — और decoder 10 को एक count पढ़ता है जिसके बाद कोई character नहीं है। अंक वाला कोई भी टेक्स्ट बिगड़ जाता है। दोनों उदाहरण टेस्ट्स में सिर्फ़ अक्षर थे, इसलिए वे इसे कभी देख ही नहीं सकते थे।

Shrinking

Hypothesis को पहली बार में "0" नहीं मिला था। उसे कोई लंबी, बेढंगी string मिली, फिर उसने उसे shrink किया: छोटे और सरल inputs आज़माए, हर उस input को रखा जो अब भी फ़ेल होता था, जब तक कि कोई भी और सरल input फ़ेल नहीं हुआ। हर failing input को दर्ज करके आप इसे होते हुए देख सकते हैं:

python
from hypothesis import given, seed, settings
from hypothesis import strategies as st

from rle import decode, encode

failing = []


@seed(2026)
@settings(database=None)
@given(st.text())
def round_trip(text):
    if decode(encode(text)) != text:
        if text not in failing:
            failing.append(text)
        raise AssertionError


try:
    round_trip()
except AssertionError:
    pass

for text in failing:
    print(repr(text))
text
'wfê\x9f\x0c\U000b6082Ñ/\x879\x9dÂ\n'
'\x03\U000a1acb\U0003c4fc1'
'´\U000cd3dc\U00081acfÔ÷\U0007122f\x89ê½5'
'l𨺮\nÝÛ\x894\x8b'
'\U00083189𤾨«ñï\xa0^ó8A\x96\x88=\x05'
'0000'
'000'
'00'
'0'

(@seed random चुनावों को स्थिर कर देता है ताकि यह run दोहराया जा सके; database=None Hypothesis को याद रखी गई failure दोबारा चलाने से रोकता है।) पहली failure तेरह characters का शोर है जिसमें एक 9 छिपा है। आप मिनटों तक उसे घूरते रहते। Shrink किया गया उदाहरण पूरी कहानी एक character में कह देता है: अंक इसे तोड़ देता है।

Hypothesis failing उदाहरणों को एक .hypothesis/ फ़ोल्डर में सहेज भी लेता है, ताकि अगला run सबसे पहले "0" आज़माए। इस फ़ोल्डर को .gitignore में जोड़ दीजिए।

सुधार count और character के बीच एक separator रख देता है, ताकि कोई अंक कभी count का हिस्सा न समझा जाए। और Hypothesis का खोजा हुआ उदाहरण @example से पक्का कर दिया जाता है, ताकि वह हर बार, हर मशीन पर चले:

python
import re


def encode(text: str) -> str:
    """'aaab' -> '3:a1:b': each run becomes count, a colon, then the character."""
    out = []
    for match in re.finditer(r"(.)\1*", text, flags=re.DOTALL):
        run = match.group(0)
        out.append(f"{len(run)}:{run[0]}")
    return "".join(out)


def decode(encoded: str) -> str:
    pairs = re.findall(r"(\d+):(.)", encoded, flags=re.DOTALL)
    return "".join(char * int(count) for count, char in pairs)

दोनों उदाहरण टेस्ट नए format ("3:a1:b") में बदल जाते हैं, और property में एक लाइन जुड़ जाती है:

python
from hypothesis import example, given
from hypothesis import strategies as st

from rle import decode, encode


@given(st.text())
@example("0")  # the input Hypothesis found; now it is checked on every run
def test_decoding_an_encoding_gives_back_the_original(text):
    assert decode(encode(text)) == text
text
...                                                                      [100%]
3 passed in 0.01s

Properties उदाहरणों की जगह नहीं लेतीं। दोनों उदाहरण टेस्ट format को इस तरह दर्ज करते हैं कि पढ़ने वाला एक नज़र में समझ जाए; property उन कोनों की रखवाली करती है जिनके बारे में किसी ने सोचा नहीं। दोनों का इस्तेमाल कीजिए।

Flaky टेस्ट

Flaky टेस्ट वह है जो उसी कोड पर कभी पास और कभी फ़ेल होता है। यह किसी टेस्ट के न होने से भी बुरा है: लोग "re-run" दबाना सीख जाते हैं और लाल पर भरोसा करना छोड़ देते हैं। लगभग हर flaky टेस्ट चार में से किसी एक कारण से आता है, और हर एक का ऐसा सुधार है जो कारण को छिपाने के बजाय उसे हटा देता है।

Shared state और टेस्ट्स का क्रम

python
_users = set()


def register(name):
    _users.add(name)


def count():
    return len(_users)
python
from registry import count, register


def test_registering_a_user_counts_them():
    register("rahim")

    assert count() == 1


def test_a_new_registry_is_empty():
    assert count() == 0
text
$ pytest -q --tb=no
.F                                                                       [100%]
=========================== short test summary info ============================
FAILED test_registry.py::test_a_new_registry_is_empty - assert 1 == 0
1 failed, 1 passed in 0.01s

$ pytest -q "test_registry.py::test_a_new_registry_is_empty"
.                                                                        [100%]
1 passed in 0.01s

suite में फ़ेल, अकेले में पास। module-level set एक टेस्ट से अगले टेस्ट तक बचा रहता है, इसलिए दूसरा टेस्ट वह देखता है जो पहला पीछे छोड़ गया। क्रम बदलिए, या pytest-xdist से टेस्ट्स समानांतर चलाइए, और नतीजा बदल जाता है। ग़लत सुधार है टेस्ट्स का क्रम बदलना। सही सुधार shared state को हटा देता है: registry को एक object बनाइए, और fixture के ज़रिए हर टेस्ट को एक नया object दीजिए:

python
class Registry:
    def __init__(self):
        self._users = set()

    def register(self, name):
        self._users.add(name)

    def count(self):
        return len(self._users)
python
import pytest

from registry import Registry


@pytest.fixture
def registry():
    return Registry()  # a fresh one for every test


def test_registering_a_user_counts_them(registry):
    registry.register("rahim")

    assert registry.count() == 1


def test_a_new_registry_is_empty(registry):
    assert registry.count() == 0
text
..                                                                       [100%]
2 passed in 0.01s

जब आप कोड नहीं बदल सकते, तब आख़िरी रास्ता है एक autouse fixture जो हर टेस्ट से पहले और बाद में state को reset करे। हर टेस्ट अकेले में भी पास होना चाहिए, और किसी भी क्रम में भी।

समय

python
from datetime import datetime


def greeting(now=None):
    now = now or datetime.now()
    return "Good morning" if now.hour < 12 else "Good afternoon"
python
from greeting import greeting


def test_greets_the_morning():
    assert greeting() == "Good morning"  # true only before noon

वही कोड, उसी पल, अलग-अलग time zones पर सेट दो मशीनों पर चलाया गया:

text
$ TZ=Europe/London pytest -q
.                                                                        [100%]
1 passed in 0.01s

$ TZ=Asia/Dhaka pytest -q --tb=no
F                                                                        [100%]
=========================== short test summary info ============================
FAILED test_greeting.py::test_greets_the_morning - AssertionError: assert 'Go...
1 failed in 0.01s

सुधार फ़ंक्शन के signature में पहले से मौजूद है: now को बाहर से पास किया जा सकता है। जो टेस्ट घड़ी को नियंत्रित करता है, वह boundary के दोनों तरफ़ जाँचता है, और किसी भी घंटे पर एक ही जवाब देता है:

python
from datetime import datetime

from greeting import greeting


def test_before_noon_it_says_good_morning():
    assert greeting(now=datetime(2026, 1, 5, 9, 30)) == "Good morning"


def test_from_noon_on_it_says_good_afternoon():
    assert greeting(now=datetime(2026, 1, 5, 12, 0)) == "Good afternoon"
text
..                                                                       [100%]
2 passed in 0.01s

समय को पास करना किसी भी mock से सरल है। जब आप signature नहीं बदल सकते, तब अध्याय दस का monkeypatch घड़ी को बदल देता है।

Randomness

python
import random


def pick_winner(names, rng=random):
    return rng.choice(names)
python
from raffle import pick_winner


def test_picks_a_winner():
    assert pick_winner(["rahim", "karim", "salma"]) == "rahim"

छह runs, कहीं कोई बदलाव नहीं:

text
$ for i in 1 2 3 4 5 6; do pytest -q | tail -1; done
1 failed in 0.01s
1 failed in 0.01s
1 failed in 0.01s
1 passed in 0.01s
1 passed in 0.01s
1 passed in 0.01s

दो सुधार, दो तरह के सवालों के लिए। वह assert कीजिए जो हर नतीजे के लिए सच हो — विजेता प्रतिभागियों में से ही कोई है। या एक seeded generator पास करके randomness को अपने नियंत्रण में लीजिए:

python
import random

from raffle import pick_winner


def test_the_winner_is_one_of_the_entrants():
    names = ["rahim", "karim", "salma"]

    assert pick_winner(names) in names


def test_the_same_seed_picks_the_same_winner():
    names = ["rahim", "karim", "salma"]

    first = pick_winner(names, rng=random.Random(42))
    second = pick_winner(names, rng=random.Random(42))

    assert first == second
text
$ for i in 1 2 3; do pytest -q | tail -1; done
2 passed in 0.01s
2 passed in 0.01s
2 passed in 0.01s

बाहरी दुनिया

चौथा कारण है हर वह चीज़ जो आपके नियंत्रण में नहीं है: असली नेटवर्क call, असली server, कोई sleep(0.1) जो "काफ़ी होना चाहिए"। सुधार वही हैं जो पिछले अध्यायों में आए — boundary पर mock कीजिए, साझा फ़ोल्डर के बजाय tmp_path इस्तेमाल कीजिए, तय समय के बजाय किसी शर्त का इंतज़ार कीजिए। एक ही सिद्धांत चारों कारणों पर लागू होता है: टेस्ट जिस चीज़ पर भी निर्भर है, उसे टेस्ट ख़ुद बनाए या नियंत्रित करे।

हर बग के लिए एक regression टेस्ट

योजना वाली टेबल के प्रश्नचिह्न पर लौटते हैं। किसी ने उसका जवाब नहीं दिया, और बग रिपोर्ट आ जाती है: `"!!!"` शीर्षक वाली एक पोस्ट `/posts/` पर सहेजी गई, और उसने index page को overwrite कर दिया।

slugify को छूने से पहले एक ऐसा टेस्ट लिखिए जो रिपोर्ट को दोहराए, और उसे फ़ेल होते देखिए। इससे साबित होता है कि टेस्ट इसी बग को पकड़ता है:

python
import pytest

from slugs import slugify


def test_a_title_with_no_letters_or_digits_is_rejected():
    # Bug: slugify("!!!") returned "", and the post was saved at /posts/
    with pytest.raises(ValueError, match="no letters or digits"):
        slugify("!!!")
text
$ pytest -q test_slugs_regressions.py
F                                                                        [100%]
=================================== FAILURES ===================================
______________ test_a_title_with_no_letters_or_digits_is_rejected ______________

    def test_a_title_with_no_letters_or_digits_is_rejected():
        # Bug: slugify("!!!") returned "", and the post was saved at /posts/
>       with pytest.raises(ValueError, match="no letters or digits"):
E       Failed: DID NOT RAISE ValueError

test_slugs_regressions.py:8: Failed
=========================== short test summary info ============================
FAILED test_slugs_regressions.py::test_a_title_with_no_letters_or_digits_is_rejected
1 failed in 0.01s

फिर इसे ठीक कीजिए — slugify का अंत यह बन जाता है:

python
def slugify(title: str) -> str:
    slug = NOT_ALLOWED.sub("-", _to_ascii(title).lower()).strip("-")
    if not slug:
        raise ValueError(f"title has no letters or digits: {title!r}")
    return slug
text
....                                                                     [100%]
4 passed in 0.01s

regression टेस्ट और तीनों उदाहरण टेस्ट पास होते हैं। लेकिन पूरी suite चलाइए, और property टेस्ट आपत्ति करते हैं:

text
FAILED test_slugs_properties.py::test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens
FAILED test_slugs_properties.py::test_slugifying_a_slug_changes_nothing - Val...
2 failed, 4 passed in 0.01s

Hypothesis ने ख़ाली string आज़माई, और slugify("") अब exception raise करता है। यह कोई नया बग नहीं है — यह contract का जानबूझकर बदलना है, और properties ने उसे नोटिस किया है। अच्छी बात है। उन्हें नए contract के मुताबिक़ अपडेट कीजिए: ऐसे शीर्षक जिनमें कम से कम एक अक्षर या अंक हो।

python
import re
import string

from hypothesis import given
from hypothesis import strategies as st

from slugs import slugify

# Any text, with at least one plain letter or digit somewhere inside it.
titles = st.builds(
    lambda before, word, after: before + word + after,
    st.text(),
    st.text(alphabet=string.ascii_letters + string.digits, min_size=1),
    st.text(),
)


@given(titles)
def test_a_slug_only_holds_lowercase_letters_digits_and_inner_hyphens(title):
    slug = slugify(title)

    assert re.fullmatch(r"[a-z0-9]+(-[a-z0-9]+)*", slug)

test_slugifying_a_slug_changes_nothing भी इसी तरह @given(titles) पर आ जाता है।

text
......                                                                   [100%]
6 passed in 0.01s

हर बग के लिए एक टेस्ट क्यों? क्योंकि जो बग एक बार हो चुका है, उसने साबित कर दिया है कि उसे करना आसान है, और उस कोड को अगली बार छूने वाला व्यक्ति शायद उसे फिर से कर देगा। क्या ग़लत हुआ था, यह बताने वाला comment टेस्ट का हिस्सा है: एक साल बाद, यही इकलौता रिकॉर्ड होगा कि यह अजीब-सा दिखने वाला case यहाँ क्यों है।


पूर्ण उदाहरण

एक password-strength checker, इस अध्याय की हर चीज़ से बना हुआ। पहले योजना:

| Case | Input | Expected | |---|---|---| | ख़ाली | "" | weak | | हर प्रकार, पर बहुत छोटा | "Ab1!xyz" (7) | weak | | पर्याप्त लंबा, एक प्रकार | "abcdefgh" | weak | | पर्याप्त लंबा, दो प्रकार | "abcdefg1" | medium | | 12 लंबा, तीन प्रकार | "abcdefghij1!" | strong | | 12 लंबा, दो प्रकार | "abcdefghijk1" | medium |

passwords.py:

python
def _kinds(password: str) -> int:
    """How many of the four kinds of character the password uses."""
    return sum([
        any(c.islower() for c in password),
        any(c.isupper() for c in password),
        any(c.isdigit() for c in password),
        any(not c.isalnum() for c in password),
    ])


def strength(password: str) -> str:
    """Rate a password as 'weak', 'medium' or 'strong'."""
    if len(password) < 8:
        return "weak"
    kinds = _kinds(password)
    if len(password) >= 12 and kinds >= 3:
        return "strong"
    if kinds >= 2:
        return "medium"
    return "weak"

test_passwords.py:

python
import pytest
from hypothesis import given
from hypothesis import strategies as st

from passwords import strength

RANK = {"weak": 0, "medium": 1, "strong": 2}


@pytest.mark.parametrize(
    "password, expected",
    [
        ("", "weak"),                      # nothing at all
        ("Ab1!xyz", "weak"),               # every kind, but only 7 long
        ("abcdefgh", "weak"),              # 8 long, one kind
        ("abcdefg1", "medium"),            # 8 long, two kinds
        ("abcdefghij1!", "strong"),        # 12 long, three kinds
        ("abcdefghijk1", "medium"),        # 12 long, only two kinds
    ],
)
def test_strength_follows_the_length_and_variety_rules(password, expected):
    assert strength(password) == expected


@given(st.text(), st.text())
def test_adding_characters_never_makes_a_password_weaker(password, extra):
    before = strength(password)

    after = strength(password + extra)

    assert RANK[after] >= RANK[before]


def test_a_long_password_of_one_kind_is_still_weak():
    # Bug: "aaaaaaaaaaaaaaaaaaaa" (20 letters) was rated "medium"
    assert strength("a" * 20) == "weak"
text
$ pytest -v
collected 8 items

test_passwords.py::test_strength_follows_the_length_and_variety_rules[-weak] PASSED [ 12%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[Ab1!xyz-weak] PASSED [ 25%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefgh-weak] PASSED [ 37%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefg1-medium] PASSED [ 50%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghij1!-strong] PASSED [ 62%]
test_passwords.py::test_strength_follows_the_length_and_variety_rules[abcdefghijk1-medium] PASSED [ 75%]
test_passwords.py::test_adding_characters_never_makes_a_password_weaker PASSED [ 87%]
test_passwords.py::test_a_long_password_of_one_kind_is_still_weak PASSED [100%]

============================== 8 passed in 0.01s ===============================

तीन फ़ैसले ध्यान देने लायक़ हैं।

टेबल की rows boundaries पर बैठी हैं — 7 और 8 characters, 11 और 12, दो प्रकार और तीन। बग boundaries पर रहते हैं: जो < असल में <= होना चाहिए था, वह किसी range के बीच में दिखाई ही नहीं देता।

Property वह कहती है जो कोई उदाहरण नहीं कह सकता: characters जोड़ने से कोई password कभी कमज़ोर नहीं होता। कोई भी यह टेस्ट हर password के लिए हाथ से नहीं लिखेगा, और Hypothesis इसे सैकड़ों passwords पर जाँचता है, जिनमें ऐसे Unicode अक्षर भी हैं जो न upper case हैं न lower case।

और कोई भी टेस्ट _kinds को नहीं छूता। उस तक strength के ज़रिए पहुँचा जाता है; हो सकता है कल वह रहे ही नहीं।


जब यह काम न करे

refactor के बाद AttributeError: 'Cart' object has no attribute '_items' टेस्ट object का अंदरूनी हिस्सा पढ़ रहा था। उसे public methods के ज़रिए जाँचने के लिए दोबारा लिखिए — यानी वह जो कॉल करने वाला देख सकता है — और वह अगले refactor में भी बचा रहेगा।

कोई टेस्ट अकेले में पास होता है लेकिन पूरे run में फ़ेल होता है (या इसका उल्टा) Shared state: कोई module-level list, dictionary या cache; किसी तय जगह पर रखी फ़ाइल; कोई environment variable जो set किया गया और कभी unset नहीं हुआ। टेस्ट्स का क्रम बदलने के बजाय हर टेस्ट से वह ख़ुद बनवाइए जिसकी उसे ज़रूरत है (fixtures, tmp_path, monkeypatch)।

hypothesis.errors.FailedHealthCheck: It looks like this test is filtering out a lot of inputs. 0 inputs were generated successfully, while 50 inputs were filtered out. कोई .filter() (या assume()) Hypothesis के बनाए लगभग सब कुछ को फेंक देता है — उदाहरण के लिए st.integers().filter(lambda n: n % 1000 == 7)। इसके बजाय जो वैल्यूज़ चाहिए, उन्हें सीधे बनाइए: st.integers().map(lambda n: n * 1000 + 7)।

hypothesis.errors.FlakyFailure: Hypothesis test_depends_on_earlier_runs(n=-25617) produces unreliable results: Failed on the first call but did not on a subsequent one टेस्ट ने एक ही input के लिए अलग जवाब दिया। generated arguments के बाहर की कोई चीज़ calls के बीच बदल गई — कोई counter, कोई module-level list, घड़ी। Hypothesis हर failure को दोबारा चलाता है, इसलिए छिपी state वाला टेस्ट shrink नहीं हो सकता।

hypothesis.errors.DeadlineExceeded: Test took 300.07ms, which exceeds the deadline of 200.00ms. default रूप से हर generated उदाहरण को 200 ms के भीतर पूरा होना चाहिए, क्योंकि सौ धीमे उदाहरण मिलकर suite को धीमा कर देते हैं। टेस्ट को तेज़ बनाइए, या अगर धीमापन वाक़ई ज़रूरी है, तो @settings(deadline=...) से सीमा बढ़ाइए।

नया TDD टेस्ट पहली बार चलाते ही पास हो जाता है यह उस कोड के बारे में कुछ भी साबित नहीं कर रहा जो आप लिखने वाले हैं। या तो वह व्यवहार पहले से मौजूद है (ठीक है — टेस्ट को documentation की तरह रख लीजिए), या टेस्ट वह नहीं जाँच रहा जो आप सोच रहे हैं। कुछ पल के लिए कोड को जानबूझकर तोड़िए और पक्का कीजिए कि टेस्ट लाल होता है।