# Durable Testing > Software testing fundamentals for developers and AI agents. The 7 testing principles, the SDLC and STLC, and the test pyramid, with hand-drawn diagrams and rules agents can follow. Built by Durable Quality (https://durableqa.xyz/). # The 7 testing principles Codified by the ISTQB, these seven principles hold for any language, framework or methodology. Treat them as constraints: a test strategy that ignores one of them will fail in a predictable way. | # | Principle | In one line | |---|---|---| | 1 | Testing shows the presence of defects, not their absence | Passing tests reduce risk; they never prove correctness. | | 2 | Exhaustive testing is impossible | Prioritise by risk and use test design techniques. | | 3 | Early testing saves time and money | Shift left: test requirements and designs, not just code. | | 4 | Defects cluster together | A few modules hold most defects; focus there. | | 5 | Tests wear out | Repeated tests stop finding new bugs; refresh them. | | 6 | Testing is context dependent | Rigour and technique depend on the product's risk. | | 7 | Absence-of-errors fallacy | Bug-free software can still be the wrong software. | ## 1. Testing shows the presence of defects, not their absence Testing lowers the probability that defects remain. A green suite proves the tested paths work under the tested conditions. It never proves the software is correct. **Example:** "42 checks passed, covering checkout and three payment errors. Refunds were not tested." ## 2. Exhaustive testing is impossible Inputs, states, orderings and environments multiply. A form with five fields of ten values each has 100,000 combinations (10^5) before you consider timing, browsers or data. **Techniques:** equivalence partitioning, boundary value analysis, pairwise testing and decision tables. ## 3. Early testing saves time and money A defect caught in a requirement costs a conversation. The same defect caught in production costs a hotfix, a rollback and user trust. Testing activities should start as soon as there is something to review ("shift left"). **In practice:** testability reviews, acceptance criteria before code, and static analysis plus unit tests on every commit. Note: the often-quoted 1x → 100x cost multipliers come from older studies and vary widely. The direction is reliable; the exact ratios are not. ## 4. Defects cluster together A small number of modules usually hold most of the defects, often close to the Pareto split of roughly 80% of defects in 20% of modules. **Where to look:** modules with high complexity, recent churn, unclear ownership or a history of defects. ## 5. Tests wear out Running the same tests over and over finds fewer new defects each time. They still catch regressions, but they stop discovering anything. Older syllabi call this the "pesticide paradox". **Techniques:** property-based testing, exploratory sessions, and mutation testing to find tests that no longer catch anything. ## 6. Testing is context dependent A medical device, a banking API and a mobile game need different techniques, rigour and coverage. There is no single correct test strategy. **Inputs to the strategy:** risk, regulation, users and release cadence. ## 7. Absence-of-errors fallacy Software that passes every test can still fail if it solves the wrong problem or is hard to use. Verification asks "did we build it right?". Validation asks "did we build the right thing?". **Techniques:** stakeholder acceptance testing, usability checks, beta feedback and product analytics. ## Rules for agents - Never claim software is bug-free. Report what was tested, how, and what was out of scope. - Rank areas by risk before writing tests. Use partitions and boundaries instead of enumerating inputs. - Write or update tests in the same change as the code they cover. - When you find a defect, search the same module and similar code for siblings. - Refresh stale suites: vary inputs, add tests for new risks, remove redundant tests. - Match rigour to context, and state the context you assumed. - Confirm the feature meets the user's actual need, not only its spec. Source: https://durableqa.xyz/durable-testing/testing-principles --- # SDLC & STLC The **Software Development Life Cycle (SDLC)** describes how software is planned, built, released and maintained. The **Software Testing Life Cycle (STLC)** describes how it is tested. STLC runs inside SDLC, and it starts when requirements exist, not when code is finished. | Aspect | SDLC | STLC | |---|---|---| | Focus | Building the product | Verifying and validating the product | | Owner | The whole team: product, design, engineering | QA and test engineers, together with developers | | Starts | With a business need or idea | As soon as requirements exist | | Output | Working software | Evidence of quality: results, defects, a closure report | ## SDLC: the six phases ```text Requirements → Design → Implementation → Testing → Deployment → Maintenance ▲ │ └────────────────────────── next release ───────────────────────┘ ``` | # | Phase | Goal | Key outputs | Quality activities | |---|---|---|---|---| | 1 | Requirements | Decide what to build and why | User stories, acceptance criteria | Review for ambiguity, testability and missing edge cases | | 2 | Design | Decide how to build it | Architecture, data models, API contracts | Design reviews, threat modelling, testable interfaces | | 3 | Implementation | Build it | Source code, unit tests | Code review, static analysis, unit tests, TDD | | 4 | Testing | Verify and validate it | Test results, defect reports | Integration, system, regression and acceptance testing | | 5 | Deployment | Release it to users | Release build, release notes | Smoke tests, canary checks, a tested rollback plan | | 6 | Maintenance | Operate and improve it | Patches, enhancements | Monitoring, incident analysis, regression tests for fixes | Every SDLC model uses these phases. Waterfall runs them once, in order. Agile runs the whole loop every sprint. The V-model pairs each build phase with a test level. DevOps compresses the loop with CI/CD so it can run many times a day. ## STLC: the six phases ```text 1. Requirement analysis 2. Test planning 3. Test case development ─┐ often in parallel 4. Environment setup ─┘ 5. Test execution ⇄ defect → fix → retest 6. Test cycle closure ``` **Entry criteria** say when a phase may start. **Exit criteria** say when it is done. Write both down and check them before moving on. | # | Phase | Entry criteria | Activities | Deliverables | |---|---|---|---|---| | 1 | Requirement analysis | Requirements or stories are available | Identify testable requirements, raise questions, define scope, assess automation | Requirements traceability matrix (RTM), clarified questions | | 2 | Test planning | Requirements analysed | Choose strategy, scope, tools, roles, schedule; assess risks; estimate effort | Test plan, effort estimate | | 3 | Test case development | Test plan approved | Write test cases and automation scripts, prepare test data, peer review | Reviewed test cases, test data, scripts | | 4 | Environment setup | Architecture known; often parallel to phase 3 | Provision environments, seed data and accounts, smoke-test the environment | Ready environment, smoke-test results | | 5 | Test execution | Test cases, environment and a testable build are ready | Run tests, log defects, retest fixes, run regression | Execution report, defect reports, updated RTM | | 6 | Test cycle closure | Execution done or exit criteria met | Evaluate coverage against exit criteria, collect metrics, hold a retrospective | Test closure report, lessons learned | ## How they fit together: the V-model The V-model makes the link between the two lifecycles explicit. Each build phase on the left produces the basis for a test level on the right. Tests for a level are designed when its partner phase finishes, long before they run. ```text Requirements .......................... Acceptance testing System design ..................... System testing Architecture design ........... Integration testing Module design ............. Unit testing Coding (left arm: build / verification ↓) (right arm: test / validation ↑) ``` - **During requirements**, analyse testability and draft acceptance tests. - **During design**, plan system and integration tests against the contracts. - **During implementation**, write unit tests with the code, and prepare environments and data. - **During testing**, execute, log defects and retest. **At release**, smoke test and close the cycle. ## Rules for agents - Identify the current SDLC phase before choosing a testing activity. - For every requirement or story, write acceptance criteria and test conditions before writing code. - Do not start execution until entry criteria are met: a testable build, a ready environment, defined test cases. - Link every test to a requirement so coverage gaps are visible. - Log defects with steps to reproduce, expected and actual results, environment, severity and priority. - Close every cycle with a report: what ran, what passed, open defects and remaining risks. Source: https://durableqa.xyz/durable-testing/sdlc-stlc --- # The testing triangle (test pyramid) Popularised by Mike Cohn in *Succeeding with Agile* (2009). Write many small, fast, isolated tests at the bottom and a few broad, slow, realistic tests at the top. The higher a test sits, the more it costs to write, run and debug. ```text /\ / \ UI / E2E few · slow · costly /----\ / \ Integration some /--------\ / \ Unit many · fast · cheap /------------\ Anti-pattern: the ice-cream cone ( manual + E2E ) ← most tests \ integration/ \ unit / \ / \ / \____/ ``` The shares (often quoted as ~70% unit, ~20% integration, ~10% E2E) are a rule of thumb, not a target. What matters is the direction: push each check down to the lowest level that can catch the defect. ## The three layers | Layer | Scope | Speed | Catches | Example tools | |---|---|---|---|---| | Unit | One function or class; dependencies faked | Milliseconds | Logic errors, edge cases, regressions in isolated code | Vitest, Jest, pytest, JUnit | | Integration | Several components, or code plus a real dependency (database, HTTP, queue) | Seconds | Contract mismatches, wiring, serialisation and query bugs | Testcontainers, Supertest, Pact | | UI / E2E | The whole system through the UI or public API, as a user would | Seconds to minutes | Broken user journeys, configuration and deployment problems | Playwright, Cypress, Selenium | ## Why the shape works - **Feedback speed.** Thousands of unit tests run in seconds, so developers run them on every save. - **Cost.** Lower tests are cheaper to write and survive refactors of unrelated code. - **Reliability.** Fewer moving parts mean fewer flaky failures from timing, network or test data. - **Diagnosis.** A failing unit test points at one function. A failing E2E test points at the whole stack. ## Anti-patterns - **The ice-cream cone** inverts the pyramid: most effort goes into manual and E2E tests, with few unit tests underneath. The pipeline is slow, failures are flaky and hard to locate, and teams start ignoring red builds. - **The hourglass** has plenty of unit and E2E tests but almost no integration tests, so the seams between components, where many real bugs live, go untested. - **Duplicate coverage** asserts the same rule at every layer. It triples the maintenance cost without adding confidence. ## Variants - **Testing trophy** (Kent C. Dodds) adds static analysis as the base and makes integration tests the largest layer. It suits front-end apps, where most bugs appear when components work together. - **Test honeycomb** (Spotify) centres microservice testing on integration tests of each service through its real interfaces, with few implementation-detail tests. The right shape follows your architecture. The underlying rule does not change: prefer the fastest, most isolated test that can still catch the defect. ## Rules for agents - Test at the lowest layer that can catch the defect. - Give new logic unit tests, new boundaries (database, HTTP, queue) integration tests, and only critical user journeys E2E tests. - Keep unit tests fast and deterministic: no network, no real clock, no shared state. - When an E2E test fails, reproduce the cause with a lower-level test before fixing it. - Quarantine and fix flaky tests. Never retry them into green silently. - Do not assert the same behaviour at multiple layers. Source: https://durableqa.xyz/durable-testing/test-pyramid