Soft rules for testing software

A collection of rules I find useful for creating reliable tests. These are all guidelines more than hard rules, but following them closely has provided me a good baseline to work from when designing tests.

Table of contents

  1. Black-Box testing
  2. Mock testing
  3. Seams between worlds
  4. Interfaces
  5. Being confident with tests
  6. Shift-Left testing
  7. Prefer debugability over testability
  8. Testing more than code

Black-Box testing

Arguably the most important rule is to always black box test, never call a private function or read a private field in a test. If you find yourself needing to verify "invisible" properties of a data structure such as memory usage or internal invariants it might be worth using dependency injection and having public validate methods. Additionally if you find yourself being unable to exercise all paths of a private method in testing that may be a sign of dead code or overly complex logic. If you are certain that all paths are reached in real world usage of the method, try recording inputs to the method in the real world and then use those in your tests. If all else fails extracting the method to a package private class or detail namespace and testing there is preferable to making the function public.

Mock testing

Continuing on from black-box testing is mock testing, mock testing is the antithesis of black-box testing where you assert that methods are invoked with specific parameters in specific orders. I argue that mock tests are always more harmful than any other form of testing, if you imagine your functions and classes as blocks of lego then mock testing is superglue. Any code change requires changing a dozen mock tests and more importantly you have no clue if the code change works. Mock tests verify nothing other than that the code is exactly what you wrote when you wrote the test. Tests are supposed to provide a level of confidence that a change has been made without changing existing behaviour in an undesirable way, mock tests allow no code changes to happen and therefore provide no confidence as the tests immediately fail on any change.

Virtual dependencies over mocks

My favoured solution to avoid mocks is to implement virtual dependencies that behave as if they were the real service, but without making actual api calls. Ideally any test could have the virtual service swapped out with the real service client implementation and would still pass. This means that all validation should be done using the services public api, although this may not be possible when testing cases that require fault injection. Here is a not too contrived example to demonstrate the difference between a virtual dependency and a mock.
                
@Component
public interface GlacierClient {
    Vault describeVault(String name);

    Vault createVault(String name);
}
                
                
@Component
@AllArgsConstructor
public class LogArchiver {
    public static final String LOG_ARCHIVE_VAULT = "LogArchiveVault";

    @Autowired
    private final GlacierClient glacier;

    public void createLongTermLogStorage() {
        Vault vault = glacier.describeVault(LOG_ARCHIVE_VAULT);
        if (valut == null) {
            vault = glacier.createVault(LOG_ARCHIVE_VAULT);
        }
    }
}
                
                
public class LogArchiverTest {
    @Mock
    private GlacierClient glacierClient;

    @Autowired
    private LogArchiver logArchiver;

    @Test
    public void testCreateLongTermStorageCreatesVaultIfNeeded() {
        // If the vault doesnt exist
        expect(glacierClient.describeVault(LogArchiver.LOG_ARCHIVE_VAULT)).andReturn(null);
        // then it should be created
        expect(glacierClient.createVault(LogArchiver.LOG_ARCHIVE_VAULT)).andReturn(new Vault(LogArchiver.LOG_ARCHIVE_VAULT));

        logArchiver.createLongTermLogStorage();
        replay(glacierClient);
    }

    @Test
    public void testCreateLongTermStorageDoesNotCreateVaultTwice() {
        // If the vault doesnt exist then the vault should not be created again
        expect(glacierClient.describeVault(LogArchiver.LOG_ARCHIVE_VAULT)).andReturn(new Vault(LogArchiver.LOG_ARCHIVE_VAULT));

        logArchiver.createLongTermLogStorage();
        replay(glacierClient);
    }
}
                
            
This demonstrates very basic mocking to ensure that a function calls specific apis, it also demonstrates that the test is tightly coupled to not only the functions interface, but its implementation and the interface of its consumed dependencies. This makes changing code hard and will require rewriting all tests when changing the behaviour of this function, which ensures that new manual testing will be required as the test code can no longer be trusted since its changed. In this way mock testing is good at preventing regressions as making any change becomes such a chore that people choose to not change code if they can at all help it. Additionally mock tests necessitate duplicate tests when you come to write your full end-to-end tests, as mocks peer inside dependencies to see what happened it is impossible to directly migrate them into a system test where you dont have access to the internals of your dependencies. Using virtual dependencies can aleviate all of these issues, and also comes with its own benefits. Assuming the same implementation here is an example of virtual dependencies.
                
@Component
public class VirtualGlacierClient implements GlacierClient {
    private Connection jdbcConnection;

    private static class VaultRecord {
        public String name;
    }

    public VirtualGlacierClient() {
        this.jdbcConnection = DriverManager.getConnection("jdbc:sqlite:glacier.db");
        try (Statement statement = connection.createStatement()) {
            statement.execute("CREATE TABLE IF NOT EXISTS vaults (id INTEGER PRIMARY KEY, name TEXT, CONSTRAINT ak_name UNIQUE (name))");
            statement.execute("DELETE FROM vaults;");
        }
    }

    @Override
    public Vault describeVault(String name) {
        try (PreparedStatement statement = connection.prepareStatement("SELECT * FROM vaults WHERE name = ?")) {
            statement.setString(1, name);
            try (ResultSet results = statement.executeQuery()) {
                if (results.next()) {
                    String name = results.getString("name");
                    return new Vault(name);
                }
            }
        }

        return null;
    }

    @Override
    public Vault createVault(String name) {
        try (PreparedStatement statement = connection.prepareStatement("INSERT INTO vaults (name) VALUES (?)")) {
            statement.setString(1, name);

            // Throws a unique key constraint if a duplicate vault is created
            statement.executeUpdate();

            return new Vault(name);
        }
    }
}
                
                
public class LogArchiverTest {
    @Autowired
    private VirtualGlacierClient glacierClient;

    @Autowired
    private LogArchiver logArchiver;

    @Test
    public void testCreateLongTermStorageCreatesVaultIfNeeded() {
        // The vault should not exist beforehand
        Assertions.assertNull(glacierClient.describeVault(LogArchiver.LOG_ARCHIVE_VAULT));

        logArchiver.createLongTermLogStorage();

        // The vault should be created
        Assertions.assertNotNull(glacierClient.describeVault(LogArchiver.LOG_ARCHIVE_VAULT));
    }

    @Test
    public void testCreateLongTermStorageDoesNotCreateVaultTwice() {
        // Create the vault beforehand
        glacierClient.createVault(LogArchiver.LOG_ARCHIVE_VAULT);

        logArchiver.createLongTermLogStorage();

        Assertions.assertNotNull(glacierClient.describeVault(LogArchiver.LOG_ARCHIVE_VAULT));
    }
}
                
            
While this code is obviously longer than the version with mocks I argue it comes with many benefits.

Seams between worlds

In the context of testability people often write an interface with the express intent of using it as a test seam, I would argue this is always an antipattern. Rather than using an interface as a test seam consider only using seams between your code and the outside world as points of injection for tests. Afterall the thing you care about is not that a specific method was invoked with a specific parameter, you care that a database row was added or that an email was sent. Testing only observable behaviour is key to writing reusable and easily changeable code, and is a fundemental part of why many systems work at all. Databases only query as fast as they do due to doing giant amounts of hidden work and optimization, similarly compilers only produce such fast assembly because they can generate code that behaves "as if" it was what you wrote. This disconnect between observable behaviour and implementation is key to writing maintainable code and being confident in your changes. If compilers were not allowed to optimize a loop because perhaps someone would disassemble the loop and see that its vectorized rather than copying individual ints they would be very slow, the same applies to all code. Using interfaces for testing breaks this principal.

Interfaces

Always use the least powerful abstraction: Never write an interface when a class would do, never write a class if a function would do.

Being confident with tests

If you need to change a test when fixing a bug, you have lost confidence in the test and it has become a hinderance. Any change to a test necessitates manual testing to verify the new test is correct, which can be time consuming and costly in many cases.

Prefer debugability over testability

Easily debuggable code and easily testable code are sometimes at odds with each other, as 80% of codes lifetime is spent being maintained and often under high stress its best to write code that easy to debug rather than code that is easy to test. Always write code as if the next person who needs to read it is reading it after 14 hours of oncall while handling an outage, and they're a very unreasonable person who has your home address. To reiterate what was said in the Interfaces section every level of indirection adds mental overhead and context switching, the cost of jumping to the definition of a function is less than the cost of jumping to a definition thats an interface and then having to find all implementations of that interface and determine which is the currently instantiated one.

Testing more than code

Traditionally testing is reserved for the behaviour of the code, rather than other properties like structure or style. But with so many languages having support for reflection, or good tooling for writing custom lints it is a better time than ever to explore testing the architecture of your software. Libraries like archunit and tools such as clang-tidy being readily available make it easier than ever to enforce codebase specific standards. Consider using reachability analysis to ensure no class in one package ever calls into another, prevent usage of banned library functions, keep all your rest controllers in one package and prevent them from spreading around. Preventing entropy like this can help with long term maintanence and keep a swarm of engineers from turning a codebase into soup. Additionally consider using a languages own facilities such as javas project jigsaw or c++ modules to ensure nobody can depend on your implementation details. Project jigsaw also allows you to use jlink and jpackage, which can produce custom jdk images tailored to your application and strip out unused modules and classes. This can shrink deployment sizes drastically, I personally have seen reductions from 100mb uberjars to 15mb after modularizing a codebase and its dependencies.