Sunday, May 3, 2015

... about convincing people NOT to use event sourcing



I work at a company where they have embraced the idea of using CQRS with event sourcing (ES). It's seen as such a great pattern that they decided to use it for everything. It has become their favorite solution to every problem and axon framework has become the golden hammer in their toolbox to whack every nail.

That's weird because for years, we've been trying to convince management what wonderful things we could do with CQRS+ES, if they only would let us... Now that they're finally convinced, it feels strange to do the opposite. To tell them what a terrible idea this is when used as the solution for the wrong problem.

When using CQRS+ES, your events become the contract. Spatial and temporal. Your events define how other bounded contexts communicate with you and it's how your own future bounded context will communicate with you. So it's super important to get those boundaries right and to get those events right. Sure, you can upcast an old obsolete event to one or more newer events and axon has support for that, but that can become a real pain when you need to do too many of those.

Other things that are super easy to do using a non ES solution, like correcting bad data, becomes hard using ES. In a non ES solution, customer support just executes some SQL update statements in your production database and the customer is happy. In an ES solution customer support needs to publish a corrective event in order to let the view do the correct projection.

So ES can become quite complex. But when I tell people ES is complex, they reply: "No it's not complex. With axon framework it's easy. You just need to define the commands, aggregates, events and projections. Axon takes care of everything else." All the infrastructure, all the nitty gritty details of publishing, storing, replaying events is handled by this wonderful framework. However, they're only talking about development complexity and fail to see the operational complexity. It might be easy to program your CRUD-like application but once it's in production, things might become a mess.

I'm not saying CQRS+ES doesn't solve some very hard problems in a very elegant way.
  • For instance when facing a domain where the time factor is of great importance (what was the state of an aggregate 5 weeks ago).
  • Or where you will need to do a lot of projections in the future that you don't know now. 
  • Or where you have a high collaborative domain where your read model cannot keep up with the concurrent writes and you need to separate them in order to scale.
Traceability is not in that list! You don't NEED event sourcing to do traceability. You can do the same, simply by using a tracelog. Just write who did what in a separate table and you're done. 
Furthermore, when your aggregates are only created and never change afterwards, ES is probably not a good idea. Just store the state of the aggregate in the database and project your views based on that.
Finally, when all you have is some CR(U)(D), don't bother doing event sourcing. Your events will only be called CreateXXXEvent, UpdateXXXEvent or DeleteXXXEvent. and won't contain any relevant business information.

Greetings
Jan

Monday, February 23, 2015

... about ugly UUID's in my tests



In my current project we use UUID's as id's for all aggregates. The good part of UUID's is that it's guaranteed to be unique. The ugly part is that it looks ... well ... ugly.

Writing tests with lots of UUID's clutters the test codebase and you lose focus on what it is you're testing.

A common use case when writing tests is that I want to assert that some id is the same as some expected id. For instance I want to test a filter - named VeryComplexFilter - that keeps users older than 40 and sorts them by age. The code is written in groovy:

Loading code....

I know. It's not the most compelling code, but lets continue for the sake of the example. My unit test creates a list of users and asserts the outcome of my filter. I'm using a data-driven test using spock because I want to test different inputs in the same test. Spock does a great job here.

Loading code....

In the where-block, I create the different input lists and define the expected  result. Each input list contains pairs of UUID-and-age tuples. In the when-block, each element of the input list is converted into a User object and handed to the filter under test. Finally in the then-block, the id's of the result list is compared against the expected results. When I run the test I get the following output:


You can see that the number of UUID's I have to create, really makes my test ugly. I could pregenerate them and then it would look like this:

Loading code....

It's the same test but I pregenerate the UUID's in a static list and refer to them from within the where-clause. However, what I really want, is to use some symbolic notation like a string that refers to a UUID. The test than would look much simpler:

Loading code....

In this where-block the input list contains a lists of strings like 'a-35' that I can refer to in the expectedResults. Now I can specify, for instance, that - provided with a list of users with ages 35, 52, 15, 79 and 41 - the filter must return the 2nd, the 5th and the 4th user. I do this by setting ...

  • the input list to "a-35", "c-52", "d-15", "e-79", "b-41" and 
  • the expectedResult to 'bce'. 
Every time I refer to uuid 'a' or 'b' or 'c', I want the system to remember what UUID I'm talking about. I can easily do that with the help of groovy's memoization:

Loading code....

I just call a closure named uuid that returns me the same UUID, every time I call it with the same input. Executing uuid.call('a') twice will always give me the same UUID. Executing uuid.call('b') gives me a new UUID.

What if I want to return something else instead of UUID's? I can do the same creating a user:

Loading code....

I can generalize this idea of creating something and return the same complex object, every time I call it with the same input parameter. Here's a method creating a complex object:

Loading code....

Here's how I call it and expect to get the same object for the same input value:

Loading code....

For this to work I need to be able to say: "This is how you create my complex object":

Loading code....

The above snippet registers a way to create the object and also passes the key that I can use to specify what I want. Once that's done, I can call it, like I did using the uuid closure call.

Loading code....

You can find a full gist here.

I don't know how far I'll take this trick in my tests. The fact that you can return the same - potentially complex - object and use some short symbol to reference it, helps a lot when trying to create some clear specification. However, there is some hocus-pocus going on, so it's not something I would apply everywhere... I think.

Greetings
Jan

Thursday, February 5, 2015

... about type safe closures in groovy

I've been using groovy for a couple of years now. I love it. It saved me from several nervous breakdowns coding in the pre-jdk8 world.

However, there's one thing that starts to annoy me.


It's closures. 
At least when passed as a parameter in a method, mimicking higher order functions.

In the following example, there's a CustomerNotifier that sends an email to a customer. It accepts a CustomerId and a closure creating the email. The closure is called with the Customer and a string indicating the environment ( we don't want to send emails to the real customers in a dev environment, do we):

Loading code....

In the method signature, there's no way to express what input parameters the closure needs. I can express the result, but not the input parameters.

That's annoying because now, the caller of that method, must look into the implementation to know what arguments will be passed to the closure. To make things worse, your IDE will not be able to assist you because it too won't have a clue of the parameters it expects. It's like typing blindfolded.

I could use @ClosureParams. This annotation provides extra information about the parameters a closure needs, assisting the IDE. However it just makes your code really ugly:

Loading code....

Lately I started using a library called functional java. It facilitates programming functionally in java and can be used by all poor souls that aren't allowed to program in scala but do want to use a more functional approach.

Among the basic stuff it contains interfaces for functions with 1 up till 8 input parameters. With these I can clearly define what input parameters my function needs.


Loading code....

Together with groovy's implicit closure coercion, I can use closures as typed functions while my IDE will resolve the input parameters for me and warn me when using the wrong type.

Loading code....

I'll definitely be using the functionalJava library more as it works great with groovy. There's also a groovy library building on top of functionalJava at https://github.com/mperry/functionalgroovy. It makes working with groovy even smoother but up till now I haven't really felt the need to use it.

You can find a gist of the complete example here.

Greetings
Jan

Friday, November 28, 2014

... about layers and maven modules. Lots of them.

tl;dr: stop scattering your codebase into maven modules divided by layers. 






I wish I could go back in time. Back to the time where I was still a junior in software development. The way things worked was to look at the existing code and copy it. When someone used an abstract factory I would try to do the same. When some of my colleagues used a visitor pattern to do double dispatch, I would use them as often as I could. I blindly accepted the patterns that were used, made them my own and tried to apply them everywhere.

That's why I was called ... a junior developer.



When I started working as a developer, aspects were the next big thing that would change everything. Hibernate did some amazing things compared to ejb2. Spring blew our minds with inversion of control. JSF was a better Struts and Seam was a better JSF. There were all these crazy new technologies that changed the way we would build software forever.



And then there was Maven. Ant was old school. Maven was hot. It gave us a way to have dependencies at runtime that were not available at compiletime allowing us to do some funky stuff with maven modules. 



It was a DIP dream come true. It felt good and it felt intelligent. We were able to protect ourselves (and those juniors) against violations made against the dependency inversion principle. 


Somehow, you could create maven modules containing services using other services that were implemented in another maven module. But those modules didn't depend on each other. Instead they depended on yet another module containing only interfaces. 

So we created a maven module per layer:

  • one containing web controllers for the UI layer, 
  • one containing application services, 
  • one containing the domain (which we called 'model' because we didn't understand the difference), 
  • one containing the infrastructure

... and finally  
  • one to tie all modules together in a war
Each layer had its own responsibility. The controllers would only accept requests and turn them into calls to services in the application layer. The application service would call domain objects from the repositories. The interface would be in the domain layer while the implementation would reside in the infrastructure. The UI layer would not have access to the domain layer and the domain layer would not have access to any other layer. Nice. Clean. Separated. 



However. It didn't stay with those 4+1 modules. 

  • We had dto's that would transfer state from the domain layer all the way to the UI. So we created a module containing dto's.
  • We had converters that would convert the aggregate to/from the dto. These obviously needed to stay in their own maven module.
  • We had value objects that were accessible in the web layer. You can't put them in the domain layer because that's not accessible from the UI layer. The solution? ... add another maven module.
  • We had code that didn't belong anywhere but that was used everywhere. The infamous shared kernel maven module was born.
  • There was code that would handle exceptions. 
  • Code that would handle web requests. 
  • Code for handling queries.
  • Rule engines.
  • Email generators.
  • Logging code.
  • Hibernate user types.
  • Aspects. 
  • ...


Memory is failing me, but what I can remember from those days is we built a lot of maven modules just because we could. When we had a piece of code that didn't really belong anywhere, we would create a new maven module. Most modules only had 1 or 2 packages each with a small number of classes. Some didn't even contain code but only contained poms tying dependencies together to manage the dependency complexity.

The result was that every project had at least 60+ modules while the bigger projects would contain more than 140 modules. 


Now, there's 2 problems with this:

1. when writing a class, it becomes very difficult to know where it belongs to. Is a validation service part of the application layer or is it a domain service. In what package should we put an event listener. The result is we put it in the shared kernel module that grows and grows and contains all sorts of things.


2. sometimes we do want the domain layer to access a dto or a command directly. Sometimes an application service needs access to some infrastructure code. Sometimes we want to query something directly in the UI layer. Sometimes the strict dependency rules we set up don't allow us to do what we want. Instead of relaxing those rules, we put those classes in yet other maven modules that do have the needed dependencies to the dto's, infrastructure, ...

Here's a law: Enforcing dependency rules by separating your application into maven modules always leads to more and more modules just to satisfy those strict rules. 


What's even more terrible is that in order to do something simple, like storing a key-value in a database, you need to go through all these layers. You need to start with a post arriving in the UI that translates the key-value into a dto that sends it to an application service that creates a some aggregate and tells a repository to save it. You can't just call some sql in the UI contollers because you just don't have a dependency on a dataSource in your web layer.


What you end up with, is an application that is totally scattered. Your application contains components or bounded contexts, but they're completely spread out over those 140+ modules. You cannot see what pieces belong together to form a component. What started as a noble intention to have a clean separation ends in a big ball of modules.


Over the last 8 years I grew older and hopefully a little wiser. I don't like to separate my code into different maven modules. Especially not by layers. I like to keep all the code together as much as I can.

When I do separate my code it's done by functionality, by component, by bounded context. Inside a maven module you would have code that belongs in the UI layer, the application layer, the domain layer and the infrastructure layer. I still have those layers, but it's indicated by putting them in different packages. 

Oooohh, but can't some junior developer do something stupid like directly use an implementation of a repository instead of the interface?  The answer is yes. Developers can do stupid things. But they can do stupid things in the 140+ module project too. Chances are they will do far more stupid things in a complicated project than in one that is separated by functionality.

Ok.

Here's the problem. 

Those projects that I worked on when I was a junior, are still around. In the meantime a new generation of developers has arrived and they're working on that same codebase, taking the same patterns that were used back then as a reference of how things should be done, making the same mistakes and assuming having a lot of maven modules is the right thing to do when you're doing some serious programming.


Here's another problem:

I don't seem to be able to sell the idea - maven modules per component instead of per layer - to most of my colleagues. They seem so indoctrinated by separating software into maven modules by layers that they don't see the wreckage it does to the clarity of the program. The idea that we need to provide these strict rules in order to obey the dependency inversion principle is so strong that they don't understand that it's detrimental to the maintainability of the code. 

Truth be told, because of this doctrine, I never even worked an a software project that was not separated by layers. I was just never able to convince my colleagues to do otherwise. But I really hope to do so. Very soon. Because it's the best protection against the ever increasing entropy in our codebase.

Greetings
Jan

Sunday, September 22, 2013

... about what ubiquitous language we should speak

DDD talks about the use of a ubiquitous language; a common language spoken by the domain experts, functional analysts and developers. As a consequence, the code that is developed should also be written in the same language, using the nouns and verbs existing in the ubiquitous language.

The code is the language. The language is the code. Hallelujah!


I live in Belgium.


Belgium is small. When you choose the right time of the day, you can drive all the way through Belgium in 2 or 3 hours, even without breaking any speed limits. 

The upper part speaks Dutch, the lower French. My native tongue is Dutch. My English skills are OK, but could be better. In France, I could order a meal and buy a croissant, but I wouldn't say my French is good.

The domain experts we work with speak Dutch, French or English. The clients we work for, are local companies or international ones.

All our code is written in English. It's the universal (!= ubiquitous) language. However when talking to our domain experts we speak the language they speak. That is Dutch, French or English. So every time we talk to the business, we take the verbs and nouns that is used in our code and translate it into the language they speak. 

That's not very ubiquitous, is it?

So, lets start an evolution. When we're dealing with a local company that uses Dutch as the language, we use dutch verbs and nouns and they end up in our code. 

We would have an entity like Vliegmachien that is persisted in a VliegmachienRepository.

The Vliegmachien entity would contain method names like stijgOp or maakNoodlanding and all would be fine.

Note that there's no reason to use Dutch terms for technical stuff like Service or Repository . That wouldn't serve any purpose. They can remain in English

When we're dealing with an international company, we would use English as the ubiquitous language and we would create an entity called Airplane containing methods like liftOff or makeEmergencyLanding.

Case closed! Everybody happy.

OK, here comes the trouble: 

Sometimes we have a Dutch speaking domain expert that leaves the company and is replaced by a French speaking domain expert.

Or this one:
the project is now in maintenance mode and it's not handled by the original Dutch speaking team anymore. Maintenance is now done by someone living in Ukraine.

For these reasons I see no other possibility than to use English as the ubiquitous language. Unless you can foresee that your customer and your team will always be speaking the same language.

Talking to domain experts in a language that is not yours and is not theirs is not that hard, but you tend to miss a lot of the small details and subtleties hidden in the language. You could"google-translate" to-from Dutch-English, but most of the time you end up with verbs like "assignSomthing", "updateSomething" or the dreaded "setSomething". It's poor, plain English, spoken by all those foreigners that aren't living in the US or aren't part of the British Commonwealth. However, most of the time it's the best we can do.

To make a short story long: English speaking developers have a head start on others when capturing the ubiquitous language.

Greetings...
Jan