You don't need an effect system
How we got here
A couple of months ago something rather unusual happened: I got permission to rewrite an old production codebase. Nothing wild inside, just a few services bouncing messages around and calling other places. I wrote most of it years ago, and most of it was god awful on account of the fact that I had no idea what I was doing at the time.
Like any real codebase written in Haskell, it relied on an effect system to do… well, things.
Indeed, like most of the community, I couldn't properly formulate what an effect system does. Yes, I know there are at least ten of them, all at odds with one another, yet seemingly completely interchangeable. Smarter people have narrowed the goals down to tracking effects, mocking and internal consistency, which strongly implies that plain IO is incapable of these things, and that's something I could neither confirm nor deny.
So now, being able to reassemble an entire system from the ground up, the question I got to ask was…
Am I using the effect system for anything?
- Do I need it to pass arguments around?
No, I can do that manually.
- Does it make for a safer codebase?
No, tracking effects does not preclude anyone from adding bad IO to an effect implementation, from importing unsafePerformIO, or from finding a sum of an infinite list.1
- Am I using it outside of IO?
No, for pure functions that do mutation the ST monad works just fine.
**Is there any benefit to tracking IO as a
separate effect?**
No, and in fact the opposite: there are small effects all around the codebase that I'd rather not track.
For example, random number generation is commonplace, but the steps necessary to generate a random value are different every time. The only shared behavior is the use of generator state, and even then it's unclear if passing it around has any benefits over using the global generator.
- Have I made use of higher-order effects?
No, none of the effects I'm using need to overlap or nest. I'm not precisely doing rocket science over here; I'd prefer that whatever I'm using doesn't come with a whole separate manual.
After throwing out pretty much every single feature, I finally found the one thing I was using the effect system for: error handling. Which upon closer inspection turned out to be the execution of some set of other effects followed by…
Early return
This may well be the most embarrassing problem in all of Haskell. Over and over again people waltz in with the exact same basic question: "How do I return from a function early?".
And time and time again they're hit with the same three options:
Make it so that the error is the last statement in
the function, using Either or continuations.
Code looks awful and constantly drifts to the right.
- Throw an exception.
*Roughly equivalent to burying a landmine in your backyard.*
Use an effect system.
???
The catch is that effect systems don't have some special third way of handling errors, they merely wrap the other two approaches. Notably, ExceptT is a faithful implementation of the first approach, threading an Either through every action. That's obviously very inefficient and is known to not compose well, so I'd prefer the second option.
Type-safe exceptions
To do this we'll need some way to carry the knowledge that a specific resource (in our case an exception) may only be used (thrown) within a specific function. This, unsurprisingly, is a problem that has been solved decades ago for file handles and raw pointers through the use of bracket. Though in our case there's nothing to allocate, the value is on the type level:
{-# LANGUAGE RoleAnnotations #-}
module Early
( Early
, leave
, runEarly
) where
import Control.Exception
import Data.Typeable
type role Early nominal
data Early a = Early
data ReturningEarly a = ReturningEarly a
instance Typeable e => Show (ReturningEarly e) where
show = displayException
instance Typeable e => Exception (ReturningEarly e) where
displayException (ReturningEarly _) = "Early return exception"
leave :: Typeable e => Early e -> e -> IO a
leave _Early e = throwIO $ ReturningEarly e
runEarly :: Typeable e => (Early e -> IO a) -> IO (Either e a)
runEarly f = catch (Right <gt; f Early) (\(ReturningEarly e) -> pure $ Left e)
And the module can then be used like this:
import Control.Monad
import Early
example :: Early () -> IO Bool
example early = do
putStrLn "Ran this action"
when True $ do
leave early ()
putStrLn "Didn't run this action"
pure True
ghci> runEarly example
Ran this action
Left ()
There are no caveats to this code beyond those that come with using bracket.
Intermission
And just like that I ran out of reasons to use an effect system.
There's not much of a story to tell from this point on: I shaped every other effect much like I had shaped Early and everything fell into place nicely. The rest of the post are my findings, structured to the best of my ability.
Effects in plain IO
Finding a definition for the word "effect" is unfortunately as tedious as finding one for "effect system". If I am to trust some people on Reddit, an "effect" is a convention to only access specific side effects through a common interface tracked on the type level.
Both Early and "handles" from the "handle pattern" fit this definition:
They have interfaces (leave,
createUser) and implementations (runEarly, withHandle). Compare to the similar effect/handler separation in eff.
- They're tracked on the type level, contrast
createUser :: Handle -> Text -> IO User
listDirectory :: FileSystem :> es => FilePath -> Eff es [FilePath]
leave :: Early e -> e -> IO a
throwError :: Error e :> es => e -> Eff es a
More generally, an effect in plain IO is a data type that serves as a bridge between interface and implementation functions. The data type's constructor is as such an implementation detail and should not be exported.
Effects are tracked as function arguments; this is quite different from effect systems, which prefer constraints. One benefit of this is that we don't have to deal with the complexities of disambiguating, reordering or reinterpreting effects.
Module structure
Effect systems tend to put all of their definitions into a single module, but there is no hard requirement for this. For example, Early can be broken into
module Effect.Early.Leave (Early, leave) where
module Effect.Early.Runner (Early, runEarly) where
module Effect.Early.Internal (Early, leave, runEarly) where
This can be used to track function access at module level; particularly useful if an effect has multiple interfaces (say, multiple programs rely on the same effect data type) and/or multiple implementations (say, mocking).
Error handling
Effects may throw exceptions, but if they do they're also responsible for catching them. Much like the data type constructors, exceptions are an implementation detail.
Effect implementations are generally not invoked alone however, they're stacked into a terrifyingly large pile at the very edge of the program. In our case stacking does work out of the box, but each layer pushes the successful case deeper into the chain:
runFoo :: (Foo -> IO a) -> IO (Either FooError a)
stack
:: (Foo -> Bar -> Qux -> IO a)
-> IO (Either FooError (Either BarError (Either QuxError a)))
stack f =
runFoo $ \foo ->
runBar $ \bar ->
runQux $ \qux ->
f foo bar qux
This can be solved rather nicely by using slightly more complicated types:3
runFoo :: (Foo -> IO (Either e a)) -> IO (Either (Either e FooError) a)
lift :: IO a -> IO (Either Void a)
lift = fmap Right
stack
:: (Foo -> Bar -> Qux -> IO a)
-> IO (Either (Either (Either (Either Void QuxError) BarError) FooError) a)
stack f =
runFoo $ \foo ->
runBar $ \bar ->
runQux $ \qux ->
lift $
f foo bar qux
Duplicate effects
Passing around two effects with the same name would be wildly confusing; passing something like a Tagged "helper" Database would be a massive nuisance. If a user needs two of the same effect, they should wrap the functions on their side into nice newtyped names (e.g. HelperDatabase).
Duplicate exception handling in implementation functions can be addressed in a similar way by providing an extra argument and matching on that, essentially letting the user call one effect "database 1" and the other "database 2".
Real-world effects
Let's look at all of the categories of effects I ended up with (or without) in my codebase.
Argument passing, fancier
Configuration data is generally read at application start before invoking the implementation stack, and is passed into it as either function arguments or "reader" effects.
Trying to implement a "reader" outside of an effect system results in a bunch of redundant wrapping:
data Conf =
Conf
{ bar :: Int
, baz :: Bool
, qux :: String
}
getBar :: Conf -> Int
getBar = bar
runConf :: Conf -> Conf
runConf = id
I thus prefer to keep configuration data types in separate modules and to pass them around as arguments.
Mocking
An effect that can be mocked is simply a product of functions:
data Database =
Database
{ createUser :: Text -> IO (Maybe User)
, getUserMail :: User -> IO [Mail]
}
runRealDatabase :: RealDatabase -> Database
runRealDatabase db =
Database
{ createUser = Database.createUser db
, getUserMail = Database.getUserMail db
}
Effects that never fail—like logging and metrics collection—are special cases of mocking:
import Data.ByteString.Builder
import Data.Text (Text)
import Data.Text.Encoding
import System.IO
newtype Logger = Logger (Builder -> IO ())
note :: Logger -> Text -> IO ()
note (Logger f) = f . encodeUtf8Builder
runLogger :: Handle -> Logger
runLogger handle =
Logger $ \msg ->
hPutBuilder handle $ msg <> "\n"
Tracking resources
I'll use the database effect as an example. Here's a rough outline of the implementation function:
import Control.Exception
import Database.PostgreSQL.Simple
import GHC.Stack
import System.IO
newtype Database = Database Connection
runDatabase
:: ConnectInfo
-> (Database -> IO (Either e a))
-> IO (Either (Either e DatabaseError) a)
runDatabase connInfo f =
mask $ \unmask -> do
conn <- connect connInfo
let cleanup = close conn
ei <- unmask (Right <gt; f (Database conn))
`catch` \ex ->
case fromException ex of
Just dbEx -> pure $ Left (dbEx :: DatabaseException)
Nothing -> _ -- (4) Failure above.
case ei of
Right (Right a) -> _ -- (1) Success.
Right (Left e) -> _ -- (2) Failure below.
Left ex -> _ -- (3) Failure at this level.
For the effect to behave the same regardless of its position within the implementation stack cases (1), (2) and (4) should all use the same cleanup function.
Interface functions may use the resource directly, although in libraries that use exceptions liberally it's more convenient to funnel all uses through a helper function:
data DatabaseException = DatabaseException CallStack DatabaseExceptionKind
data DatabaseExceptionKind = DatabaseSqlException SqlError
| _
handlesPostgres :: HasCallStack => Database -> (Connection -> IO a) -> IO a
handlesPostgres (Database conn) f =
let rethrow :: (e -> DatabaseExceptionKind) -> e -> IO a
rethrow kind = throwIO . DatabaseException callStack . kind
in f conn
`catches`
[ Handler $ rethrow DatabaseSqlException
-- , -- however many more of these are necessary --
]
createUser :: Database -> Text -> IO (Maybe User)
createUser db =
handlesPostgres db $ \conn ->
_ conn
Tying everything together
Interfaces
Taking argument passing to the extreme, one will inevitably end up with a function that looks like
endpoint
:: AuthConf -> ServiceConf
-> Logger -> Metrics -> Early ServerError -> Database -> Cache -> Messaging
-> AuthHeader -> EndpointRequest -> IO EndpointResponse
endpoint authConf serviceConf logger metrics early database cache messaging
authHeader EndpointRequest {..} = do
To combat this the "handle pattern" post proposes nesting effects; the "ReaderT design pattern" post instead suggests carrying all the effects in a record. Both of these solutions are wrong.
GHC has warnings for unused arguments, and since we treat effects as arguments, we can use it to point out unused effects. To leverage this, the code must be structured in such a way that each statement uses at most one effect. A bunch of statements then bundle into a function, functions bundle into larger functions, until we reach the aforementioned endpoint.
As an example, creating a user actually requires at least three effects:
newUser :: Database -> Text -> IO (Maybe User)
createUser :: Logger -> Early ServerError -> Database -> Text -> IO User
createUser logger early database qux = do
mayUser <- newUser database qux
case mayUser of
Just user -> pure user
Nothing -> do
note logger "Could not create a user entry"
leave early err500
This approach works in the opposite direction too: if we need to find out what caused a DatabaseError, we only need to track which functions are passed the Database argument.
Implementations
Running the implementation stack is as straightforward as in any effect system:
main :: IO ()
main = do
conf <- readConfiguration
let logger = runLogger stderr
metrics <- runMetrics (getMetricsConf conf)
withDatabasePool (getDatabaseConf conf) $ \dbPool -> do
withMessagingEnv (getMessagingConf conf) $ \msgEnv -> do
note logger "Initialization complete"
consumeMessagesForever msgEnv $ \newMsg ->
runStack logger metrics dbPool msgEnv $ \early database messaging ->
process logger metrics early database messaging newMsg
runStack
:: Logger -> Metrics -> DatabasePool -> MessagingEnv
-> (Early () -> Database -> Messaging -> IO ())
-> IO ()
runStack logger metrics dbPool msgEnv f = do
ei <- runEarly $ \early ->
runDatabase logger metrics dbPool $ \database ->
runMessaging logger metrics msgEnv $ \msg ->
lift $
f early database msg
case ei of
Right () -> _ -- Success
Left err ->
case err of
Left (Left (Right msgError)) -> _ -- Messaging error
Left (Right dbError) -> _ -- Database error
Right () -> _ -- Early return
process
:: Logger -> Metrics -> Early () -> Database -> Messaging
-> NewMessage -> IO ()
The shape remains almost the same when using servant. The only quirk is that argument passing renders all the fancy hoisting functionality completely unusable, so the stack has to be invoked inside each endpoint manually.
runStack
:: Logger -> Metrics -> DatabasePool -> MessagingEnv
-> (Early ServerError -> Database -> Messaging -> IO a)
-> Handler a
runStack logger metrics dbPool msgEnv f =
MkHandler $ do
ei <- runEarly $ \early ->
runDatabase logger metrics dbPool $ \database ->
runMessaging logger metrics msgEnv $ \msg ->
lift $
f early database msg
case ei of
Right a -> pure $ Right a
Left err ->
case err of
Left (Left (Right msgError)) -> _ -- Messaging error
Left (Right dbError) -> _ -- Database error
Right srvError -> pure $ Left srvError
type API = CountAPI :<|> _
server
:: Logger -> Metrics -> DatabasePool -> MessagingEnv
-> Server API
server logger metrics dbPool msgEnv =
countServer logger metrics dbPool msgEnv
:<|> _
type CountAPI = "count"
:> ReqBody '[JSON] CountRequest
:> Get '[JSON] CountResponse
countServer
:: Logger -> Metrics -> Early ServerError -> DatabasePool -> MessagingEnv
-> Server CountAPI
countServer logger metrics early dbPool msgEnv request =
runStack logger metrics dbPool msgEnv $ \early database messaging ->
countEndpoint logger metrics early database messaging request
countEndpoint
:: Logger -> Metrics -> Early ServerError -> Database -> Messaging
-> CountRequest -> IO CountResponse
Conclusion
There is only one unique feature effect systems—or, more generally, custom monads designed to supercede IO—provide: they can do anything in between the lines [of code]. And I don't think that's a good thing.
Here are the advantages of running effects in plain IO:
- Works with any Haskell 98 (or later) compiler;
- Straightforward implementation;
- No extra dependencies;
- No weird type errors;
- Compiles fast;
- Runs fast.
Here are the disadvantages:
- Some functions will be verbose.
I hope this knowledge is used for things.