Introduction
Various tools and patterns improve performance while you author and design your workspace. For other performance tuning tips and tricks, see Performance Tuning FME.
Workspace Authoring
A few FME tools help a workspace run faster while you author and debug it, speeding up the authoring process.
Data Caching (previously Feature Caching)
A best practice for authoring (and debugging) a workspace is to make small changes and run it often to test results. This makes it easier to pinpoint a problem when one occurs, rather than having to check a large number of untested readers, writers, and transformers.
Frequently running a workspace can be time-consuming. This is where Data Caching comes in.
Data Caching stores the results of a workspace at every step of the translation. Having this data at hand helps you to inspect the results at each stage. However, running a translation doesn't mean running an entire workspace. Instead, you can use cached results instead of re-running earlier parts of a workspace.
This technique saves time when re-running a workspace, although it consumes disk space and resources when writing cache. Additionally, it can give a false impression of performance when a workspace is put into production without caching.
Data Caches are not stored within the workspace; they are a setting in FME Workbench. If you use FME Workbench to run your workspaces after they are authored, turning off Data Caching will improve performance.
Performance tools, such as Parallel Processing, also won’t work if Feature Caching is enabled. For optimal performance, run a workspace in either the FME Quick Translator or via the Command Line. Data Caches are not published to FME Flow.
A best practice for authoring (and debugging) a workspace is to make small changes and run it often to test results. This makes it easier to pinpoint a problem when one occurs, rather than having to check a large number of untested readers, writers, and transformers.
Frequently running a workspace can be time-consuming. This is where Feature Caching comes in.
Feature Caching stores the results of a workspace at every step of the translation. Having this data at hand helps you to inspect the results at each stage. However, it also means you don't have to run the entire workspace again. Instead, you can use cached results instead of re-running earlier parts of the workspace.
This technique saves time when re-running a workspace, although it consumes disk space and resources when writing the cache. Additionally, it can give a false impression of performance when you put a workspace into production without caching.
Feature Caches are not stored within the workspace; they are a setting in FME Workbench. If you use FME Workbench to run your workspaces after they are authored, turning off Feature Caching will improve performance.
Performance tools, such as Parallel Processing, also won’t work if Feature Caching is enabled. For optimal performance, run the workspace in either the FME Quick Translator or via the Command Line. Feature Caches are not published to FME Flow.
Record Counts (previously Feature Counts)
Record Counts are the numbers that appear on connections (links) when you run a workspace in FME Workbench. The real-time animation of these counts often indicates which transformers are blocking your data or slowing the workspace.
A combination of Data Caching and Record Counts shows record counts and caches the data at each link on the workspace, which can be particularly useful. In 2025.1, you can turn off Record Counts, which might improve performance (first image). In 2025.2 and newer, this can be turned off in FME Options > Translation > Enable Record Counting (second image).
Feature Counts are the numbers that appear on connections when you run a workspace in FME Workbench. The real-time animation of these counts often indicates which transformers are blocking your data or slowing the workspace.
Feature Caching and Feature Counts both show feature counts and cache the data at each link on the workspace, which can be particularly useful. In newer versions of FME, You can turn off Feature Counts, which might improve performance.
Workspace Design Patterns
Design your workspace for maximum performance as you create it. Various design patterns can improve performance.
Attribute Cleaning
During a translation, FME holds your data either in memory (physical or virtual) or in a cache on disk. Obviously, performance is improved by reducing the amount of unnecessary data being held, both features and the components of those features.
One particular aspect is an excess of attributes. Often, not all attributes on the source data are required for processing or output. In this scenario, it's better to avoid reading those attributes or use a transformer to remove them as soon as possible in your translation.
Design Pattern:
- Read Data
- Remove Excess Attributes
- Process Data
In other words, carry through only the geometry and attributes you intend to keep in the output. You can use an AttributeManager, AttributeRemover, AttributeKeeper, or GeometryRemover transformer to remove excess components, and you should use it as early in the translation as possible.
Note that the AttributeKeeper transformer can create bulk features, which will significantly speed up downstream processing. This feature is highly recommended when more than two transformers follow the AttributeKeeper.
In particular, lists (an attribute with multiple values) or geometry stored as attribute values can take up resources because they tend to carry larger amounts of data. For example, a feature joined to 1,000 records and storing those records as a list is now the equivalent of 1,000 records!
Another way to reduce attributes is to not read them at all. See the section on databases (below) for more information.
Data Filtering
Similarly to excess attributes, excess records (previously features) use valuable system resources and should be removed as early in the translation as possible.
Design Pattern:
- Filter Data
- Process Data
In other words, filter out unwanted data so it doesn't incur unnecessary processing.
In this example, the author measures the area of features and then filters the data:
This is entirely the wrong way around. The workspace wastes time measuring features that are later discarded (Tester:Failed).
Reducing Duplication
A common bad design pattern (sometimes called an Anti-Pattern) is a chain of duplicate transformers. This is often less efficient than processing data entirely in a single transformer.
For example, chaining ExpressionEvaluator transformers this way, where each carries out a different step of one overall calculation, is not the most efficient way to process data. It would be far better to condense the actions into a single AttributeManager transformer.
Tip: A chain of duplicate transformers is an anti-pattern, and you should investigate whether there is a better way to achieve your goal.
Memory Reduction Strategies
The FME Engine passes features through in a mix of ways to maximize performance.
Some transformers in FME Workbench operate on one feature at a time. These are known as Feature-Based Transformers. They can operate on one feature at a time because the process they carry out does not need different features to interact.
Other transformers work on groups of features. These are known as Group-Based Transformers. Group-based transformers process multiple features simultaneously; for example, they may intersect many line features to produce a topological network.
Creating bulk features ahead of a group-based transformer with an AttributeKeeper improves performance because, while the data waits to be processed, it occupies less memory.
Obviously, any transformer that works on a group of features must hold them all in memory at a time, incurring processing costs. This is known as Feature Holding, and should ideally be considered when designing an overall FME strategy.
Tip: The FME transformers documentation contains information about whether a transformer is Group or Feature-Based, and whether it is a Feature Holding transformer.
Fortunately, many group-based transformers can reduce their memory footprint under certain conditions by holding less data. However, the workspace author must set a parameter to confirm these conditions are met.
For example, the DuplicateFilter transformer separates features containing a duplicate attribute key. The transformer can run more efficiently if it knows the data is already sorted by key, but the workspace author must set the Input is Ordered parameter to confirm this.
Tip: Many feature-holding transformers have parameters that can improve performance under the right conditions. Group By (previously Group-By Mode) is one such parameter and is common to many transformers. Clippers Arrive First (previously Clippers First) is a parameter that only applies to one transformer (the Clipper). Inspect the documentation closely to look for ways to reduce the load on group-based transformers.
Writer Order
Multiple writers in a workspace execute in a specific order. The first writer in the list writes first, while subsequent writers cache data until it is their turn.
Performance is hindered by caching data. Therefore, it makes sense to move the writer handling the most data to the top of the list so its data isn't cached.
Writers can be ordered by dragging them up and down the list in the Navigator window.
Additionally, the Workspace Parameters > Order and Redirect> Order Writers By allows you to set the order in which writers are executed (either the order in the Navigator or the order in which features arrive).
Tip: When you have multiple writers in a workspace, always ensure the one getting the larger amount of data is the first writer in the list. It is also worth Creating Bulk Mode features with the AttributeKeeper prior to any writer, except the first writer, to reduce the storage footprint.
Tip: For peak performance, tiles output from the RasterTiler or WebMapTiler should be written in the order they are output from these transformers.
Database Design Patterns
Reading from and writing to databases provide unique opportunities for performance tuning. In general, processing can be quicker in the database itself than in the FME Engine. See the tutorial series Let the Database Do the Work for step-by-step instructions.
When reading data from a database, the ideal pattern is to filter the data as it is read, rather than in FME:
Design Pattern:
- Filter Data
- Read Data
- Process Data
Or...
Design Pattern:
- Read Data (with a Filter)
- Process Data
A "filter" can be applied by using the SQL Statements and SQL WHERE clauses available in most FME database readers.
Another way to "filter" data is to not read unnecessary attributes from the database table. These can be turned off in FME feature types by unchecking the attribute in the User Attributes tab:
When joining data, SQL Joins can be quicker than the FeatureJoiner transformer. Create a database materialized view for better performance and to simplify your workspace (though Database Administrators may not allow it).
Legacy: For ArcSDE 10.2 or older, SQL Statements are only supported for non-spatial tables. For spatial tables use: sdetable -o create_view to create a view that contains a spatial column in the join.
You can create complex table joins by combining sdetable -o create_view with SQL ALTER VIEW.
With 10.3+, you can create views directly in SQL.
Tip: Performance improves in some cases by handing off processing to a database.
Maximize Use of System Resources
It may seem strange to suggest using as many resources as possible, but this is acceptable as long as they are doing useful work. For example, if your computer has an 8-core CPU, it makes sense to split your work into eight parts, so that each core can do its share. Restricting the process to a single core does not make the best use of system resources.
For FME Flow, the recommendation is to have one engine per core, although this can vary depending on the type and size of data being processed.
In terms of memory, it can make sense to read the entire contents of a database table in one query - regardless of whether every feature is required or not - in order to have the required records immediately available to FME. Reading individual records on demand, though it incurs lower network traffic, may overall be a slower strategy.
Partitioning and Load Balancing
An extremely large amount of data can be difficult to process in a single task. You may want to consider partitioning the data into groups (e.g., maybe on the basis of a geographic region) and processing each separately.
For example, you could divide the entirety of a Canada-wide dataset into a separate group per province, using a WHERE clause to select the required data. This way, the data is divided into ten different groups.
At this point, either each group is processed consecutively (so that approximately only 10% of the data is processed at any one time by FME Engine) or each group is processed concurrently over a number of engines.
Miscellaneous Design Tips
Tip: For maximum performance, in an arithmetic editor, use the functions @add(), @mult(), and @div() instead of using the equivalent operators (+ * and /). These functions work at a lower, much faster level of processing. Plus, they have the added bonus of handling nulls better.