7.1 Modern Design Experience & The Unified Target Concept
Key Takeaways
The Modern Design Experience uses Use Application/Browser and Unified Target descriptors instead of the classic Open/Attach activities and single selectors.
A Unified Target can hold a strict selector, fuzzy selector, Computer Vision, image, and native text; UiPath ranks selectors first, Computer Vision second, and image third.
Default targeting methods differ by technology: desktop apps start with strict and fuzzy selectors, web browsers with the fuzzy selector only, and Computer Vision and image start off.
Input modes are Hardware Events, Simulate, Window Messages, and Chromium API; all but Hardware Events work without focus, and Chromium API works only for Chromium elements.
Activity settings override Use Application/Browser settings, which override the UI Automation Modern project settings.
7.1 Modern Design Experience & The Unified Target Concept
Core Concept: In enterprise Robotic Process Automation (RPA), user interface automation represents both the most powerful capability and the primary source of operational brittleness. UiPath's Modern Design Experience addresses this vulnerability through the Unified Target framework. Rather than relying on a single, fragile point of identification (such as a static XML selector), Unified Targeting combines several targeting methods (strict selector, fuzzy selector, Computer Vision, image, and, for some technologies, native text) into one descriptor, and uses them in a documented order of preference at runtime.
Historically, RPA developers working in UiPath Studio operated within the Classic Design Experience. Classic automation treated web browsers and desktop applications through separate, disjointed activity packages (such as Open Browser versus Open Application, or Attach Browser versus Attach Window). If a web developer altered a button's CSS class or modified an HTML element ID during a routine software release, the robot's single selector failed instantly, throwing a selector-not-found exception (for classic activities, UiPath.Core.SelectorNotFoundException) and halting unattended processing. The Modern Design Experience replaces this legacy architecture with unified containers, automated multi-technology targeting, integrated visual anchors, and centralized project governance.
1. Architectural Comparison: Modern vs. Classic Experience
The transition from Classic to Modern UI Automation represents a paradigm shift across design time, execution mechanics, and maintenance overhead.
| Architectural Dimension | Classic Design Experience | Modern Design Experience (Unified Target) |
|---|---|---|
| Container Scope | Divided into separate Attach Browser, Attach Window, Open Browser, and Open Application scopes. | Unified under Use Application/Browser, managing both desktop executables and browser sessions uniformly. |
| Targeting Paradigm | Single-technology targeting (primarily static XML selector). Anchors required a separate Anchor Base wrapper activity. | Unified Target descriptor embedding Strict Selector, Fuzzy Selector, Image, and Native Text with integrated multi-anchors. |
| Runtime Resilience | Binary: either the selector matches exactly or the activity fails upon timeout. | Ranked targeting methods: selectors first, then Computer Vision, then image, each only if enabled. |
| Input Method Selection | Configured per activity via discrete Boolean checkboxes (SimulateClick, SendWindowMessages). | Inherited hierarchically from Use Application/Browser or Project Settings; includes specialized Chromium API mode. |
| Data Scraping & Wizards | Fragmented tools: Screen Scraping (for text/OCR) and Data Scraping (for structured tabular data). | Unified Table Extraction wizard for structured multi-page grids and modern App/Web Recorder. |
| Object Repository | Limited, unsupported, or manual UI element library integration. | Native, deep integration with Object Repository for reusable, version-controlled UI Descriptors across projects. |
| Element Synchronization | Relies on separate Element Exists, Find Element, or Wait Element Appear coupled with conditional If branches. | Built-in Check App State activity providing dual-branch execution (Target appears / Target does not appear) and Verify execution. |
| Default Design Mode | The default in older Studio versions. | The default for new projects in current Studio releases; toggled via Project Settings. |
Enabling Modern Experience in UiPath Studio
In modern Studio installations, new projects default to the Modern Experience. For legacy projects or specific enterprise templates, the experience can be configured via Project Settings > General > Modern Design Experience:
- Setting the toggle to Yes enables modern activities, the Unified Target selection screen, modern recorders, and Object Repository binding.
- Setting the toggle to No restores classic activities in the Activities panel.
- Developers can access classic activities within a modern project by clicking the Filter icon in the Activities panel and selecting Show Classic.
2. The Unified Target Framework
The cornerstone of the Modern Experience is the Unified Target engine. When a developer indicates an element on screen using the modern selection wizard, Studio does not simply generate an isolated XML string. Instead, it extracts and binds four distinct targeting layers into a composite descriptor:
+-----------------------------------------------------------------------+
| UNIFIED TARGET DESCRIPTOR |
| |
| 1. STRICT SELECTOR |
| <html app='chrome.exe' title='Invoicing Portal' /> |
| <webctrl tag='BUTTON' id='btn_submit_order' aaname='Submit' /> |
| |
| 2. FUZZY SELECTOR |
| <webctrl tag='BUTTON' aaname='Submit' matching:aaname='fuzzy' /> |
| [Fuzzy level: 0.70 on a 0-to-1 similarity scale] |
| |
| 3. IMAGE AUTOMATION |
| [Visual Template: 120x36 px PNG] [Accuracy: 0.80] |
| [OpenCV Feature Matching & Pixel Bounding Box] |
| |
| 4. NATIVE TEXT |
| Text: "Submit Order" [OCR / Screen Scraping Extraction Engine] |
| |
| 5. INTEGRATED ANCHORS (1 to N References) |
| - Anchor 1: Strict/Fuzzy Selector for "Order Summary" Header |
| - Anchor 2: Visual Icon for Shopping Cart |
+-----------------------------------------------------------------------+
Layer 1: Strict Selector
The Strict Selector is a deterministic, hierarchical XML fragment that matches application and DOM attributes precisely. It relies on exact tag names, IDs, CSS classes, automation IDs, and accessibility names (aaname).
- Strengths: Maximum execution speed, zero ambiguity, completely deterministic.
- Weaknesses: Vulnerable to minor dynamic changes, such as autogenerated framework IDs or localized label updates.
Layer 2: Fuzzy Selector
The Fuzzy Selector relaxes exact string matching by comparing attribute values by similarity instead of equality. Each fuzzy attribute has a level between 0 and 1; the closer to 1, the more exact the match must be.
- Accuracy Threshold: If an attribute's similarity meets or exceeds its configured level, the candidate is accepted.
- Resilience: If a button's text shifts from
"Submit Order"to"Submit Orders"or"Submit order", a fuzzy selector with an accuracy of0.70successfully matches the element where a strict selector would fail.
Layer 3: Image Automation
The Image Automation layer captures a pixel-accurate visual snapshot of the target element. At runtime, UiPath utilizes OpenCV pattern matching algorithms to locate the visual template within the application window.
- Accuracy Threshold: Similar to fuzzy selectors, visual matching evaluates pixel similarity on a scale from
0.0to1.0(defaulting to0.80). - Utility: Acts as an essential fallback for non-standard UI frameworks, graphic canvases, legacy mainframe emulators, and virtual desktop infrastructure (VDI) environments (such as Citrix or VMware Horizon) where underlying DOM or accessibility nodes are inaccessible.
Layer 4: Native Text
The Native Text targeting method locates an element by the text the application renders, where the technology exposes it.
Layer 5: Computer Vision
Computer Vision uses UiPath's AI models to recognize UI elements such as buttons, fields, and labels from a screenshot. It is the secondary targeting method after selectors and is useful when a remote or custom interface exposes no reliable selector. In projects created with recent package versions it is off by default and must be enabled.
3. How the Targeting Methods Are Used at Runtime
When you indicate an element, Studio records every enabled targeting method for it: strict selector, fuzzy selector, Computer Vision, image, and, for some technologies, native text. UiPath's current UI Automation documentation describes a ranking that reflects each method's targeting power and resilience:
- Primary: selectors (the strict selector or the fuzzy selector).
- Secondary: Computer Vision.
- Tertiary: image (disabled by default).
At design time, Studio marks the method that is currently leading, simulating what will happen at runtime. A project setting, Wait for primary targeting method until timeout, controls whether the robot keeps waiting for the selector-based method instead of switching to the others.
Default targeting methods per technology
New projects start with these defaults, which you can change in Project Settings > UI Automation Modern or per activity in the selection screen:
| Targeting method | Desktop apps | Web browsers | Java | SAP |
|---|---|---|---|---|
| Strict selector | On | Off | On | On |
| Fuzzy selector | On | On | On | Off |
| Computer Vision | Off | Off | Off | Off |
| Image | Off | Off | Off | Off |
Projects created with older UIAutomation package versions had Computer Vision switched on by default for desktop, web, and Java targets, so an existing project may behave differently from a new one.
Anchors and timeouts
Anchors are checked for whichever method finds the element: the candidate must sit in the recorded position relative to its anchors. If no enabled method finds a valid match before the activity Timeout expires, the activity throws an exception. Healing Agent, covered in the Advanced UI Automation chapter, adds a further recovery layer on top of this.
Note
In the selection screen you can switch individual targeting methods on or off for one element. For example, turn off image targeting where visual lookalikes could produce a false match.
4. Runtime Execution Modes (Input Methods)
Locating an element is only half of the automation process; the robot must also deliver input events (clicks, text input, keyboard shortcuts) to that element. Modern UI Automation provides four runtime execution modes, each with distinct technical mechanics, speed ratings, and environment constraints:
+---------------------------------------------------------------------------------------+
| INPUT METHOD MECHANISMS |
| |
| 1. CHROMIUM API ──► Browser debugger APIs (Chrome DevTools Protocol) |
| [Works without focus | Chromium elements only] |
| |
| 2. SIMULATE ──► Accessibility APIs of the target technology |
| [Works without focus | Sends text in one action] |
| |
| 3. WINDOW MESSAGES ──► Win32 PostMessage / SendMessage (WM_LBUTTONDOWN, WM_CHAR) |
| [Works without focus | Desktop apps, Windows only] |
| |
| 4. HARDWARE EVENTS ──► OS Level Mouse Driver & Keyboard Scan Code Injection |
| [Foreground Only | Slowest | Maximum Compatibility] |
+---------------------------------------------------------------------------------------+
1. Chromium API
- Mechanics: Performs actions through the browser's debugger APIs (the Chrome DevTools Protocol) in Chromium-based browsers such as Chrome and Edge.
- Capabilities: Works even when the target browser is not in focus and sends all text in one action.
- Limitations: Works only for Chromium elements; it cannot drive native desktop windows or non-Chromium browsers such as Firefox.
2. Simulate
- Mechanics: Directly invokes the target application's internal event handling mechanisms or DOM methods (for example, executing
HTMLElement.click()or dispatching a DOM change event in web applications, or triggering Win32 button notification codes). - Capabilities: Operates fully in the background without stealing focus or moving the mouse pointer. When typing, it populates the target element's text property instantly in a single operation rather than typing character-by-character.
- Limitations: Does not support hardware-level hotkeys (such as
Ctrl+Alt+Delor physical function keys), mouse hover states, or drag-and-drop gestures that require continuous physical coordinate tracking.
3. Window Messages
- Mechanics: Posts Windows operating system messages directly into the target control's message queue (
HWND) via Win32 API functions (PostMessage/SendMessage), sending messages such asWM_LBUTTONDOWN,WM_LBUTTONUP,WM_CHAR, orWM_SETTEXT. - Capabilities: Supports background execution for many traditional desktop applications (C++, WinForms, legacy ERPs) that do not support Simulate.
- Limitations: Does not support universal web DOM events; some custom owner-drawn desktop controls ignore synthetic Win32 messages.
4. Hardware Events
- Mechanics: Injects physical mouse and keyboard scan codes directly into the Windows operating system input stream. The robot physically moves the mouse pointer to the element's screen coordinates and generates genuine hardware click and keypress events.
- Capabilities: 100% compatibility across all UI frameworks, custom Canvas elements, games, and Citrix/VDI environments. Generates full mouse hover triggers and supports all keyboard combinations.
- Limitations: Requires the application to remain in the active foreground with an unlocked desktop session. The robot steals user control, rendering the runner machine unusable for human operators during execution. Unsuitable for concurrent background automations.
Input Method Comparison
| Input method | How it works (UiPath documentation) | Works in the background | Typical use |
|---|---|---|---|
| Hardware Events | Real mouse and keyboard input sent to the operating system; full behavioral emulation, though some events may occasionally be lost | No; the window needs focus | Canvas apps, hotkeys, hover menus, anything the other modes cannot reach |
| Simulate | Uses accessibility APIs; sends all text in one action | Yes | Browsers, Java, and SAP, where supported |
| Window Messages | Sends Win32 messages; sends all text in one action | Yes | Classic desktop applications (Windows only) |
| Chromium API | Uses the browser debugger APIs; works only for Chromium elements | Yes | Chrome and Edge pages |
Some options depend on the input mode. For example, Type Into's Empty field, Click before typing, and Delay between keys options cannot be used with Simulate.
5. Modern UI Automation Project Settings
Project Settings > UI Automation Modern holds the defaults that new activities in the project use. They include timeouts and delays, the input mode, the window attach mode of new Use Application/Browser activities (Application instance searches the application's windows and pop-ups; Single window searches only the indicated window), the browser automation mode (UIAutomation v26.10 and later), and the targeting methods enabled for desktop, web, Java, and SAP targets.
Configuration Inheritance
Activity properties follow a simple hierarchy:
- Activity level. A value set directly on the activity, for example Input mode = Hardware Events on one Click, wins.
- Container level. Activities set to Same as App/Browser take the value from their enclosing Use Application/Browser activity.
- Project level. The container itself falls back to the project settings.
This lets a team switch many activities from Hardware Events to Simulate or Chromium API by changing one container or one project setting, and then test the few activities that need an exception.
How does a fuzzy selector decide whether an on-screen element matches the recorded attribute value?
It hashes the attribute value and requires an exact hash match.
It compares attribute values by similarity and accepts a candidate whose similarity meets the configured fuzzy level, a value between 0 and 1.
It runs OCR over the whole desktop and applies a regular expression.
It evaluates an XPath wildcard against the browser's root document.
A web page is redeployed and an input's HTML id changes, while its label and position stay the same. The Type Into activity uses the default web targeting methods plus a label anchor. What happens at runtime?
The activity fails immediately because the strict selector no longer matches.
Studio opens the Repair wizard on the robot machine.
The fuzzy selector, which is on by default for web targets, can still identify the field, and the anchor confirms its position, so the action succeeds.
The robot switches the input mode to Window Messages to reach the field.
Which Modern UI input method operates by communicating directly with Chromium-based browsers via the Chrome DevTools Protocol (CDP), providing high execution speed and 100% background automation capability without stealing desktop focus?
Chromium API
Hardware Events
Simulate
Window Messages
Sections you finish are checked off in the contents.