- Overview
- The form control class hierarchy
- Finding form controls
- Clicking, typing, focus and selection
- Checkboxes and radio buttons
- File inputs
- Buttons and form submission
- Value handling
- getValue(), getValueAttribute() and getRawValue()
- Three ways to set a value
- HtmlNumberInput: parsing and sanitization
- HtmlTextArea: the dirty value flag and reset
- Constraint validation
- willValidate(): is this control barred from validation?
- isValid(), isValidValidityState() and the individual validity checks
- checkValidity(), reportValidity() and the 'invalid' event
- Custom validity and validation messages
- HtmlButton: validation specifics
- HtmlSelect: validation specifics
- HtmlSelect: the size attribute and default selection
- HtmlSelect and required
Overview
This page is a guide to working with HTML form controls through HtmlUnit's plain Java API. It covers three things, in order: the general shape of the form-control classes and the everyday operations you'll use on them; the different, easily-confused ways to read and write a control's value; and the full HTML Constraint Validation API as HtmlUnit implements it.
The form control class hierarchy
Every form control in a parsed page is represented by a subclass of
HtmlElement. The two most important base types are:
HtmlInput-- the base class for every<input>element, regardless of itstypeattribute.HtmlSelectableTextInput-- anHtmlInputsubclass that adds text-selection and simulated-typing support, used by every text-like input type.
HtmlUnit creates a different concrete class per type attribute value, so
that each type can implement exactly the constraints and behavior real browsers give
it. The most commonly used ones:
HtmlTextInput(type="text", and the default whentypeis absent or unrecognized)HtmlNumberInput,HtmlEmailInput,HtmlUrlInput,HtmlTelInput,HtmlSearchInput,HtmlPasswordInput-- all text-like, each with its own additional constraint (see Constraint validation below)HtmlCheckBoxInputandHtmlRadioButtonInput-- boolean "checked" controls rather than text controlsHtmlFileInput-- represents a file-upload control; see File inputs belowHtmlSubmitInput,HtmlResetInput,HtmlButtonInput,HtmlImageInput-- the various button-like input typesHtmlHiddenInput-- never a candidate for constraint validation and never focusable, but otherwise a plain value holderHtmlDateInput,HtmlTimeInput,HtmlColorInput,HtmlRangeInputand the other less common HTML5 types each have their own class as well.
Outside the <input> family, HtmlSelect,
HtmlTextArea and HtmlButton (the standalone
<button> element, not <input type="submit">) are
the other major form-control classes; both value handling and validation apply to
them too, and are covered later on this page.
If a script changes an input's type attribute at runtime, HtmlUnit
internally replaces the Java object backing that DOM node with an instance of the
correct class for the new type, carrying the old value across via
adjustValueAfterTypeChange() -- each subclass decides for itself what,
if anything, of the old value still makes sense under the new type (for example,
switching into type="file" always clears the value, since a
real browser can never carry a text value over into a file selection).
Finding form controls
The usual DOM-style lookups work as expected --
page.getHtmlElementById(String),
page.getElementByName(String), or the generic
DomElement traversal/query methods. HtmlForm additionally
offers form-scoped convenience lookups that are usually more direct when you already
know which form you're working with:
final HtmlForm form = page.getFormByName("login");
final HtmlTextInput username = form.getInputByName("username");
final List<HtmlInput> allChecked = form.getInputsByValue("on"); // e.g. matching checkboxes
final List<HtmlRadioButtonInput> options = form.getRadioButtonsByName("plan");
HtmlForm.getElements() / the JavaScript-visible elements
collection return every "listed" control associated with the form -- inputs, selects,
textareas, buttons, fieldsets and objects -- in document order, regardless of whether
they're physically nested inside the <form> tag or associated with
it purely via a form="..." attribute elsewhere in the page.
Clicking, typing, focus and selection
click() simulates a real mouse click, including whatever side effects a
real click has for that specific control -- a submit button submits its form, a reset
button resets it, a checkbox toggles its checked state, and so on. This is almost
always the right method to call to simulate user interaction with a button, rather
than calling a lower-level state-changing method directly.
type(String) (available on any HtmlSelectableTextInput or
HtmlTextArea) simulates real keystroke-by-keystroke typing, character by
character, including firing the same key and input events a real browser would, and
respecting whatever character-level restrictions the control type imposes (for
example, a number input silently ignores non-numeric keystrokes unless the relevant
browser-version feature says otherwise). This is different from -- and usually more
faithful to real-browser behavior than -- calling setValue(String),
which behaves like scripted assignment (comparable to JavaScript's .value =)
rather than user input. See Three ways to set a value below for the concrete
differences this can produce.
focus() and removeFocus() move keyboard focus to or away
from a control, the same way clicking into a field or tabbing away from it would.
select() (on selectable text controls) selects the entire current value,
and getSelectedText(), getSelectionStart()/
setSelectionStart(int) and getSelectionEnd()/
setSelectionEnd(int) give fine-grained access to the current text
selection/cursor position, mirroring the same-named JavaScript properties.
Checkboxes and radio buttons
HtmlCheckBoxInput and HtmlRadioButtonInput don't hold a
free-form text value the way other inputs do -- their state is a simple boolean,
accessed via isChecked() / setChecked(boolean). Calling
setChecked(true) on a radio button automatically unchecks every other
radio button in the same named group within the same form, exactly as a real browser
would.
isDefaultChecked() reflects the checked attribute as parsed
from markup -- the state the control reverts to on a form reset -- independently of
whatever the live isChecked() state currently is, the same relationship
getDefaultValue() has to getValue() for text controls.
final HtmlRadioButtonInput yearly = form.getRadioButtonsByName("plan").stream()
.filter(r -> "yearly".equals(r.getValueAttribute()))
.findFirst().orElseThrow();
yearly.setChecked(true);
File inputs
HtmlFileInput deliberately does not support setting its value
to an arbitrary string the way a text input does -- real browsers never allow a page
to pre-fill a file chooser's selection for security reasons, and HtmlUnit follows the
same rule. Instead, use one of:
setFiles(File...)-- point the control at one or more real files on disk, as if the user had picked them in a native file-chooser dialog. Passing more than one file requires themultipleattribute to be present, or anIllegalStateExceptionis thrown.setDirectory(File)-- recursively selects every file in a directory, for controls with the non-standardwebkitdirectoryattribute.setData(byte[])/getData()andsetContentType(String)/getContentType()-- an HtmlUnit-only convenience with no browser equivalent at all: lets a test supply the bytes (and MIME type) to actually submit, without needing a real file on disk. When set, this in-memory data is used during form submission instead of reading the selected file's real contents.
getValue() on a file input never returns the real file path -- it
returns the empty string if nothing is selected, or the same
C:\fakepath\filename-style placeholder real browsers use, for the same
privacy reason setting an arbitrary value isn't allowed in the first place.
Value handling
The remainder of this page covers two topics in detail: how to read and write a
control's value correctly, and how HtmlUnit's constraint-validation API works. Both
apply across every text-like HtmlInput subtype, HtmlTextArea,
and (with its own specific rules, covered separately) HtmlSelect.
getValue(), getValueAttribute() and getRawValue()
These three methods are easy to confuse because they can return the same string in the simplest cases -- but they answer three different questions, and they diverge as soon as a script (or a real user) interacts with the field.
getValueAttribute() reads the literal value="..."
attribute as it currently stands in the HTML/DOM tree. This is the control's
default value -- what the field would reset back to on a form reset. It
never changes just because someone typed into the field or called
setValue(); it only changes if the attribute itself is modified (via
setValueAttribute(), or the equivalent defaultValue
property from JavaScript).
getRawValue() returns the control's current, unsanitized internal
value -- whatever is "really" stored right now, decoupled from the DOM attribute.
The first time a script calls setValue() or the user types into the
field, the control's dirty value flag is set, and from that point on
getRawValue() reflects the live edited content, independent of whatever
the value attribute says.
getValue() is what most code should actually call. For a plain text
input it is effectively the sanitized form of getRawValue(). For some
control types (see below for HtmlNumberInput) it additionally normalizes
or rejects content that getRawValue() would return unchanged.
public void value() throws Exception {
try (WebClient webClient = new WebClient()) {
final HtmlPage page = webClient.getPage("https://www.wetator.org/testform/");
final HtmlTextInput input = page.getHtmlElementById("project");
// before any edit, all three agree
System.out.println(input.getValueAttribute()); // "initial"
System.out.println(input.getRawValue()); // "initial"
System.out.println(input.getValue()); // "initial"
input.type("my fancy HtmlUnit project");
// the attribute never moved -- it is still the DEFAULT value
System.out.println(input.getValueAttribute()); // "initial"
// but the live value has
System.out.println(input.getRawValue()); // "edited"
System.out.println(input.getValue()); // "edited"
}
}
Calling reset() (the same method a form's reset algorithm calls on every
one of its controls) clears the dirty flag and re-derives the live value from the
current default value -- not necessarily the value the page had when it was
first loaded, since a script may have changed defaultValue/the
value attribute in the meantime.
input.setDefaultValue("changed-default");
input.reset();
System.out.println(input.getValue()); // "changed-default", not "initial"
Three ways to set a value
A text-like input's value can be changed through three genuinely different code
paths, and it's worth being explicit about them because they don't always apply the
same rules -- this matters most for HtmlNumberInput (see the next
section), but the general distinction applies more broadly:
- The initial
valueattribute, parsed once when the page loads. setValue(String)(and the equivalent JavaScript.valuesetter) -- scripted assignment, comparable to a program directly overwriting the field, with no simulation of individual keystrokes.type(String)-- simulated real typing, character by character, which can expose different intermediate states than a singlesetValue()call would, and respects per-keystroke input filtering thatsetValue()does not.
For most control types these three paths end up agreeing on the final result. Where they diverge, it's almost always because a specific control type applies extra sanitization or leniency to one path but not another -- see the next section for a concrete, fully worked-out example.
HtmlNumberInput: parsing and sanitization
HtmlNumberInput layers extra behavior on top of the three methods above,
because a number input's textual content has to be checked against the HTML
specification's own, fairly strict, floating-point-number grammar -- which is
not the same grammar as a Java or JavaScript numeric literal.
The grammar (implemented by the internal HtmlNumberParser helper) accepts:
- an optional single leading
-(a leading+is not accepted -- this is a common surprise, since it differs from JavaScript's ownNumber()/parseFloat()), - one or more ASCII digits, optionally with a single decimal point and further digits,
- an optional exponent (
eorE), which -- unlike the number's own leading sign -- does accept both+and-, - nothing else: no whitespace anywhere, no thousands separators, no
NaN/Infinity, no hex/binary/octal notation.
getValue() on a number input returns the value trimmed and
locale-normalized if it parses successfully, or the empty string if it does not
(unless the rare JS_INPUT_NUMBER_ACCEPT_ALL browser feature is active for
the configured BrowserVersion, in which case the raw text is returned
even when unparseable). getRawValue() always returns the literal
characters that are currently in the field, valid or not.
final HtmlNumberInput number = page.getHtmlElementById("qty");
number.setValue("42");
System.out.println(number.getValue()); // "42"
number.setValue("not-a-number");
System.out.println(number.getValue()); // "" -- sanitized away
System.out.println(number.getRawValue()); // "not-a-number" -- untouched
number.setValue("+5");
System.out.println(number.getValue()); // "" -- leading '+' is never valid
This is the clearest example of the three value-setting paths genuinely disagreeing.
Take the string "1." -- a digit followed by a decimal point with nothing
after it:
- Set via the initial
valueattribute, some configured browser versions silently correct it to"1", while others reject it outright to the empty string. - Set via
setValue("1."), it is rejected outright to the empty string, uniformly, with no browser-specific leniency at all. - Reached via
type("1.")(simulating the user typing1then.), the field's live value after those two keystrokes reports"1"-- the longest already-valid prefix of what has been typed so far -- not the raw"1."and not the empty string. A further keystroke continues from the full typed text, so typing2next correctly produces"1.2", not"12".
If your test needs to simulate a user filling in a number field, prefer
type(String) over setValue(String) whenever the exact,
keystroke-by-keystroke intermediate state matters -- for example, when asserting on
checkValidity() partway through typing, rather than only after the field
is fully filled in.
HtmlTextArea: the dirty value flag and reset
A <textarea>'s value works differently from an
<input>'s in one important respect: an input's value always lives
in a plain string attribute, but a textarea's default value is its child
text content -- the literal text between the opening and closing tags. HtmlUnit
exposes the live value through getText() / setText(String)
(mirroring the JavaScript value property), and the default through
getDefaultValue() / setDefaultValue(String) (mirroring
defaultValue), the same pairing as for HtmlInput.
The relationship between the two is governed by a dirty value flag, exactly
as the HTML specification defines it. While the flag is clean (the textarea has never
been edited by script or by the user), getText() is always computed
fresh from the current child text content -- so mutating the DOM children directly
(for example, via setDefaultValue(), or lower-level DOM manipulation)
is immediately reflected in getText(). The moment setText()
is called (or the user types into the field), the flag becomes dirty, and from then
on the live value is held independently -- child-content changes no longer affect it
at all, until a reset clears the flag again.
final HtmlTextArea comments = page.getHtmlElementById("comments"); // <textarea>seed</textarea>
System.out.println(comments.getText()); // "seed"
comments.setText("edited"); // dirties the flag
System.out.println(comments.getText()); // "edited"
// once dirty, further changes to the underlying default no longer
// affect the live value
comments.setDefaultValue("new-default");
System.out.println(comments.getText()); // still "edited"
System.out.println(comments.getDefaultValue()); // "new-default"
reset() clears the dirty flag and re-derives the live value from
whatever the child text content currently is -- not the page's originally
parsed text, and not necessarily whatever getDefaultValue() last
returned either, if the children were mutated some other way while the flag was
still dirty (in which case those mutations had no visible effect on
getText() until this exact moment, since they happened while the flag
was dirty and therefore ignored).
comments.reset();
System.out.println(comments.getText()); // "new-default" -- picks up the current child content
Like text-type inputs, setting the value to the value it already holds is treated as
a near no-op: the dirty flag still becomes (or stays) set, but the text entry cursor
is not moved, unlike a genuine change, which always moves the cursor to the
end of the new value. readonly is a real, spec-supported attribute for
<textarea> (unlike for <select> or
<button>, covered below) and bars the control from constraint
validation the same way disabled does.
Cloning a textarea (cloneNode(boolean)) correctly carries the dirty flag
and live value across to the clone as independent state -- a clone of a dirtied
textarea keeps the dirtied value, decoupled from its own (independently cloned) child
content, and typing into the original afterward has no effect on the clone or vice
versa.
Constraint validation
Every control that can participate in HTML's Constraint Validation API implements
ValidatableHtmlElement. The methods on that interface answer three
related but distinct questions:
- Does this control take part in validation at all? --
willValidate() - If it does take part, does its current value satisfy its own
constraints? --
isValid()/isValidValidityState()and the individualisXxxValidityState()/hasXxxValidityState()methods - Combining both of the above, is the control valid right now, and why
not if it isn't? --
checkValidity(),reportValidity(),getValidationMessage()
It is important to understand that question 2 on its own deliberately
ignores question 1. A disabled, empty, required text input is "invalid" in the pure
constraint-checking sense (isValueMissingValidityState() returns
true), but it is also barred from validation entirely
(willValidate() returns false), and the two are combined
only at the higher-level methods. This is exactly how real browsers behave: a
disabled field never blocks form submission and never shows a validation message,
no matter how "wrong" its content looks.
willValidate(): is this control barred from validation?
willValidate() returns false if the control is barred from
constraint validation, and true otherwise. A control is barred when any
of the following apply:
- it has the
disabledattribute set, directly or inherited from an enclosing<fieldset disabled>-- except for controls that are themselves inside that fieldset's first<legend>child, which are exempt from the fieldset's disabling, - it is a text-like input or a textarea with the
readonlyattribute set (readonly on a<select>or<button>has no such effect, since neither element supports a realreadonlyattribute), - it is an
<input type="hidden">, - it is a
<button>withtype="button"ortype="reset"(onlytype="submit"buttons -- the default -- participate in validation), - it is a
<fieldset>or<object>element -- these always reportfalse, since neither is itself a candidate for constraint validation, regardless of any attribute.
final HtmlTextInput required = page.getHtmlElementById("required");
final HtmlTextInput disabled = page.getHtmlElementById("disabledRequired");
System.out.println(required.willValidate()); // true
System.out.println(disabled.willValidate()); // false, even though it's also required+empty
isValid(), isValidValidityState() and the individual validity checks
isValid() (and the near-identical isValidValidityState())
answer the pure constraint question, without regard to
willValidate(). Internally each returns true only if none
of the following individual checks report a problem. Each check corresponds directly
to one property on the standard DOM ValidityState object, and each is
individually callable, which is useful for building a custom error message without
relying on the generic text from getValidationMessage():
isCustomErrorValidityState()--trueifsetCustomValidity(String)was called with a non-empty message and it has not since been clearedisValueMissingValidityState()--trueif the control has arequiredattribute and no value (for<select>, this specifically means the auto-selected first option has an empty value and nothing else was explicitly selected)hasPatternMismatchValidityState()-- value present but doesn't match thepatternattributeisTooLongValidityState()/isTooShortValidityState()-- value longer/shorter thanmaxlength/minlength(only once the control's value has actually been edited by the user or a script -- an over-long initial attribute value alone does not trigger this)hasTypeMismatchValidityState()-- value doesn't match the required syntax for the control's type (e.g.email,url)hasRangeOverflowValidityState()/hasRangeUnderflowValidityState()-- numeric value outsidemax/min(number/range inputs)isStepMismatchValidityState()-- numeric value not aligned to thestepattribute, measured frommin(or zero, ifminis absent)hasBadInputValidityState()-- the user's input could not be converted at all (for example, a number input whose text simply doesn't parse as a number). This is distinct from a range or step violation: a value that fails to parse can never be "too big" or "misaligned", it just isn't a number in the first place.
final HtmlNumberInput qty = page.getHtmlElementById("qty");
qty.setValue("not-a-number");
System.out.println(qty.hasBadInputValidityState()); // true
System.out.println(qty.hasRangeOverflowValidityState()); // false -- there's no number to compare
System.out.println(qty.isStepMismatchValidityState()); // false -- same reason
System.out.println(qty.isValid()); // false, because of badInput
checkValidity(), reportValidity() and the 'invalid' event
These are the two methods application code (and HtmlUnit's own form-submission logic)
should normally call, rather than combining willValidate() and
isValid() by hand every time.
checkValidity() performs static validation: it returns
true if the control is either barred from validation or actually valid,
and false otherwise. If it returns false, a cancelable
invalid DOM event is fired on the control -- but no focus change happens.
reportValidity() does exactly the same static check and fires the same
event, but additionally moves focus to the control if it is invalid. This mirrors
the real-browser distinction between a script silently checking validity
(checkValidity()) and a user-facing "please fix this field"
interaction (reportValidity()).
A <form>'s own checkValidity()/
reportValidity() run this same check across every one of the form's
associated controls. Barred controls are skipped entirely -- neither counted as
passing nor failing -- and every genuinely invalid control (not just the first one
found) fires its own invalid event. reportValidity() at
the form level focuses only the first invalid control it finds.
final HtmlForm form = page.getHtmlElementById("myForm");
if (!form.checkValidity()) {
// at least one control is genuinely invalid; every failing control
// has already had its 'invalid' event fired
}
Custom validity and validation messages
setCustomValidity(String message) lets application code (or a test)
mark a control as invalid for a reason the built-in constraints don't express, by
supplying a non-empty message. Passing an empty string clears it again. A non-empty
custom message always wins over every other constraint -- but, like every other
constraint, it is still ignored entirely if the control is barred from validation
(willValidate() is false).
getCustomValidity() returns the raw string most recently passed to
setCustomValidity(), or the empty string if none is currently set.
getValidationMessage() returns a human-readable description of why the
control currently fails validation, or the empty string if it is valid or barred.
If a custom validity message is set, that exact string is returned unchanged.
Otherwise, the message describes whichever single built-in constraint is failing,
checked in this priority order: custom error, pattern mismatch, range overflow,
range underflow, step mismatch, too long, too short, type mismatch, bad input, value
missing. The wording for each built-in case is generic by design (for example,
"Please fill out this field." for a missing value) -- some control
types with control-specific data available (such as a number input's actual
min/max) may produce a more specific message that
interpolates the relevant value.
final HtmlTextInput email = page.getHtmlElementById("email");
email.setCustomValidity("We already have an account with this address.");
System.out.println(email.checkValidity()); // false
System.out.println(email.getValidationMessage()); // "We already have an account with this address."
email.setCustomValidity("");
System.out.println(email.checkValidity()); // back to whatever the built-in constraints say
HtmlSelect: validation specifics
HtmlSelect has its own value-and-selection model, since a
<select> doesn't hold a single text value the way an
<input> does -- it holds a set of <option>
children, zero or more of which are selected.
getOptions(), getOptionSize(), getOption(int)
and getSelectedOptions() all search the entire descendant tree
of the select, not just its direct children -- so options grouped inside one or more
<optgroup> elements are found correctly, exactly as in a real
browser:
<select id="country">
<optgroup label="Europe">
<option value="de">Germany</option>
<option value="fr">France</option>
</optgroup>
<optgroup label="Asia">
<option value="jp">Japan</option>
</optgroup>
</select>
final HtmlSelect select = page.getHtmlElementById("country");
System.out.println(select.getOptionSize()); // 3 -- both optgroups are searched
getSelectedOptions() behaves differently depending on whether the
select allows multiple selections (isMultipleSelectEnabled()):
- in multiple-select mode, it returns every option whose
isSelected()istrue, in document order; - in single-select mode, if more than one option is (erroneously) marked selected, only the last one encountered in document order is returned; if nothing at all is currently selected, an empty list is returned -- it does not fall back to "the first option" on its own.
That last point is worth calling out explicitly: a single-select <select>
with at least one option is normally never in a truly unselected state, because
HtmlUnit automatically selects the first option for you once, when the page finishes
loading (if nothing else was explicitly selected), and again automatically after
removing or replacing options via removeOption()/
replaceOption(). But if every option is deselected some other way (for
example, calling setSelectedIndex(-1) directly, or removing the
selected state from script), getSelectedOptions() will
genuinely return an empty list until something re-establishes a selection.
final HtmlSelect select = page.getHtmlElementById("country");
// right after page load, single-select mode: the first option is
// auto-selected if nothing else was
System.out.println(select.getSelectedOptions()); // [Germany]
select.setSelectedIndex(-1);
System.out.println(select.getSelectedOptions()); // [] -- empty, not [Germany]
select.setSelectedIndex(2);
System.out.println(select.getSelectedOptions()); // [Japan]
getSelectedIndex() / setSelectedIndex(int) give you the
same information as a plain integer index into getOptions() rather than
a list of HtmlOption objects; getSelectedIndex() returns
-1 when nothing is selected, matching the empty-list case above.
Setting a negative or out-of-range index simply results in nothing being selected.
HtmlSelect: the size attribute and default selection
getSize() returns the parsed size attribute as an
int, defaulting to 0 if the attribute is absent or isn't a
valid non-negative integer. This intentionally diverges from the HTML specification
itself, which defines the default rendered size as 1 for a
single select and 4 for a multiple select when the attribute is absent
-- getSize() is a raw attribute accessor, not a rendering computation, so
treat a return value of 0 as "no explicit size was set", not as a claim
about how the control would actually be drawn.
Whether a select gets a default selection at all depends on both
isMultipleSelectEnabled() and this size value: HtmlUnit auto-selects the
first option, once, when the page finishes loading, only if the select is in
single-select mode and its size is 1 or less (i.e. it renders as
a dropdown, not a scrollable list box) and nothing was already explicitly
selected in the markup.
HtmlSelect and required
A <select required> follows the same
isValueMissingValidityState() mechanism described earlier, but the rule
for what counts as "missing" is specific to how a select's default selection works.
It only applies to a single-select select with size 1
or less (the same rendering condition that triggers auto-selection, described above):
the value is considered missing if the option that ends up selected is the
first option in the list, and that first option's effective value is the
empty string, with nothing else explicitly selected instead.
<select required>
<option value=""></option>
<option value="a">Option A</option>
</select>
For the select above, with nothing explicitly selected, the auto-selected first
option has an empty value -- so isValid() is false and
checkValidity() reports a missing value, exactly like a browser's native
validation would show a "please select an item" message. Explicitly selecting the
second option (either in markup with selected, or via
setSelectedIndex(1)) clears the missing-value state immediately.
A multiple-select, or a single-select with size greater than
1, is only considered to have a missing value if literally nothing
is selected at all -- there is no "empty first option" special case for those, since
browsers don't auto-select anything for them in the first place.

