go-02-word-frequency
1.000
Challenge Ā· difficulty 2/5
# Word frequency
Implement **`solution.go`** in `package challenge` exporting:
```go
func WordFrequency(text string) map[string]int
```
Count how many times each word occurs in `text` and return the counts in a map.
Rules:
- Split `text` into tokens on **whitespace** (spaces, tabs, newlines).
- For each token, strip any **surrounding ASCII punctuation** (leading and trailing).
Punctuation in the middle of a token is kept.
- **Lowercase** each word before counting.
- If, after stripping, a token is empty, skip it (do not count an empty string).
- For empty or whitespace-only input, return an **empty, non-nil** map (length 0).
ASCII punctuation is the set of characters for which Go's `unicode.IsPunct` returns true
together with symbols such as `+`, `<`, `=`, etc. For this challenge, treat a byte as
"punctuation to strip" when it is an ASCII byte that is **not** a letter or digit.
Examples:
- `WordFrequency("the cat sat on the mat")` ā
`{"the": 2, "cat": 1, "sat": 1, "on": 1, "mat": 1}`
- `WordFrequency("Hello, hello! HELLO.")` ā `{"hello": 3}`
- `WordFrequency("don't stop")` ā `{"don't": 1, "stop": 1}` (interior apostrophe kept)
- `WordFrequency(" ")` ā `{}` (empty, non-nil map)
tests/solution_test.go
package challenge
import (
"reflect"
"testing"
)
func TestWordFrequency(t *testing.T) {
cases := []struct {
name string
in string
want map[string]int
}{
{
name: "simple repeats",
in: "the cat sat on the mat",
want: map[string]int{"the": 2, "cat": 1, "sat": 1, "on": 1, "mat": 1},
},
{
name: "punctuation and case",
in: "Hello, hello! HELLO.",
want: map[string]int{"hello": 3},
},
{
name: "interior apostrophe kept",
in: "don't stop don't",
want: map[string]int{"don't": 2, "stop": 1},
},
{
name: "tabs and newlines as whitespace",
in: "a\tb\nc a",
want: map[string]int{"a": 2, "b": 1, "c": 1},
},
{
name: "leading and trailing punctuation stripped",
in: "(go) [go]; {GO}",
want: map[string]int{"go": 3},
},
{
name: "digits are words",
in: "42 42 forty-two",
want: map[string]int{"42": 2, "forty-two": 1},
},
{
name: "token that is only punctuation is skipped",
in: "hi --- !!! bye",
want: map[string]int{"hi": 1, "bye": 1},
},
{
name: "empty input",
in: "",
want: map[string]int{},
},
{
name: "whitespace only",
in: " \t \n ",
want: map[string]int{},
},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
got := WordFrequency(c.in)
if got == nil {
t.Fatalf("WordFrequency(%q) returned nil map, want non-nil", c.in)
}
if !reflect.DeepEqual(got, c.want) {
t.Errorf("WordFrequency(%q) = %v, want %v", c.in, got, c.want)
}
})
}
}
func TestWordFrequencyEmptyIsNonNil(t *testing.T) {
got := WordFrequency("")
if got == nil {
t.Fatal("WordFrequency(\"\") = nil, want non-nil empty map")
}
if len(got) != 0 {
t.Errorf("WordFrequency(\"\") = %v, want empty map", got)
}
}
System prompt
You are an expert programmer. Solve the task exactly as specified. Output your solution as fenced code blocks using the required file name(s) and the exact function/type signatures requested. Prefer correctness; do not include prose outside code unless asked.Proposed solution
To solve this problem, we need to count the frequency of each word in a given text string, following specific rules for tokenization and punctuation handling. The solution involves splitting the text into tokens, processing each token to remove surrounding punctuation, converting tokens to lowercase, and counting their occurrences.
### Approach
1. **Tokenization**: Split the input text into tokens using whitespace as the delimiter. This ensures that each word is separated correctly, even if there are multiple spaces, tabs, or newlines.
2. **Punctuation Handling**: For each token, strip any leading and trailing punctuation. Punctuation is defined as any ASCII character that is not a letter or digit.
3. **Lowercasing**: Convert each processed token to lowercase to ensure case-insensitive counting.
4. **Counting**: Use a map to count the occurrences of each processed token. If a token becomes empty after processing, it is skipped.
### Solution Code
```go
package challenge
import (
"strings"
)
func isAlphanum(c rune) bool {
return (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9')
}
func WordFrequency(text string) map[string]int {
frequency := make(map[string]int)
tokens := strings.Fields(text)
for _, token := range tokens {
start := 0
n := len(token)
for start < n && !isAlphanum(rune(token[start])) {
start++
}
end := n - 1
for end >= 0 && !isAlphanum(rune(token[end])) {
end--
}
if start > end {
continue
}
trimmed := token[start:end+1]
lower := strings.ToLower(trimmed)
frequency[lower]++
}
return frequency
}
```
### Explanation
1. **Tokenization**: The `strings.Fields` function is used to split the input text into tokens, which handles any whitespace and trims empty strings.
2. **Punctuation Handling**: Each token is processed to remove leading and trailing non-alphanumeric characters. This is done by iterating from the start and end of the token until alphanumeric characters are found.
3. **Lowercasing**: The processed token is converted to lowercase to ensure that words like "Hello" and "hello" are counted as the same word.
4. **Counting**: A map is used to keep track of the frequency of each word. If a token is empty after processing, it is skipped, ensuring only valid words are counted.
This approach efficiently handles the requirements and ensures that the solution is both correct and optimal.Test output (stdout)
ped","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205381612+02:00","Action":"run","Package":"challenge","Test":"TestWordFrequency/digits_are_words"}
{"Time":"2026-07-08T00:25:34.205382422+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/digits_are_words","Output":"=== RUN TestWordFrequency/digits_are_words\n"}
{"Time":"2026-07-08T00:25:34.205383742+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/digits_are_words","Output":"--- PASS: TestWordFrequency/digits_are_words (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205385062+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequency/digits_are_words","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205386012+02:00","Action":"run","Package":"challenge","Test":"TestWordFrequency/token_that_is_only_punctuation_is_skipped"}
{"Time":"2026-07-08T00:25:34.205386842+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/token_that_is_only_punctuation_is_skipped","Output":"=== RUN TestWordFrequency/token_that_is_only_punctuation_is_skipped\n"}
{"Time":"2026-07-08T00:25:34.205388012+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/token_that_is_only_punctuation_is_skipped","Output":"--- PASS: TestWordFrequency/token_that_is_only_punctuation_is_skipped (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205389252+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequency/token_that_is_only_punctuation_is_skipped","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205390192+02:00","Action":"run","Package":"challenge","Test":"TestWordFrequency/empty_input"}
{"Time":"2026-07-08T00:25:34.205390992+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/empty_input","Output":"=== RUN TestWordFrequency/empty_input\n"}
{"Time":"2026-07-08T00:25:34.205392142+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/empty_input","Output":"--- PASS: TestWordFrequency/empty_input (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205394112+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequency/empty_input","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205395042+02:00","Action":"run","Package":"challenge","Test":"TestWordFrequency/whitespace_only"}
{"Time":"2026-07-08T00:25:34.205395872+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/whitespace_only","Output":"=== RUN TestWordFrequency/whitespace_only\n"}
{"Time":"2026-07-08T00:25:34.205396982+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency/whitespace_only","Output":"--- PASS: TestWordFrequency/whitespace_only (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205398952+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequency/whitespace_only","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205400072+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequency","Output":"--- PASS: TestWordFrequency (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205401262+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequency","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205402302+02:00","Action":"run","Package":"challenge","Test":"TestWordFrequencyEmptyIsNonNil"}
{"Time":"2026-07-08T00:25:34.205403282+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequencyEmptyIsNonNil","Output":"=== RUN TestWordFrequencyEmptyIsNonNil\n"}
{"Time":"2026-07-08T00:25:34.205404572+02:00","Action":"output","Package":"challenge","Test":"TestWordFrequencyEmptyIsNonNil","Output":"--- PASS: TestWordFrequencyEmptyIsNonNil (0.00s)\n"}
{"Time":"2026-07-08T00:25:34.205407442+02:00","Action":"pass","Package":"challenge","Test":"TestWordFrequencyEmptyIsNonNil","Elapsed":0}
{"Time":"2026-07-08T00:25:34.205408732+02:00","Action":"output","Package":"challenge","Output":"PASS\n"}
{"Time":"2026-07-08T00:25:34.205620323+02:00","Action":"output","Package":"challenge","Output":"ok \tchallenge\t0.001s\n"}
{"Time":"2026-07-08T00:25:34.205638563+02:00","Action":"pass","Package":"challenge","Elapsed":0.001}
Reasoning (chain-of-thought)
Okay, I need to solve this problem where I have to write a Go function called WordFrequency that takes a string and returns a map of word frequencies. Let's see the requirements.
First, the function is supposed to split the text into tokens based on whitespace. So I'll need to split the string into words wherever there's a space, tab, newline, etc. In Go, I can use strings.Fields, which splits on any whitespace and handles multiple spaces correctly.
But wait, the problem says to split on whitespace, so maybe I should use strings.Split with a space as the delimiter, but that might not handle all whitespace correctly. Oh right, strings.Fields is better because it splits on any whitespace and trims the results, so empty strings from multiple spaces won't be included. Hmm, but I need to process each token, so maybe I should split into tokens and then process each one.
Next, for each token, I need to strip any surrounding ASCII punctuation. So leading and trailing punctuation should be removed, but punctuation in the middle stays. How do I determine what's punctuation? The problem says to consider a byte as punctuation if it's an ASCII byte that's not a letter or digit. So I can't just use unicode.IsPunct because that includes some non-ASCII punctuation as well. Wait, the problem says to treat a byte as punctuation to strip if it's an ASCII byte that is not a letter or digit. So for each character in the token, if it's an ASCII letter or digit, it's kept; otherwise, it's considered punctuation and stripped from the start and end.
So for example, the token "Hello," becomes "Hello" after stripping the comma. The token "don't" should remain "don't" because the apostrophe is in the middle.
So the plan is: for each token, process each character from the start until we find a letter or digit, then from the end backwards. So for the token, we'll create a new string that starts after the leading non-alphanumerics and ends before the trailing ones.
How to implement this? Maybe for each token, create a new string by trimming the leading and trailing non-alphanum characters.
In Go, I can loop through each character in the token, find the first index where the character is a letter or digit, then find the last such index, and take the substring from first to last+1.
Alternatively, I can use a regex to trim the leading and trailing non-word characters. Wait, but word characters in Go's regex are [a-zA-Z0-9_], but the problem says letters and digits, not underscores. Hmm, so maybe I need to create a custom regex that matches leading and trailing non-alphanum characters.
So for each token, I can apply a regex substitution to remove leading and trailing non-alphanum. The regex pattern would be something like ^[^a-zA-Z0-9]+ or [^a-zA-Z0-9]+$. So I can use a regex that matches these and replaces them with empty.
Wait, but in Go, the syntax is a bit different. So the pattern would be `^[^a-zA-Z0-9]+` and `[^a-zA-Z0-9]+$`. So I can use two separate Replace operations, or combine them into one.
Alternatively, I can write a single regex that matches both leading and trailing non-alphanum. Like `^[^a-zA-Z0-9]+|[^a-zA-Z0-9]+$`.
So for each token, I can do:
trimmed := regex.ReplaceAllString(token, pattern)
Wait, no, because the pattern is to remove those parts. So for each token, I can do:
trimmed := string.TrimSpace(token) // Wait, no, because that only removes whitespace. But I need to remove leading and trailing non-alphanum, not just whitespace.
So perhaps the regex approach is better.
Let me think: for each token, I can create a new string by trimming leading non-alphanum, then trimming trailing non-alphanum.
So for example, token is "!!!hello!!!" ā becomes "hello".
Another example: token is "a.b" ā remains "a.b".
So the steps for each token:
1. Trim leading non-alphanum characters.
2. Trim trailing non-alphanum characters.
How to implement this in Go.
Option 1: Iterate from the start until a letter or digit is found, then from the end backwards.
Let's see. For a given token, say s:
start := 0
for start < len(s) && !isAlphanum(s[start]) {
start++
}
end := len(s) - 1
for end >= 0 && !isAlphanum(s[end]) {
end--
}
if start > end {
// the token is empty after trimming
} else {
trimmed = s[start:end+1]
}
Where isAlphanum is a helper function that checks if a character is a letter or digit.
Yes, that's a possible approach.
So I'll write a helper function:
func isAlphanum(c byte) bool {
return (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9')
}
Wait, but in Go, strings are accessed as characters, which are of type rune, which is an alias for int32. So perhaps I should cast the character to byte to check if it's ASCII.
Wait, but the problem says to consider only ASCII bytes. So perhaps I should check if the character is an ASCII letter or digit.
So the helper function could be:
func isAlphanum(c rune) bool {
return (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9')
}
Wait, but rune is an int32, so comparing directly with 'a' etc. is okay.
So for each token, I can loop through the start and end indices as described.
Once I have the trimmed token, I need to lowercase it. So the next step is to convert the trimmed string to lowercase.
Then, if the resulting string is empty, I skip it. Otherwise, I add it to the map.
So putting it all together:
The function will:
1. Split the input text into tokens using strings.Fields, which splits on any whitespace and returns a slice of tokens, ignoring empty strings.
Wait, no. Wait, strings.Fields splits the string into a slice, omitting empty tokens. So for example, " " will result in an empty slice. So that's good because in the example where the input is " ", the output is an empty map.
Wait, no. Let me check: strings.Fields(" ") returns an empty slice. So in that case, the for loop over tokens will not process anything, and the map will be empty, which is correct.
So the steps:
- Split the text into tokens using strings.Fields. So tokens = strings.Fields(text)
- For each token in tokens:
a. Trim
ā¦[4141 chars elided]ā¦
like "!!!", it's trimmed to empty, so it's skipped.
What about a token like "a1b2"? It remains as is, lowercased.
So the code should be correct.
Now, let's write the code.
First, the helper function isAlphanum.
But wait, in Go, functions can't be defined inside another function, so I need to define it outside. But since the solution is to write a function in solution.go, perhaps I can define the helper inside the same package.
So in solution.go, the code would be:
package challenge
import (
"strings"
)
func isAlphanum(c rune) bool {
return (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z') || (c >= '0' && c <= '9')
}
func WordFrequency(text string) map[string]int {
m := make(map[string]int)
tokens := strings.Fields(text)
for _, token := range tokens {
start := 0
n := len(token)
for start < n && !isAlphanum(rune(token[start])) {
start++
}
end := n - 1
for end >= 0 && !isAlphanum(rune(token[end])) {
end--
}
if start > end {
continue
}
trimmed := token[start:end+1]
lower := strings.ToLower(trimmed)
m[lower]++
}
return m
}
Wait, but in Go, rune is an alias for int32, and token is a string, so token[start] is a byte. So when I pass rune(token[start]), it's converting the byte to an int32, which is correct.
But wait, in Go, when you have a byte value, say 'a' (97), and you convert it to rune, it becomes 97 as an int32. So the comparison in isAlphanum is correct.
Testing this code with the examples should work.
Another thing to consider: what about Unicode letters? The problem says to treat only ASCII bytes as punctuation. So the code correctly ignores non-ASCII letters, but the problem says that the input is processed as per ASCII. So the code is correct.
Wait, but the problem says that the function should process the text as per ASCII punctuation. So the code correctly handles only ASCII letters and digits.
Yes.
So this should solve the problem.