8 NumPy 1D-Arrays

(CSE331) Python for Data Science

Author
Affiliation

Md Rasel Biswas

IASDS, University of Dhaka

NumPy (Numerical Python) provides arrays and tools for numerical computing. Its array operations let us calculate with many values at once, usually without writing Python loops.

NumPy is a third-party package and is commonly included with Anaconda. Use the installation instructions from Lecture 7 if it is unavailable. By convention, import it as np:

import numpy as np

1 Creating a One-Dimensional Array

The main NumPy object is the n-dimensional array, or ndarray. A 1D array stores a sequence of values; a 2D array arranges values in rows and columns.

my_list = [4, 1, 7, 3, 5]
my_array = np.array(my_list)

print(my_list)
print(my_array)
print(type(my_array))
[4, 1, 7, 3, 5]
[4 1 7 3 5]
<class 'numpy.ndarray'>

For large numerical collections, NumPy arrays often use less memory and support faster calculations than Python lists. Each array has one data type, called its dtype.

1.1 Inspecting an array

print(my_array.ndim)   # 1: number of dimensions
print(my_array.shape)  # (5,): length of each dimension
print(my_array.size)   # 5: total number of elements
print(my_array.dtype)  # integer dtype; exact width can depend on the system
1
(5,)
5
int64

These are attributes, so no parentheses are needed. The comma in (5,) identifies a one-item tuple. For a 1D array, len(my_array) and my_array.size are equal.

2 Data Types and Conversion

NumPy chooses a common dtype from the supplied values, or we can specify one.

integers = np.array([8, 4, 5, 2])
decimals = np.array([8, 4, 5, 2], dtype=float)
mixed_numbers = np.array([1, 2.5, 3])
print(integers.dtype, decimals.dtype)
print(mixed_numbers)  # integers converted to floating-point values
int64 float64
[1.  2.5 3. ]

Assigning a new value does not automatically change the array’s dtype:

integers[2] = 7.9
print(integers)  # [8 4 7 2]: 7.9 was truncated to 7
[8 4 7 2]

Use .astype() to convert the array before storing fractional values:

converted = integers.astype(float)
converted[2] = 7.9
print(converted)
print(integers)  # original array remains integer-valued
[8.  4.  7.9 2. ]
[8 4 7 2]

Converting afterward cannot recover a fractional part that has already been discarded. A value that cannot be converted, such as "hello" assigned to an integer array, raises an error.

3 Indexing and Slicing

Indices start at 0. A negative index counts from the end. A slice excludes its stop position.

x = np.array([10, 20, 30, 40, 50, 60])
print(x[0], x[-1])
print(x[1:4])  # [20 30 40]
print(x[::2]) # [10 30 50]
print(x[::-1])
10 60
[20 30 40]
[10 30 50]
[60 50 40 30 20 10]

Assign to a position or slice to change existing values:

x[0] = 15
x[1:3] = [25, 35]
print(x)
[15 25 35 40 50 60]

3.1 Views and copies

Unlike a Python list slice, a basic NumPy slice shares data with the original array. Such an array is called a view.

a = np.array([10, 20, 30, 40])
part = a[1:3]
part[0] = 99
print(a)  # [10 99 30 40]
[10 99 30 40]

Use .copy() when changes should be independent:

separate = a[1:3].copy()
separate[0] = -1
print(separate)  # [-1 30]
print(a)        # unchanged by the preceding assignment
[-1 30]
[10 99 30 40]

b = a gives another name for the same array; b = a.copy() creates independent numerical data.

4 Elementwise Arithmetic

A list comprehension applies a calculation to each item:

values = [4, 1, 7, 3, 5]
print([5 * value for value in values])
[20, 5, 35, 15, 25]

An array expresses the same calculation directly:

a = np.array(values)
print(5 * a)
print(a + 100)
print(a**2)
[20  5 35 15 25]
[104 101 107 103 105]
[16  1 49  9 25]

This is vectorized arithmetic: NumPy applies the operation to the elements. Recall that 5 * values repeats a Python list.

4.1 Two arrays

Arrays with the same shape combine element by element:

a = np.array([1, 4, 3])
b = np.array([5, 8, 2])
print(a + b)
print(a - b)
print(a * b)
print(a / b)
[ 6 12  5]
[-4 -4  1]
[ 5 32  6]
[0.2 0.5 1.5]

For two 1D arrays, their lengths must match, or one must have length 1. A scalar also works. These are simple cases of broadcasting, covered in Lecture 9.

print(a + np.array([10]))  # [11 14 13]
[11 14 13]

Lengths 3 and 4 are incompatible:

np.array([2, 1, 4]) + np.array([3, 9, 2, 7])
---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[15], line 1
----> 1 np.array([2, 1, 4]) + np.array([3, 9, 2, 7])

ValueError: operands could not be broadcast together with shapes (3,) (4,) 

This example deliberately produces a ValueError.

5 Creating Special Arrays

print(np.zeros(5))
print(np.ones(5))
print(np.zeros(5, dtype=int))
[0. 0. 0. 0. 0.]
[1. 1. 1. 1. 1.]
[0 0 0 0 0]

zeros() and ones() use a floating-point dtype by default.

Function Specify Endpoint behaviour
np.arange(start, stop, step) Step size Intended to exclude stop
np.linspace(start, stop, num) Number of values Includes stop by default
print(np.arange(2, 10, 2))
print(np.arange(2, 4, 0.25))
print(np.linspace(2, 4, 11))
[2 4 6 8]
[2.   2.25 2.5  2.75 3.   3.25 3.5  3.75]
[2.  2.2 2.4 2.6 2.8 3.  3.2 3.4 3.6 3.8 4. ]

Floating-point steps in arange() can produce rounding surprises. Use linspace() when the required number of points and endpoints are important. Set endpoint=False if the final endpoint should be excluded.

6 Array Functions

6.1 Summaries

x = np.array([3.2, 4.8, 8.7, 8.7, 6.4, 5.3])
print("Sum:", np.sum(x))
print("Product:", np.prod(x))
print("Minimum and maximum:", np.min(x), np.max(x))
print("Mean and median:", np.mean(x), np.median(x))
print("Index of minimum:", np.argmin(x))
print("Index of maximum:", np.argmax(x))
print("Distinct values:", np.unique(x))
Sum: 37.099999999999994
Product: 39435.33772799999
Minimum and maximum: 3.2 8.7
Mean and median: 6.183333333333333 5.85
Index of minimum: 0
Index of maximum: 2
Distinct values: [3.2 4.8 5.3 6.4 8.7]

argmin() and argmax() give positions, not values. In a 1D array, a tie returns the position of the first occurrence. unique() returns sorted distinct values for these numerical arrays.

Many summaries are also array methods: x.mean() and np.mean(x) give the same result here.

6.2 Standard deviation: check the denominator

NumPy’s default standard deviation uses denominator \(n\). For the usual sample standard deviation, use ddof=1 to obtain denominator \(n-1\):

\[ s = \sqrt{\frac{\sum_{i=1}^{n}(x_i-\bar{x})^2}{n-1}}. \]

print("SD using n:", np.std(x))
print("Sample SD:", np.std(x, ddof=1))
SD using n: 2.0128062223892513
Sample SD: 2.2049187437787054

R’s sd() uses the sample convention. At least two observations are required for ddof=1.

6.3 Elementwise functions

These functions transform individual values:

positive = np.array([1.0, 2.0, 4.0])
print(np.sqrt(positive))
print(np.exp(positive))
print(np.log(positive))  # natural logarithms
print(np.round(np.log(positive), 2))
[1.         1.41421356 2.        ]
[ 2.71828183  7.3890561  54.59815003]
[0.         0.69314718 1.38629436]
[0.   0.69 1.39]

For finite real logarithms, use positive inputs. np.log(0) gives -inf and negative real inputs give nan, with warnings. Round for display when appropriate; retain full precision in intermediate calculations.

7 Comparisons and Boolean Masks

A comparison returns one Boolean value for each element:

values = np.array([4, 7, 6, 3, 9, 8])
mask = values > 5
print(mask)
print(values[mask])  # [7 6 9 8]
[False  True  True False  True  True]
[7 6 9 8]

Selecting values with a Boolean array is Boolean masking. For this 1D selection, the mask must have the same length as the array.

print(np.sum(mask))   # 4: count of True values
print(np.mean(mask))  # 4/6: proportion meeting the condition
print(np.any(mask))   # True if at least one value meets the condition
print(np.all(mask))   # True if every value meets the condition
4
0.6666666666666666
True
False

The proportion interpretation assumes a nonempty array. A Boolean counts as 1 for True and 0 for False in these calculations.

7.1 Combining conditions

Use &, |, and ~ on Boolean arrays. Put each comparison in parentheses.

Operation Example
AND (values > 5) & (values < 9)
OR (values < 4) | (values > 8)
NOT ~(values > 5)
print(values[(values > 5) & (values < 9)])
print(values[(values < 4) | (values > 8)])
print(values[~(values > 5)])
[7 6 8]
[3 9]
[4 3]

Do not combine array comparisons with Python’s and or or, or write 5 < values < 9. A multi-element Boolean array also cannot be used directly as an if condition; reduce it with np.any() or np.all() when one Boolean decision is needed.

7.2 Selecting and updating

scores = np.array([72, -1, 0, 84, -1])
selected = scores[scores >= 0]
print(selected)

scores[scores < 0] = 0
print(scores)
[72  0 84]
[72  0  0 84  0]

Boolean selection creates a copy. Direct indexed assignment, as in the last operation, changes the original array. Here replacing negatives by zero is only an assignment demonstration; a missing score should not normally be treated as a recorded zero.

8 Selecting by Position: Fancy Indexing

Supply a list or array of integer positions to select items in a chosen order:

x = np.array([10, 20, 30, 40, 50, 60, 70, 80, 90])
print(x[[6, 3, 8]])  # [70 40 90]
[70 40 90]

Integer-array selection creates a copy, just like Boolean selection. Basic slicing, such as x[1:4], creates a view.

9 Worked Example: Comparing Groups

The category and value arrays describe the same observations in the same order.

category = np.array(["A", "B", "A", "C", "B", "A", "C", "A"])
value = np.array([8, 6, 3, 1, 10, 12, 6, 4])

group_a = value[category == "A"]
print("Group A values:", group_a)
print("Group A count:", group_a.size)
print("Group A mean:", np.mean(group_a))
print("Group A values above 5:", value[(category == "A") & (value > 5)])
Group A values: [ 8  3 12  4]
Group A count: 4
Group A mean: 6.75
Group A values above 5: [ 8 12]

The mean is 6.75. A combined mask can select values or be summed to count observations. Check that a selected group is nonempty before calculating its mean.

10 Missing Numerical Values: A First Look

Use np.nan to represent missing values in a floating-point array. It may also appear after an undefined numerical calculation.

recorded = np.array([72, np.nan, 0, 84])
print(np.isnan(recorded))
print(np.mean(recorded))     # nan
print(np.nanmean(recorded))  # 52.0: excludes nan, retains zero
[False  True False False]
nan
52.0

Use np.isnan(), not == np.nan, to identify these entries. An integer array cannot store np.nan; create or convert to a floating-point array first. If all entries are nan, nanmean() returns nan with a warning.

11 Exercises

  1. Create x = np.array([12, 7, 18, 5, 9, 14]). Display its shape, size, dtype, first element, last element, and first three elements.
  2. Create separate arrays containing x + 5, x**2, and x / 2.
  3. Create the even integers from 2 through 20 using arange(), then create 9 evenly spaced values from 0 to 1 inclusive using linspace().
  4. Calculate the mean, median, and sample standard deviation of x. Find the index of its maximum.
  5. Select values in x that are at least 8 and below 15. Count them and find their proportion.
  6. Select the last, first, and third elements of x in that order using integer indices.
  7. Create a slice of x and modify it. Repeat using .copy() and explain the difference.
  8. For the worked group data, find the mean of group B and count observations in group A or C whose value is at least 6.
  9. For y = np.array([10, np.nan, 14, 0, np.nan]), count missing entries and compute the mean of recorded values.

12 Summary

  • Array arithmetic is elementwise; inspect shape and dtype before calculations.
  • Basic slices share data; use .copy() for independent data.
  • Use ddof=1 for the usual sample standard deviation.
  • Boolean masks select observations; &, |, and ~ combine conditions.
  • Distinguish missing values from zero.

13 Further reading